A Top Technology Initiative Article – Sept. 2026
This year, colleague Brian Tankerley and I have had the pleasure of assisting many firms and organizations refine their AI strategy. You may recall in my July column that I used a pyramid to describe three levels of AI Adoption: LLMs for productivity, built into products, and Agents with MCPs (Model Context Protocol).
Notable commercial transactions in the past 30 days have affected my strategies and recommendations to build more complex AI strategies. The prior recommendations aren’t wrong or unworkable, but I need you to consider the impact of transactions involving OpenRouter and n8n, to name two.
Consulting for technology strategies and AI strategies has been so strong this year, I wanted you to benefit from my thinking and new learning as the summer progressed. Additionally, speaking at various events on AI has given me the opportunity to hear questions directly from attendees when things weren’t clear, or when other presenters made statements I didn’t understand or that were just wrong.
Another item that needed clarity was the build vs. buy discussion. Firms building AI agents see that as a competitive advantage, knowing they will have long-term maintenance in addition to the initial development. The “buy” crowd will find useful and functional agents available in bundles. Some will provide many of the tasks that you need to do, while others will need modification.
Vendors of agent bundles will need to provide implementation services and charge ongoing monthly fees. In effect, you will be trading money for leverage for your team. You’ll see more providers announce new agent bundles after October 15, so I’ll cover more after the announcements. Right now, three interesting ones include Auras from 34Five.ai, TaxGPT CoWork, and Netsurit.
Brian and I have recorded podcasts on some of these topics, and I would invite you to listen to What Tokens Are You Smokin’? and Review of Microsoft Agent 365. On the Accounting Technology Lab podcast, we review leading products and talk with thought leaders; we recently had a delightful time speaking with Alexis Kingbury, who wrote Accrual Intentions and built a complete CAS firm using Agents.
So, What IS AI Plumbing? That’s Orchestration!
Simply put, AI Plumbing is the way we connect AI systems. Every major vendor is trying to supply an orchestration option. For the past few years, we recommended n8n, which took a significant $60M investment from SAP and recently changed its pricing. Earlier this year, Microsoft released Agent 365, its orchestration approach available in Microsoft E7 licensing.
All the big companies have orchestration: Anthropic, OpenAI, Google, Perplexity, Amazon, and more. And what does orchestration look like? Consider this diagram:
AI Orchestrator Architecture with Multi-Model Routing

Here’s the architecture explanation: the orchestrator runs workflows and decides whether to send a task to the private Ollama server or through the router to paid cloud models. How it works: the orchestrator holds the workflow logic, so it can send sensitive or high-volume, routine tasks (classification, summarizing internal docs, embeddings) to the Ollama server on your own hardware at no per-token cost. You can buy AI servers from Dell or Lenovo right now for around $5,500 to host Ollama. It sends tasks that need frontier reasoning, live web answers (Perplexity), or a specific vendor to the router instead. The router gives you one API and one bill, handles fallbacks when a provider is down, and meters the tokens used from each cloud model.
One important detail: Ollama itself doesn’t draw tokens from Claude, ChatGPT, or the others. It runs open-weight models locally. The cloud providers are only reached through the router or directly by the orchestrator. If you want the private server to sit in the same routing layer as the cloud models, a self-hosted gateway like LiteLLM or an AI router platform can register Ollama as another endpoint alongside the cloud APIs. OpenRouter only routes to hosted providers, so it can’t reach a private Ollama box.
Azure AI is also often connected directly rather than through OpenRouter. Enterprises that want Microsoft’s data residency and compliance terms usually call their own Azure deployment from the orchestrator. That’s especially common with Agent 365, since it lives in the Microsoft stack.
But coordinating agents is only the start. As your firm or AI usage grows, your AI costs will increase. Today, we are buying AI capabilities (tokens, the unit of consumption) below cost from most providers. As we have discussed before, tokens are being throttled or restricted, just like cellular data used to be. To overcome that issue, heavy users can easily consume all their tokens in the morning. So, users have been subscribing to multiple accounts to get more tokens.
However, there is a better way (and I don’t mean like the Mandalorian code of conduct and belief system, guiding their actions, loyalty, and identity, often expressed through the mantra “This is the Way.”!). I simply prefer lower-cost, independent systems that keep client data confidential.
The way? AI Routers
Our favorite AI Router was OpenRouter, but after Stripe’s $8 billion purchase, we are reticent to keep recommending it blindly. OpenRouter had 8 million users and 400 models available on the platform as of May, per Bloomberg. I did not say OpenRouter is a bad choice; it’s simply that using this platform now makes you dependent on Stripe’s charges and whims.
While I haven’t completely settled on my new favorite and recommended routing strategy, we have various options listed below. Most organizations choose Trusted Router or Factory AI, but you should do your own due diligence. For example, Meta is reportedly working on an OpenRouter competitor called Switchboard, and Ramp launched its own AI model router, called Router, on August 19.
AI Router Products
This table is representative of AI Routers as of the date of this column. Again, I encourage you to look for current solutions when you are ready to proceed. The bottom line is that AI Routers give you access to many different AI engines, likely reducing both your dependence on a single vendor and minimizing your costs.
| PRODUCT + WEBSITE | COST | KEY BENEFITS |
| OpenRouter openrouter.ai | Pay-as-you-go; 5.5% fee on credit purchases; free tier | 500+ models and 80+ providers; unified API and billing; fallbacks; strong developer ecosystem |
| Ramp Router router.com | Routing free through 2026; inference charges apply; future pricing unannounced | Benchmark-based routing; shadow testing; cost and latency dashboard; reported average savings of about 40% |
| TrustedRouter trustedrouter.com | Provider cost + 5.5%; no subscription; BYOK available | No prompt/output logs; E2EE and zero-retention routes; privacy- and jurisdiction-aware routing |
| Factory AI factory.ai | $20 Pro; $100 Plus; $200 Max monthly; Business/Enterprise custom | Agent-native software-development platform; model choice; ZDR, SSO and enterprise governance |
| Portkey AI Gateway portkey.ai | Free Developer; $49/month Production; Enterprise custom | Routing, observability, caching, guardrails, load balancing, and key management |
| LiteLLM litellm.ai | Open source; paid enterprise offerings | Self-hosted OpenAI-compatible gateway; broad model support; budgets, fallbacks, and enterprise controls |
| Cloudflare AI Gateway developers.cloudflare.com/ ai-gateway | Free and usage-based services; related Cloudflare services may add cost | Global edge network; analytics; caching; rate controls; provider resilience and security |
| Helicone helicone.ai | Free and paid usage-based plans | LLM observability, cost tracking, prompt management, caching, and gateway routing |
| Martian Router withmartian.com | Enterprise = contact sales | Automated model selection based on quality, cost, and latency; designed specifically for routing |
| TrueFoundry AI Gateway truefoundry.com | Enterprise = contact sales | Multi-model gateway with guardrails, governance, observability, and private cloud/Kubernetes deployment |
| Kong AI Gateway konghq.com | Enterprise = contact sales; limited open-source capabilities vary by package | Extends mature API management to AI traffic with plugins, security, governance, and traffic control |
| Amazon Bedrock aws.amazon.com/bedrock | Consumption-based model and service pricing | Managed access to multiple model families; AWS security, governance, agents, and knowledge-base integration |
| Microsoft Azure AI Foundry ai.azure.com | Consumption-based Azure model and service pricing | Unified model catalog and enterprise controls; Azure identity, networking, monitoring, and governance |
| Google Vertex AI Model Garden cloud.google.com/vertex-ai | Usage-based cloud pricing | Managed Gemini and third-party models; evaluation, deployment, governance, and Google Cloud integration |
So, What Should You Do Now?
A straightforward action is to start small with a well-known product. While we are all busy, investing one hour per week in a subscription AI product will improve your productivity. The safest and least capable AI product is Microsoft Copilot. On this platform, you can safely process documents containing client data. That is not true of any of the other $20/month products. Use a paid, not free, version of the following:
- ChatGPT (OpenAI)
- ChatGPT Plus at $20 / mo.
- ChatGPT Pro – $200 / mo.
- Microsoft Copilot 365 – $30 / mo.
- Anthropic (Claude)
- Claude Pro $20 / mo. ($17/mo. annual)
- Claude Max $100–$200 / mo.
- Google Gemini
- Google AI Pro (Gemini Advanced) $19.99 / mo.
- Google AI Ultra $1,199.88 / year (~$99.99/mo.)
- Perplexity
- Perplexity Pro $20 / mo.
- Perplexity Max $200 / mo.
Learn one of the 12+ prompting methods for AI from a CPE course like K2’s AI – Better Prompts, Better Results. This course will provide a prompting spreadsheet. Refine your own prompting model for your common tasks. Consider how you want to use Gen AI products to boost productivity for yourself and across the business. Then set a strategy for your business.
Do I buy products with AI built in? Again, refer to my July column that showed the three styles of implementing AI. Finally, consider whether any vendor sells agents that meet your needs. This is a “buy” strategy. If you can’t find what you need off-the-shelf, consider a build strategy, where Claude currently has the best option, and OpenAI is trying to catch up. If you are using many agents and want more sophistication, my model above will enable unlimited token budgets and the use of all popular AI models.
While this is a lot to consider, solve a few simple problems first, then tackle a hard, time-consuming problem and continue building a library of prompts that solve more and more problems. You’ll likely convert your prompts, also called chats, into repeatable projects in ChatGPT or Claude.
You will find that some of your chats have lasting value because you have asked a question that has surfaced a particularly valuable answer. I frequently refer back to older chats where I’ve used AI to answer a complex question. It’s really the same song, second verse. Can you hear the music?
Sign in to get access to this free resource, and all of our whitepapers and reports.
Download this content today!
Register Now Already registered? Click here to Log In