By María Constanza Herrera, For Bombellii Ventures
Companies are using more AI models across teams and tasks, raising a new question: which model should handle each task, at what cost, and using how much to compute?
That matters because compute is constrained and expensive to expand. So as AI usage grows, the opportunity is not only to add more capacity, but to use existing compute more efficiently.
Multi Model AI Is Becoming the Default
For most of 2023, “AI” and “ChatGPT” were almost interchangeable: only 1 in 100 enterprise users used a genAI app, and nearly all meant ChatGPT.
That changed quickly. Instead of converging around one winner, users began adding models alongside each other. Today, 85.9% of Claude users still use ChatGPT, while companies have gone from using nearly six genAI apps on average in early 2025 to roughly 15 by August 2025.
Thinking Machines Lab’s July 2026 launch of Inkling provides another signal of this trend. Founded by former OpenAI CTO Mira Murati and backed by $2 billion, the lab is not positioning their model Inkling simply as the most intelligent model available. Inkling is designed primarily as an open weight foundation for customization. Companies can download and fine tune the model, while Thinking Machines’ Tinker platform helps adapt it to proprietary knowledge and specific workflows. This event reinforces a broader shift from model supremacy toward model specialization.
The result is a multi model world where deciding which model should handle each task, and how much compute it deserves, becomes increasingly important.
Model Efficiency Alone Will Not Solve the Compute Problem
Beyond hardware and infrastructure, the compute problem is already being tackled from a different angle: model efficiency. This isn’t a quiet, academic corner of AI. It’s one of the most active spaces right now, with big companies already placing bets and a wave of startups competing hard for the same ground, most of the movement happening in just the last year or two.
Software optimization is reducing the resources required to customize models, making fine-tuning faster and less GPU-intensive. This niche picked up real momentum in 2024 and 2025, with the broader model-customization and fine-tuning market now expanding at roughly 18-23% a year.
New model architectures are trying to deliver more capability with less hardware, designed to run across cloud servers and edge devices rather than relying only on increasingly large models. This space has already attracted strategic backing from major chip makers. AMD, for example, led a funding round in 2025 in one of the leading companies pursuing this approach.
Cloud software is improving utilization of GPUs already installed in data centers, letting developers rent capacity on demand instead of buying machines or reserving capacity long term. This category, often called “neocloud,” took off through 2024 and 2025 and is projected to grow from about $20 billion in 2026 to roughly $180 billion by 2030. It already has serious backing: NVIDIA is an investor in most of the leading players, Microsoft is one of the largest customers, and Meta and SoftBank are now building competing GPU capacity of their own.
But one question remains: with so many models available, which should handle each task? Today, developers often make that choice once and keep it fixed. This is costly because leading models can cost five times more per token than lighter versions, even though routine tasks like classifying, summarizing, or extracting don’t need that firepower.
The approach is straightforward: build the middle layer that sits between the models and the users, giving companies control over which model runs what, and when, based on cost, quality, and what the job actually requires. We’ve seen this pattern before: every time a technology multiplies and scales, a new layer emerges to orchestrate it, and AI is no exception. That layer is what stops companies from overpaying for capability they don’t need, without sacrificing quality where it actually matters.
When a disruptive technology multiplies and scales, complexity grows with it, creating the need for a new layer to coordinate, control, and optimize the underlying infrastructure.
Disruptive Technology Creates a New Need for Orchestration
Server infrastructure followed this pattern first. In the early 2000s, as companies began virtualizing servers, VMware created a coordination layer for using compute more efficiently, and that startup alone grew into a company worth $13.35 billion in annual revenue by 2023, with a market cap of $61 billion by the time Broadcom acquired it.
Later, as applications spread across multiple clouds, servers, and services, companies lost visibility into how their systems were performing and where problems were emerging. An observability layer emerged to bring infrastructure, applications, logs, and security into a single view, with players like Datadog, New Relic, and Dynatrace leading the way. That market was worth about $2 billion in 2011 and had grown to around $9 billion by 2024, roughly 4.5x in just over a decade.
The pattern is similar: the underlying technology comes first, and the coordination layer becomes valuable once operating it at scale becomes difficult. AI has reached that same moment. Models, agents, data sources, teams, and providers are multiplying while compute remains constrained, creating the need for a new orchestration layer that can coordinate all of these moving pieces efficiently, reliably, and effectively for both companies and users.
AI Orchestration Is Becoming a Critical Layer
Four capabilities are emerging around this need: visibility into economics, observability of system behavior, model selection, and trust.
1. Economics & Accountability: Know what AI costs and what it produces
Companies need to understand how AI spending is distributed across teams, products, and models, and whether that spending generates valuable outcomes. This layer connects consumption with budgets, unit economics, and ROI.
Because AI FinOps is still part of the broader cloud-cost-management category, there is not yet a reliable standalone market estimate. The broader cloud FinOps market provides a useful proxy and is projected to grow from $14.9 billion in 2025 to $26.9 billion by 2030, representing annual growth of 12.6%.
Later-stage investment activity provides additional evidence. CloudZero raised a $56 million Series C to connect cloud and AI infrastructure costs with business outcomes, while Finout raised a $40 million Series C, bringing its total funding to $85 million. Both companies are expanding traditional FinOps capabilities to help enterprises allocate and optimize AI spending across models, teams, and products.
2. Performance & Reliability: Know how AI behaves and why it fails
Companies also need to understand how models and agents behave in production. This layer traces decisions, evaluates outputs, monitors latency and quality, and identifies where failures occur across models, prompts, data, and tools.
The market for agentic AI monitoring, analytics, and observability tools is estimated at $550 million in 2025 and projected to reach $2.05 billion by 2030, growing at approximately 30% annually.
Several companies have already reached later funding stages. Arize AI raised a $70 million Series C for its AI and agent evaluation and observability platform. Galileo raised a $45 million Series B to connect offline evaluations with production monitoring and guardrails, while Fiddler AI raised a $30 million Series C to develop a control plane for monitoring and governing compound AI systems.
These companies show that reliability is becoming more than a debugging function. It is becoming the feedback system through which enterprises evaluate, improve, and control AI behavior.
3. Routing & Optimization: Decide what runs
Different tasks require different levels of intelligence. Routing software can dynamically choose between models based on quality, cost, latency, privacy, availability, and task difficulty. Over time, these systems could learn from production outcomes and continuously improve the relationship between each task, model, cost, and result.
The model-routing market remains small but is expected to grow quickly. One estimate values it at approximately $101 million in 2025 and projects it could reach $4.1 billion by 2035, representing annual growth of almost 45%.
The clearest later-stage startup signal is OpenRouter, which raised a $113 million Series B in 2026. The company sits between AI applications and model providers, managing routing, reliability, cost optimization, and compliance across more than 400 models. Its platform now serves more than eight million developers. The limited number of pure routing companies beyond Series A indicates that this category is still developing, but OpenRouter’s scale suggests it is becoming an important part of multi-model infrastructure.
4. Security & Provenance: Trust what you run
As companies use more open models, third-party fine-tunes, APIs, datasets, and autonomous agents, they need to know where those systems came from, what they contain, whether they have been modified, and what resources they can access. This creates demand for model scanning, provenance, signing, access controls, runtime protection, and continuous verification.
The broader AI trust, risk, and security management market was valued at approximately $2.3 billion in 2024 and is projected to reach $7.4 billion by 2030, growing at 21.6% annually.
Later-stage funding provides strong commercial validation. Protect AI raised a $60 million Series B in 2024 to expand its platform for securing AI models and applications. The company was subsequently acquired by Palo Alto Networks in 2025. Noma Security raised a $100 million Series B, bringing its total funding to $132 million, to secure enterprise models, data, applications, and agents across the AI lifecycle.
Together, these signals suggest that trust will not be a separate compliance exercise. It will become an operating requirement embedded directly into the systems that select, deploy, monitor, and coordinate AI.
Smarter Model Allocation Can Turn AI Efficiency into Climate Impact
The industry is spending trillions of dollars expanding the supply of compute. AI orchestration represents the other side of that equation: the software layer that determines how efficiently that capacity is used. The opportunity is not simply to reduce token costs. It is to help companies generate more reliable, secure, and valuable outcomes from every model and every unit of compute they already have.
Among the four capabilities, we think routing and optimization represents the strongest investment opportunity. Economics and observability are becoming more established, while security is increasingly being integrated into broader cybersecurity platforms. Routing remains earlier and occupies a more strategic position: it sits in the path of every AI request and determines which model receives the workload. This creates recurring usage-based revenue, proprietary performance data, and a learning advantage that improves as more requests pass through the platform.
The opportunity also carries a meaningful climate benefit. Routing can reduce token consumption by up to 60%, inference energy by 31%, and CO₂ emissions per request by up to 74%, depending on the workload and baseline. The winning control layer could therefore improve both the economics and environmental efficiency of AI by ensuring that every task uses only the intelligence and compute it actually needs.
If you are building in the orchestration layer for multi model AI, we would love to hear from you.
References
BenchLM.ai. (2026, July). OpenAI API pricing (July 2026): Model & token costs. https://benchlm.ai/openai/api-pricing
Baseten. (2026). Baseten [Company website]. https://baseten.co
Built In. (2026). Why are millions of users leaving ChatGPT for Claude? https://builtin.com
CB Insights, PitchBook, & Wikipedia. (2026). Company and acquisition data: VMware, Broadcom, AirWatch, and Okta [Data set].
electroIQ / Business of Apps. (2026). Claude vs. ChatGPT statistics 2026. https://www.businessofapps.com
Fal.ai. (2026). Fal.ai [Company website]. https://fal.ai
Goldman Sachs Research. (2024). AI data center power demand outlook.
Google. (2025). Measuring the environmental impact of a Gemini prompt. https://blog.google
Liquid AI. (2026). Liquid AI [Company website]. https://liquid.ai
McKinsey & Company. (2024). Investing in the rising data center economy.
MIT News. (2024). AI has high data center energy costs, but there are solutions. Massachusetts Institute of Technology. https://news.mit.edu
Modal. (2026). Modal [Company website]. https://modal.com
Mordor Intelligence. (2026). Cloud FinOps market size and forecast. https://www.mordorintelligence.com
Netskope. (2024). Cloud and threat report: AI apps in the enterprise 2024. https://www.netskope.com
Netskope. (2025). Cloud and threat report: Generative AI 2025. https://www.netskope.com/resources/reports-guides/cloud-and-threat-report-generative-ai-2025
Netskope Threat Labs. (2025, August). Shadow AI risks proliferate as GenAI platforms and AI agents see rapid adoption. Netskope. https://www.netskope.com
Predibase. (2026). Predibase [Company website]. https://predibase.com
Replicate. (2026). Replicate [Company website]. https://replicate.com
RunPod. (2026). Company documentation. https://runpod.io
Sakana AI. (2026). Sakana AI. https://sakana.ai
Thinking Machines Lab. (2026, July). Inkling: Our open-weights model.
Unsloth. (2026). Unsloth AI [Company website]. https://unsloth.ai
Accenture. (2024, September 17). Accenture invests in Martian to bring dynamic routing of large language queries and more effective AI systems to clients.
Arize AI. (2025, February 20). Arize AI raises $70M Series C to build the gold standard for AI evaluation and observability.
Astute Analytica. (2026, July 22). AI model router market size, share, and forecast, 2026–2035.
Cisco. (n.d.). Robust Intelligence is now part of Cisco. Retrieved August 6, 2026.
CloudZero. (2025, May 28). CloudZero raises $56M Series C to redefine cloud cost optimization in the AI era.
Fiddler AI. (2026, January 27). Fiddler raises $30M Series C to deliver the first control plane for AI.
FinOps Foundation. (2026). State of FinOps 2026.
Finout. (2025, January 29). Finout’s Round C: A milestone moment for FinOps.
Galileo. (2024, October 15). Announcing our Series B and evaluation intelligence platform.
Grand View Research. (n.d.). AI trust, risk, and security management market size report. Retrieved August 6, 2026.
MarketsandMarkets. (n.d.). Cloud FinOps market: Global forecast to 2030. Retrieved August 6, 2026.
MarketsandMarkets. (2025, October 13). AI orchestration market worth $30.23 billion by 2030.
Mordor Intelligence. (2025). Agentic AI monitoring, analytics, and observability tools market: Growth, trends, and forecasts.
OpenRouter. (2026, May 28). OpenRouter raises $113M Series B.
Palo Alto Networks. (2025, July 22). Palo Alto Networks completes acquisition of Protect AI.
Portkey. (2026, February 19). Portkey raises $15M Series A to scale the unified control plane for production AI.
Reuters. (2025, July 31). Israeli cyber startup Noma Security raises $100 million in private funding round.
Weave. (2026, July 28). Weave raises $13.5M Series A to build engineering and token intelligence.