The Rise of the AI Cost-Cutting Industry

Manos Koukoumidis, CEO of Washington-based AI company Oumi, puts it bluntly: companies are using Swiss Army knives where scalpels are needed. Too many enterprises are running niche internal tasks on the most powerful — and most expensive — models from frontier labs such as Anthropic, OpenAI and Google, a practice he calls “widely irrational and inefficient.” The result is enterprise AI bills that have ballooned far beyond what the work actually requires.

That mismatch has given birth to a fast-growing service industry dedicated to slashing those bills. The players fall into three rough camps. The “coaches” are consultancies like Sydney-based Adaptovate, which has spent the past three years helping clients from a snacks manufacturer to a 30,000-person professional services firm figure out where AI actually pays. The “measurers” are platform companies such as San Francisco-based Larridin, which tracks how employee productivity responds to token spending so clients can find the efficient midpoint. And the “builders” are startups creating the underlying technology: Oumi lets employees assemble niche models in minutes, Runware supplies cheaper inference infrastructure, Tensormesh optimizes GPU caching, and Conifer, still in Y Combinator’s summer 2026 batch, plans to route each query to the cheapest suitable model.

The timing is no accident. After a phase of “tokenmaxxing” — in which employees were encouraged to burn as many AI tokens as possible — companies are now scrutinizing every dollar. That has fueled a wave of funding: Oumi raised $10 million in seed in 2024, Larridin raised $17 million from Andreessen Horowitz, Bloomberg Beta and Google Ventures, and Tensormesh raised $20 million from AMD and Nvidia’s venture arm NVentures. Larridin’s CTO Ameya Kanitkar says traction has doubled every quarter since the start of the year, a sign of how quickly the cost problem went from afterthought to boardroom priority.

How the Cost Savers Make Their Money — and What Is Still Uncertain

The End of Tokenmaxxing

The industry’s growth is best read as a hangover from the generative AI boom. When enterprises first deployed frontier models, the incentive structure rewarded usage: teams were told to experiment, burn tokens and discover use cases. That created enormous waste but little accountability. Now procurement and finance teams are asking whether hundreds of millions in AI spending is justified — the exact question the cost-saving vendors exist to answer. This is a familiar maturation pattern for enterprise technology: novelty spending eventually collides with procurement discipline.

Advertisement

Three Business Models, One Blind Spot

The diversity of approaches suggests enterprises do not actually know where the waste sits. Coaches like Adaptovate start with strategy — how decision-making, talent structures and organizational design must change before AI scales — which implies the problem is as much managerial as technical. Measurers like Larridin attack the visibility gap, plotting productivity against token budgets to find where spend stops producing results. Builders attack unit economics directly: Tensormesh’s key-value cache lets companies accept more queries on fewer GPUs, and Conifer’s plan to split queries across models including local ones is aimed squarely at customers “focused on cost,” in cofounder Michael Jeffords’ words. That all three models are finding customers is itself evidence that enterprises lack a single, clean answer.

The ROI Story Nobody Can Verify

The most honest part of the pitch comes from the vendors themselves. Kanitkar argues AI spending should be treated as capital expenditure, with payback expected over one to two years rather than quarters — a useful correction to boards expecting immediate returns. And Runware cofounder Ioana Hreninciuc concedes the market’s numbers are unreliable: “The incentive right now is for people to say it is paying off and to find a way to justify it, and because of this, we don’t have an accurate view of the market, as nobody wants to be the canary in the coal mine.” In plain terms, reported AI ROI figures are self-reported, unaudited and probably flattering. Buyers should treat vendor claims about savings with the same skepticism they would apply to any other unverified marketing.

What Companies Should Check Before Trimming AI Spend

For enterprises carrying large AI budgets, the story suggests a practical sequence, not a blanket spending freeze:

  • Audit model fit before cutting anything: Identify which tasks genuinely require frontier models from Anthropic, OpenAI or Google and route the rest to lighter, open-source or locally run models — the core premise behind Oumi and Conifer’s products.
  • Measure before you manage: Track productivity against token spend per team, as Larridin does for clients, to find where incremental tokens stop producing output rather than imposing an arbitrary cap.
  • Set realistic payback expectations: Treat AI outlays as capital expenditure with a one-to-two-year horizon, per Kanitkar’s argument, and communicate that timeline to boards before they expect quarterly returns.
  • Optimize infrastructure before buying more: If your company runs its own GPUs, evaluate caching and inference optimization — the Tensormesh and Runware approaches — to raise throughput per GPU before purchasing additional hardware.
  • Watch Conifer’s post-Demo Day progress as a read on investor appetite for this segment: the startup had raised $1.3 million, about 20% of its target, before Y Combinator’s summer 2026 Demo Day.

Risk & Opportunity Assessment

Commercial RiskMediumGrowth is real — Larridin says traction doubles quarterly — but vendors depend on enterprises admitting waste, which Runware's cofounder says few are willing to do publicly; a pullback in AI spending could hit the cost-savers as hard as their clients.
Competitive RiskHighThe sector is crowded with overlapping coaches, measurers and builders; frontier labs keep cutting prices and adding features, while large consultancies compete with Adaptovate on strategy, pressuring the startups' differentiation.
Regulatory RiskLowNo specific regulation governs AI cost-optimization services; the main exposure is indirect, via enterprise demand for data privacy in caching and model-routing tools, which is a commercial concern rather than a compliance one.
Reputation RiskMediumThe market's own founders admit reported ROI figures are unverifiable and that nobody wants to be the canary in the coal mine; if savings claims deflate, trust in the whole category could erode quickly.
Technology DisruptionHighThe cost-saving wedge could narrow if frontier labs aggressively cut token prices or if lighter open-source models improve fast enough that enterprises no longer need specialized routing, caching and measurement layers.
Commercial OpportunityHighEnterprises poured millions into AI without measurement or procurement discipline, and the shift away from tokenmaxxing creates immediate demand for control tools — evidenced by $50 million-plus in recent funding across Oumi, Larridin and Tensormesh.