Infinity Raises $15M to Build a CUDA Alternative for AI Inference

AI infrastructure startup Infinity has raised $15 million at a $100 million valuation, drawing capital from Touring Capital, Principal VC, and a roster of researchers from frontier labs OpenAI and Anthropic. The company is building kernel-level software that allows AI models to run efficiently on virtually any chip—GPUs, SRAM, mobile processors, or systolic arrays—without rewriting code for each architecture.

The bet is aimed squarely at Nvidia’s CUDA software stack, the de facto standard that has locked developers onto its hardware for years. Because popular AI frameworks like PyTorch and TensorFlow are built atop CUDA, most startups and enterprises default to Nvidia chips. Infinity’s answer is Ignition, an AI research agent that automatically generates, tests, and optimizes the low-level code needed for inference on non-Nvidia silicon, compressing what could be months of manual kernel engineering into hours or days.

Infinity’s first named customer is AI chipmaker (and Nvidia challenger) D-Matrix, and CEO Jeremy Nixon says the startup is already in discussions with other large chip and cloud companies. The company employs a performance-based pricing model—taking a cut of the speed gains and cost savings measured in tokens per second—rather than charging an upfront license fee.

Why CUDA's Dominance Is Starting to Crack

A Software Moat That’s Hard to Cross

Nvidia’s dominance is not just about its hardware; the real lock-in is CUDA. Because the vast majority of AI training and inference workloads now run through CUDA-optimized frameworks, any serious competitor must replicate that software environment or offer a compelling way to bypass it. Infinity is betting that automated kernel generation can make alternative chips “CUDA-compatible” without the years of hand-crafted integration that have stymied past efforts.

What OpenAI and Anthropic Researchers Bring to the Cap Table

The presence of individual researchers from OpenAI and Anthropic as investors is a notable signal. These labs are among the heaviest consumers of inference compute, and their technical endorsement—even in a personal capacity—suggests that demand for a cross-platform inference stack is real. It also hints that large AI developers are exploring ways to reduce dependency on a single chip supplier during inference, where cost per token can determine the economics of an entire product.

The Self-Optimizing Agent Advantage

Ignition is designed to continuously learn and improve, adapting to new chip architectures without manual intervention. If it delivers on the claimed speedup—reducing kernel development from months to hours—it could lower the barrier for niche chipmakers trying to compete in the inference market. Still, the startup must prove its technology at scale outside of a single customer and demonstrate that performance gains hold across a wide range of hardware, including the latest accelerators from AMD, Intel, and custom ASICs.

What Infinity’s Universal Inference Stack Means for Chipmakers and Cloud Providers

The funding and early customer traction give several groups a concrete signal to act on.

  • Chipmakers designing AI accelerators should initiate a technical evaluation of Ignition. Infinity’s claims around automated kernel generation could drastically cut the software porting costs that have historically prevented them from gaining inference share, and the D-Matrix partnership shows it’s already being tested in a real product pipeline.
  • Cloud providers negotiating with Nvidia can treat the emergence of universal inference libraries as a credible alternative. The fact that Infinity is already in talks with “large cloud companies” (disclosed by the CEO) means that several hyperscalers are actively exploring multi-vendor inference stacks. Engineering leads at those firms should benchmark Ignition against their in-house CUDA-to-other-chip translation layers to decide whether to co-develop or license.
  • AI application teams with tight inference budgets can monitor Infinity’s pricing model—a share of cost savings measured in tokens per second—as a direct signal of real-world performance gains. If early benchmarks show a measurable drop in cost per 1,000 tokens on non-Nvidia hardware, it will make commercial sense to add support for Infinity’s library into their deployment pipelines.

Risk & Opportunity Assessment

Commercial RiskMediumInfinity must prove that Ignition works reliably across dozens of chip architectures and can scale beyond its first named customer, D-Matrix. Investor appetite for AI infrastructure is strong, but without a second major design win, commercial traction remains unvalidated.
Competitive RiskMediumNvidia is deeply entrenched with CUDA and has vast resources to release its own cross-platform kernel tools or tighten its software ecosystem. Other startups are also pursuing universal inference layers, increasing the pressure to differentiate quickly.
Regulatory RiskLowThe company operates in the software layer for inference, not in a directly regulated sector. No antitrust or export-control developments directly threaten its model at this stage.
Reputation RiskLowAs a pre-revenue startup with limited public visibility, a technical setback would primarily affect future fundraising rather than a broad brand. The equity backing from known researchers provides some halo effect that could be quickly dented by under-delivery.
Technology DisruptionTransformationalIf Infinity succeeds, it would sever the tight coupling between AI inference frameworks and Nvidia hardware, allowing any accelerator maker to compete on cost and performance. That would fundamentally reshape the economics of AI deployment and potentially end CUDA's years-long status as a must-have stack.
Commercial OpportunityMediumA universal inference library addresses a real pain point for chipmakers and cloud providers. However, the $100 million valuation and modest raise suggest that the market is still waiting for broader proof points. Success hinges on converting early talks with large cloud companies into commercial contracts.