Why did we start Keld
Enterprise AI spending crossed a trillion dollars in 2025, and most teams couldn't tell you where it went. The average enterprise runs more than five AI models in production, pays invoices from a dozen providers, and has no single view of what any of it costs, let alone whether the spend was efficient and necessary. That isn't an AI problem. It's a visibility-and-pricing problem, and both halves are fixable.
Static rate cards bill urgent work and work that can wait four hours at the same price. Fragmented provider accounts hide spend across teams. And compute sits idle while jobs queue at premium prices.
Today we're launching Keld: map every AI workflow your company runs, then route the jobs that don't need a premium model or a real-time response to the best-value provider of the model you need.
- Global AI infrastructure spending is projected to hit $487 billion in 2026, up from $153B in 20242, yet average enterprise GPU utilization sits at just 5%3.
- Set a deadline and a price ceiling on each job; Keld matches it to the strongest qualifying model at the best price across 100+ providers, with a guaranteed backup if the first choice can't deliver in time.
- Keld Atlas maps your AI spend across teams, projects, models and providers: free, within a day of connecting, with no routing changes.
AI spend is becoming a balance-sheet problem
IDC projects global AI infrastructure spending will reach $487 billion in 2026, more than triple the $153 billion spent in 2024, and on track to pass $1 trillion by 20292. The capital is committed; the discipline to manage it isn't. Three structural failures drive the waste.
Static pricing. Most inference is billed at flat per-token rates built for a model's peak use case, but most enterprise jobs aren't peak-case work. Classifying a ticket or extracting metadata from a video can be resolved at full quality by near-frontier models that cost a fraction of frontier pricing. Paying list price for them is like overnighting a letter that could have gone standard post. Why the same inference can cost $5 or 5¢ →
Idle compute. Cast AI's 2026 report, covering 23,000+ production clusters, found average enterprise GPU utilization sits at just 5%3. Ninety-five cents of every dollar spent on AI infrastructure earns nothing. That is not a hardware shortage but a matching problem: spare capacity never reaches the job that could use it.
Undiscovered models. The best model for a job is often not the famous one. Independent providers run models that are genuinely exceptional at one thing (drafting copy, generating audio, pulling structure out of video) but most are invisible, so work defaults to a handful of frontier names and teams overpay for power they don't need.
How Keld works
The unit of work on Keld isn't a prompt. It's a job: a deadline (the most time it can take), a ceiling (the most you'll pay), and the model or use case you need. Keld runs it on the strongest model for the task at the best available price within your bounds. And if the first-choice provider can't deliver in time, the job escalates to a guaranteed backup, so it always lands on time, at or under your price.
It starts one level up. Before Keld optimizes anything, Keld Atlas maps every AI workflow you run (spend by team, project, model, provider and use case) and then flags the workloads that don't need a premium model or a real-time response. You decide what moves.
Underneath is a real-time, neutral exchange that matches incoming jobs against live supply from 100+ providers and clears each at the best price inside its deadline: the same market structure that prices equities, FX and energy. Neutrality is what holds it together: Keld earns no margin steering you to a favored provider, so the market clears on price, deadline and performance, never on who's paying us. And because it settles only on a job's requirements and price, the marketplace never sees your prompt data. How a marketplace differs from a router or gateway →
The Keld product suite
The inference market has no standard wire protocol, no neutral matching layer, and no unified view for the teams running the compute. Keld supplies all three, in a compact suite split by audience.
For enterprises, two products work together. Keld Atlas is the control plane: it maps your spend across six dimensions (team, project, model, provider, job category, capex/opex) from your existing usage, then runs the latency-tolerant jobs asynchronously, sending only a job's requirements to the matching engine and streaming the payload straight to the matched provider, never buffering or storing it. Integrations make Keld drop-in: if you use LiteLLM, LangChain or the OpenAI SDK, you connect over the open IXP protocol without rewriting anything.
For AI model providers, Keld Trade is where they list spare capacity and manage orders, with micro-batching that paces matched jobs into a fleet at exactly the utilization it can absorb: new revenue from otherwise-idle GPUs, without eroding direct-contract pricing.
Pay for speed only when you need it
Not all inference needs to be fast. That's the insight most enterprise AI teams haven't operationalized.
Deloitte projects inference will be two-thirds of all AI compute in 2026, up from a third in 20236. Real-time work (chatbots, live coding assistants) is a minority of volume but commands a premium. The majority is overnight enrichment, classification at scale, batch extraction: jobs that can wait minutes or hours yet bill at live-session rates. On Keld, a batch job submitted with a two-hour deadline runs on the best-value provider that fits the window, not the fastest one.
Built for enterprise control from day one
Every API credential, model connection and routing rule lives in Atlas. Budget caps and quotas are set per team and project, and spend is auditable across six dimensions from the moment you connect: no shared credentials, no provider accounts in individual engineers' environments, no billing emails to personal inboxes.
The IXP protocol is open source, so you can inspect the wire format or build custom adapters. Multi-model access is built in, not a negotiated bundle: because one neutral engine connects every provider to every enterprise, you reach providers far beyond your current contracts, and the catalog widens as new ones join. The recommended starting point is Atlas: connect your credentials and you'll have a mapped view of your AI spend within a day, with no routing changes and no commitment to async execution.
The infrastructure for enterprise AI isn't short of capital. It's short of structure and discipline. That gap is why Keld exists, and it's live today. Atlas is free to start.
Who's building Keld
Dave OttenCo-founder & CEO
Federico EnniCo-founder, Product & Strategy
Doug ShoreCo-founder, EngineeringKeld is built by Dave, Doug and Federico, three technology-industry veterans who scaled JW Player to thousands of enterprise customers, serving the world's biggest publishers and broadcasters at the intersection of video streaming and advertising. But the three of us are only the start. Keld is a growing, AI-first team of industry experts who have spent their careers building platforms at high scale and enterprise quality, across distributed systems, infrastructure, data, security and developer experience. We started Keld to bring that hard-won discipline to the fragmented world of AI.
Frequently asked questions
Is Keld a router or a gateway?
No. Gateways like LiteLLM or OpenRouter route across a fixed set of static integrations and add a markup, and they do that job well. Keld is a different thing: it routes each prompt into a live marketplace of independent model providers, a dynamic and ever-changing set of companies and people competing on price and quality. A fixed set of integrations is exactly the kind of plumbing that AI will eventually optimize away; an open, competitive marketplace of independent sellers is not.
Does Keld see my prompt data?
No. Keld Atlas handles only the routing and financial metadata (deadline, ceiling price, use-case category, estimated token count), which is all the matching engine needs to find your best price. The raw prompt payload streams directly from your infrastructure to the matched provider. Keld never buffers, stores, or processes prompt content at any point.
How do the ceiling price and deadline work?
Set a maximum price per token you'll pay and a deadline: "within two hours," "by end of day," whatever the job requires. Keld finds the best-priced option that meets both. If supply exists at or below your ceiling within your window, the job runs immediately. If none can deliver inside your window, the job escalates to your configured backup, typically a direct provider at list price, so it always lands on time.
How does Keld handle AI regulations like the EU AI Act or model-origin restrictions?
It turns them into bounds on the job. AI regulation is tightening: the EU AI Act's phased obligations, data-residency rules, and a growing patchwork of restrictions on where models run and where they originate. On Keld those aren't blockers; they're curation parameters. Alongside price, deadline and quality, you can constrain routing by jurisdiction, data residency, and model or provider origin, so every job only ever clears on models and providers that meet your obligations. You set the boundary once and Keld keeps each job inside it, adapting as the rules change; the marketplace already matches only on a job's requirements, never on your prompt content. In short: you work within your boundaries, and Keld does the enforcement.
Why does Keld run a marketplace at the inference level, not the GPU level?
Because you don't actually want a GPU: you want a summarized document, a classified ticket, a transcribed call, at a quality bar and before a deadline. The unit you consume, and the one your finance team sees on the invoice, is the token, not the GPU-hour. So Keld clears the market on units of model output rather than raw hardware. That choice does three things an infrastructure market can't: it makes offers directly comparable ("one million tokens of summarization" is one yardstick every provider competes on, whereas A100-hours versus H100-hours don't map cleanly to the result you need); it carries no operational burden, since the provider keeps the scheduling and idle-capacity risk while you just get the output; and it optimizes the thing that matters, routing to the model that's genuinely best for the task (even a specialist you've never heard of) instead of cheaper silicon running the same expensive model. Renting GPUs by the hour just moves where the meter runs; pricing the work moves the bill. More on why an inference marketplace →
- IDC, AI Infrastructure Spending Caps Historic Year at $90B in Q4 2025, Q1 2026, retrieved June 23, 2026, idc.com
- Cast AI, 2026 State of Kubernetes Optimization Report, Q1 2026, retrieved June 23, 2026, cast.ai
- Deloitte, TMT Predictions 2026: Why AI's Next Phase Will Demand More Compute, Dec 2025, retrieved June 23, 2026, deloitte.com