TLDRocket
Sign in

Touchmark wants to turn AI inference into a futures market

Tech Funding News Abhinaya Prabhu

Touchmark is launching a market for AI tokens, letting buyers prepay for future inference. It could make AI costs more like commodities, with prices, forwards, and resale.

Based on reporting by Tech Funding News, Abhinaya Prabhu — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Touchmark wants to do for AI inference what futures markets did for oil, power, and other hard-to-budget inputs: turn a messy, private negotiation into something with a quote. This week, the two-person startup from Y Combinator’s Summer 2026 batch is opening a forwards market to commercial buyers and sellers. A buyer expecting heavy usage in September can pay now and lock in tokens for that month at a discount of up to 30%, with larger orders getting direct quotes that can go further.

The unit is the token, not the server. On Touchmark, a billion tokens on a named model becomes a prepaid contract with a delivery window, then gets drawn down through an API when that window arrives. Buyers pick the model, the token volume, the provider, and the delivery window, then pay in full up front. The pitch is plain: if inference is becoming a commodity, it should start trading like one.

That pitch lands because budgeting for inference has become a headache. Enterprise AI spending roughly doubled year over year in 2026, reaching an average of about $1.2 million per organization, and 78 percent of IT leaders said they were hit with charges they had not budgeted for. Agents made the problem uglier. A single autonomous workflow can burn through far more tokens than a one-off query, and it does so quietly, over hours, which is bad news for anyone trying to forecast costs from a spreadsheet.

Touchmark is also setting up a request-for-quote board at launch. Buyers can specify model or models, volume, delivery window, throughput and latency floors, with requests denominated in tokens, dedicated GPUs, or reserved throughput. Providers then compete to quote against the request. The company says supply is skewing toward open-weight families such as Kimi, GLM, and Qwen, and points to OpenRouter routing data showing Chinese open-weight models crossing into a majority of the tokens it processes by the middle of this year, usually at lower prices than U.S. frontier models.

The provider side is the other half of the bet. Touchmark’s first inference provider is Wafer, described as one of the leading inference providers by throughput. For providers, the appeal is simple enough: get revenue now against capacity that won’t be served for months. Touchmark’s co-founders, Ilia Bolgov and Roman Yanushevskyi, both talk about the current market as fragmented, opaque, and negotiated one deal at a time. That complaint has support. Just this spring, CME Group and Silicon Data said they would bring compute futures to market, ICE and Ornn announced GPU compute futures, and Jensen Huang called AI compute a new asset class. Touchmark goes one layer higher, selling the output people actually use, while still leaving the awkward model risk in place. Its answer is contracts for best-in-class capacity that resolve to whichever model tops a public benchmark at delivery time. In other words: if you want a September token, you’re really betting on what the market will think intelligence is worth by then.

My take — AI-written commentary, not fact-checked reporting

This is the kind of financial engineering AI kept inviting the minute invoices stopped looking like invoices and started looking like weather. The industry loves to pretend models are the product, but money always finds the bottleneck first. Touchmark is right that if everyone is already treating inference like a commodity, somebody will eventually build the market structure around it.

Read more about this at: Tech Funding News

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.