TLDRocket
Sign in

Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference

MarkTechPost Michal Sutter ● Covered by 2 sources

Architect launched Liquid Inference, a live auction for LLM requests. Your prompt gets bid on; the cheapest qualifying provider wins.

Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Architect Financial Technologies has rolled out Liquid Inference, an LLM router that turns every request into a live auction. The setup is simple on the surface: a developer swaps in a base URL, keeps the rest of the code, and lets providers compete to answer the prompt.

The company says providers post offers for specific models, and each request is matched against everyone quoting that model. The buyer’s rules do the filtering first. Then the lowest qualifying offer wins, and the max price is locked before the first token is generated. Billing is based on metered usage only.

That buyer control is the point. Users can set per-job cost caps, time-to-first-token limits, minimum throughput, approved regions, zero data retention, and allow lists for providers or models. There’s also an Auto mode that picks the model for a given unit of work. Harrison’s LinkedIn post adds routing-rule presets and full multi-modal support, while account holders can watch live order books, per-provider and per-model quotes, and cleared transactions. That kind of market data is not what people usually get from an LLM API.

The product comes from a trading firm, not an AI lab, and that background shows. Architect already runs the AX perpetual futures exchange, and in May 2026 it acquired a US Designated Contract Market to list GPU compute futures, pending regulatory review. The team is clearly borrowing exchange mechanics for inference: price discovery, order books, receipts, the whole apparatus. It’s a neat fit, even if it feels a little strange to see prompt routing dressed up like a trading floor.

Liquid Inference is already wired for OpenAI and Anthropic API calls, and the company says it works with tools like Claude Code, Codex, OpenCode, Cursor, Pi, and Cline. Providers register through REST and WebSocket APIs, can update quotes based on their own costs, and get paid through Stripe with itemized job records. The company also says new providers are verified in minutes, not weeks. First 500 users get $20 of free inference, and referrals earn free inference too, including 20% of referred fees.

My take — AI-written commentary, not fact-checked reporting

This is classic fintech brain meeting AI infrastructure, and that’s not an insult. The real story is that inference is getting treated like a market, which is probably healthier than pretending pricing will stay magical forever. Open ecosystems win when they make switching cheap and pricing visible; the rest is mostly marketing frosting.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.