TLDRocket
Sign in

OpenRouter launches Batch API across 70+ models for cheaper, waitable jobs

OpenRouter Blog

OpenRouter now lets people queue big AI jobs across 70+ models and pay less. Most batches finish fast anyway, so the 24-hour wait is more of a ceiling than a delay.

Based on reporting by OpenRouter Blog — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenRouter has turned on Batch API support for more than 70 models, giving users a cheaper way to run work that does not need an instant answer. The pitch is simple: send a batch, let the provider pick a time within a 24-hour window, and pay about half the usual per-token price — sometimes less.

The company says the slowest possible outcome is not the typical one. During a two-week beta, it saw more than 230,000 completed batches, with a median finish time of 7 minutes and 90% done within an hour. The 99th percentile stretched to 10.3 hours, but most jobs landed far sooner than the full day.

This is aimed at the kind of chores that do not deserve a live request-response loop: labeling text, back-filling embeddings, scoring eval sets, summarizing ticket backlogs, or running one prompt across thousands of rows overnight. The API accepts chat completions, responses, messages, and embeddings, and completed batches return their results inline. Users submit to api/v1/batches and poll GET /api/v1/batches/:id until the job is completed, failed, expired, or cancelled.

OpenRouter also says timing matters more than size in some cases. Batches submitted between 5am and noon Pacific were slower than other hours, while submissions after 6pm Pacific had a 90th percentile under 50 minutes. Even a single-request batch finished in 5 to 11 minutes depending on the hour, and batches of 1,000 or more requests completed in 12 to 21 minutes. For batches over 100 requests submitted between midnight and noon Pacific, the slowest tenth could take as long as 6.8 hours.

There are some guardrails. Each batch runs on a single provider, results come back per request so one bad row does not sink the job, and inputs and outputs are kept for 30 days unless the batch is deleted. Images and files must be public URLs, while audio, video, and OpenRouter’s own web search plugin are not supported in batch.

My take — AI-written commentary, not fact-checked reporting

This is the kind of boring infrastructure feature that quietly saves real money, which is why it matters. Everybody loves flashy model demos; far fewer people want to talk about queueing, retention, and provider routing, even though that is where the bills are paid. Batch APIs are what AI looks like once it stops pretending every prompt is an emergency.

Read more about this at: OpenRouter Blog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.