TLDRocket
Sign in

Beyond the price per token: Choosing the right OpenAI model on Amazon Bedrock for your workload

Amazon Web Services Nick McCarthy

AWS says the cheapest model per token isn’t always the cheapest per job. OpenAI models on Bedrock won when the math included accuracy, extra turns, and rework.

Based on reporting by Amazon Web Services, Nick McCarthy — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

AWS is pushing a simpler idea: stop shopping by token price alone. In a new benchmark post, the company compares OpenAI models on Amazon Bedrock with two cheaper OpenAI API baselines, and argues that the real bill depends on whether the model gets the job done, how many tokens it burns getting there, and how many turns an agent needs before it stops looping.

The test setup is meant to be reproducible, not theatrical. AWS used one OpenAI Responses API code path across five models: gpt-5.6-luna, gpt-5.6-terra, and gpt-5.6-sol on Amazon Bedrock, plus gpt-5.4-mini and gpt-5.4-nano on the OpenAI API. The Bedrock models ran with reasoning disabled, while the API baselines kept their defaults, so this is a comparison of practical deployment choices rather than a pure model showdown.

On single-shot benchmarks, the higher tiers showed up clearly. Sol solved 75 percent of AIME problems, versus mini’s 37 percent, and also led on GPQA Diamond and MMLU-Pro. But the more striking shift came from efficiency. AWS says luna was already cheaper per correct AIME answer than mini at the original price, because it used fewer billed tokens with reasoning turned off. After the July 30, 2026 price cuts on Bedrock, AWS records luna at $0.0021 per correct AIME answer, versus $0.0139 for mini in this sample.

The agent test makes the point even more sharply. AWS ran 50 DeepSearchQA questions through live web search and fetch steps, with conversation history re-sent on every turn. Mini averaged 7.6 turns and ended up with 114k input tokens per question, about 2.3 times terra’s 50k. Terra still cost less per passing answer than mini, and luna was cheaper still at $0.05 per passing answer versus mini’s $0.40. Nano’s lower token price did not save it here.

For document work, the same pattern held. On a 48-task GDPval slice, all three gpt-5.6 configurations scored better than mini and nano, with luna passing 27 deliverables to mini’s 20. After repricing, AWS says luna also had the lowest observed cost per passing deliverable in the sample, at $0.010. The broader message is blunt: if a workload is judged on success, not just token count, the pricing page is only the starting point.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI buying that gets skipped because spreadsheets are comforting and wrong. Token price is a nice little lie when the model has to retry, search again, or meet a rubric a human would actually sign off on. The industry keeps pretending the cheapest model is the cheapest option; that usually ends the moment real work enters the chat.

Read more about this at: Amazon Web Services

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.