Cekura Bench
Product Hunt Garry Tan
Cekura launched Bench, a public test for voice AI phone calls. It scores 9 models on live calls, with transcripts open so the numbers can be checked.
Based on reporting by Product Hunt, Garry Tan — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Cekura’s fifth launch is trying to do something the voice AI crowd keeps promising and rarely makes easy to verify: show its work. Cekura Bench is a public benchmark for speech-to-speech models, built around live phone calls instead of polished demos and cherry-picked clips.
The setup is blunt. It tests nine realtime voice models, including GPT Realtime 2.1, Gemini Live, Grok and Phonic, as complete phone agents. Each model goes through 82 scenarios, and each scenario is run three times. That gives the ranking more texture than a single lucky call, and it also makes the failures harder to hand-wave away.
The scorecard is not just about sounding natural. Cekura says it ranks models on reliability, data accuracy, stalled calls, response time and cost. That mix matters, because a voice agent that talks smoothly but drops calls or mangles details is still a bad agent. And the public transcripts are the real hook here: anyone can read the call log and judge whether the benchmark is being fair.
Cekura Bench is one piece of a wider set of evaluations from the company. It also covers voice agent benchmarks and STT benchmarks, with TTS benchmarks coming soon. For a market full of glossy demos, the pitch here is simple: fewer vibes, more receipts.
My take — AI-written commentary, not fact-checked reporting
This is the right instinct. Voice AI needs less leaderboard theater and more public transcripts, because a cheerful bot that forgets the point of the call is just expensive performance art. The industry loves to talk about trust; here’s a way to measure it without the fog machine.
Read more about this at: Product Hunt