TLDRocket
Sign in

Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

MarkTechPost Michal Sutter

Supabase open sourced Evals, a benchmark framework that tests AI coding agents like Claude Code and GPT models on real Supabase development tasks such as schema building and policy debugging. The benchmark uses three dimensions (products, topics, stages) with separate benchmark and regression test suites, combining deterministic checks with LLM-based scoring. The results show top models like Opus 5 and Kimi K3 achieve 100% on build tasks unaided, while skills improve smaller models' performance and reveal agents under-utilize declarative schemas and documentation.

Why it matters

Supabase has open sourced supabase/evals, an Apache-2.0 benchmark and framework that runs coding agents including Claude Code, Codex and OpenCode against real Supabase tasks — building schemas, debugging Edge Functions, fixing RLS policies — inside containerized stacks, then scores them with deterministic checks and LLM-as-a-judge. The post Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks appeared first on MarkTechPost.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.