Introducing the SWE-Lancer benchmark
OpenAI
OpenAI built a test asking: can AI actually earn real freelance coding money? Answer so far: not really, but it's closer than you'd think.
Based on reporting by OpenAI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI's newest benchmark skips the usual academic coding puzzles and goes straight for something messier: real freelance software engineering work, the kind posted on job boards with actual dollar values attached. SWE-Lancer, as they call it, poses one blunt question — could a frontier language model take on these gigs and walk away with a payout, the way a human contractor would?
The setup matters here. Instead of grading models on toy problems with clean right answers, OpenAI pulled tasks that carry real-world stakes, tallying up to a million dollars in potential earnings across the set. That's a deliberate choice. Freelance engineering work is full of ambiguity, half-written specs, legacy code nobody wants to touch, and clients who don't always know what they're asking for. It's a much rougher test than anything on a typical leaderboard.
What's notable isn't just the framing but the honesty of it. OpenAI isn't pretending its models have already cracked this. The benchmark exists because nobody currently knows how close frontier systems are to functioning like an actual paid engineer, end to end — writing code, fixing bugs, navigating vague requirements, and producing something a client would accept and pay for. That's a very different bar than passing a coding interview question.
This fits a pattern we've seen from OpenAI and rivals over the past year: benchmarks are quietly shifting away from abstract reasoning tests toward economically grounded ones. Instead of asking whether a model can solve a riddle, the industry is starting to ask whether it can do a job that someone would otherwise pay a human to do. SWE-Lancer is one of the more direct attempts at putting a price tag on that question.
My take — AI-written commentary, not fact-checked reporting
I like that OpenAI is finally tying benchmarks to actual money instead of vibes and leaderboard bragging rights — that's the only kind of test that tells you anything about job displacement. But don't mistake a benchmark for a verdict; until these models can handle a confused client at 11pm without hallucinating a fix, freelancers aren't obsolete, they're just getting a very expensive new competitor to watch.
Read more about this at: OpenAI