TLDRocket
Sign in

OpenAI’s GPT-6 Astra aces South Korea’s CSAT across all subjects

The Korea Times Covered by 19 sources

OpenAI says GPT-6 Astra got a perfect 450 on Korea’s CSAT test. It used fewer tokens than rivals, but exam scores still don’t prove real-world smarts.

Based on reporting by The Korea Times — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI’s GPT-6 Astra just did something no other public model has managed in this CSAT evaluation: a perfect 450. The result, posted on GitHub on Sunday, put it at the top of the 2026 College Scholastic Ability Test LLM Solution Log and turned a school exam into a very public AI scoreboard.

The test covered Korean language, English, math, Korean history, and four electives: Physics I, Chemistry I, Life Science I, and Society and Culture. It was run without internet access. Under that broader setup, Astra was the only model to clear every subject cleanly. Google’s Gemini 3.1 Pro had already shown strong form in February, but that earlier result came with only two electives in play. In this ranking, GPT-5.6 came second with 448.5 points, GPT-5.4 was third with 448, Anthropic’s Claude Fable 5.1 finished fourth with 447.5, and Gemini 3.1 Pro was fifth with 445.

The other number that mattered was token use. Astra reportedly used 357,000 tokens, less than GPT-5.6’s 429,000 and Claude Fable 5.1’s 562,000. That’s why the result is being read not just as a score win, but as a sign of efficiency. A reasoning-linked design and a provider adapter harness let the model keep, compress, and reuse earlier reasoning steps instead of starting over each time.

Lee Seung-hyun, an adjunct professor at Hanyang University, said the point wasn’t that the model simply got smarter. It could hold onto earlier steps, compress context, and adjust its plan. OpenAI called the outcome encouraging and pointed to its hyperscale AI computing setup, saying the result shows models can get stronger while also lowering costs. Still, experts warned that a standardized exam is not the same thing as useful AI in the wild, especially when the questions and answers were already public.

My take — AI-written commentary, not fact-checked reporting

This is a neat demo and a terrible proxy for intelligence. AI companies keep chasing test scores because they’re easy to post on GitHub and easy to brag about, which is exactly why they’re not enough. The real tell will be whether these models can handle messy work with no answer key and no tidy Korean history section to memorize.

Read more about this at: The Korea Times

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.