EDINET-Bench: A Japanese Financial Benchmark Using Securities Reports
Sakana AI
Sakana AI built EDINET-Bench, a Japanese benchmark testing LLMs on spotting accounting fraud in real securities reports. It's now accepted at ICLR2026 and shows top LLMs barely beat basic logistic regression at catching fraud.
Based on reporting by Sakana AI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Sakana AI just got a paper accepted into ICLR2026, and it's not about flashy chatbots or coding agents. It's about whether large language models can catch companies cooking their books, using Japan's own regulatory filings as the test bed.
The project, called EDINET-Bench, pulls from EDINET, the Financial Services Agency's free disclosure system where roughly 41,000 securities reports from Japanese listed companies sit going back a decade. Sakana AI's team scraped ten years of these filings and, critically, ten years of corrected reports too, about 6,700 of them, to build a dataset of confirmed fraud and error cases. After running an LLM to flag which corrections actually involved accounting fraud or unintentional misstatements, they landed on roughly 600 labeled cases. A manual review found most held up, though a small percentage were corrections for unrelated reasons.
Why build this at all? Japan doesn't lack interest in fraud detection, but researchers without access to expensive proprietary financial datasets have had almost nothing to work with. English-language benchmarks like FinBen have pushed into harder, more realistic financial tasks, but Japan's regulatory and accounting quirks don't translate cleanly from English-trained systems. A model that aces a US-style benchmark isn't guaranteed to spot fraud patterns specific to Japanese filings.
The results are humbling. Even frontier LLMs, given balance sheets and cash flow statements in a zero-shot setup, topped out around 0.7 ROC-AUC, only modestly better than a coin flip and roughly in line with a plain logistic regression model. Feeding in text sections like business descriptions helped some, and in a few cases models leaned on auditor names in their reasoning, an interesting but slightly unsettling behavior that raises fairness questions about whether models are quietly judging audit firms rather than the numbers themselves.
Sakana AI is upfront that this is a narrow test. Real auditors work with far more than one filing, pulling in earnings briefings, internal documents, and outside context that this benchmark doesn't yet simulate. The team has open-sourced both the dataset on HuggingFace and the underlying tool, edinet2dataset, on GitHub, so anyone can regenerate or expand the benchmark as new filings arrive. That openness, more than any leaderboard number, might be the actual contribution here.
My take — AI-written commentary, not fact-checked reporting
I like this precisely because it's not another benchmark designed to make a model look good. A 0.7 AUC ceiling is a useful gut check against the assumption that LLMs are already ready to replace forensic accountants, and the auditor-name shortcut is a small but real warning sign about what these models actually latch onto. Open datasets built from public regulatory filings, rather than paywalled financial data, are exactly the kind of infrastructure this space needs more of, and it's telling that a Japanese lab had to build it themselves rather than wait for a Western one to bother.
Read more about this at: Sakana AI
Related stories
Research-Grade EdgeBench Analysis: AI Agent Benchmarking, Leaderboard Analytics, Scaling Laws, and Evaluation Metrics
MarkTechPost · 1 month ago ·
27
Introducing LifeSciBench
OpenAI · 3 months ago ·
11
CoffeeBench: Long-horizon Task Benchmark for LLM Agents in Multi-agent Economic Environments
Sakana AI ·
36