TLDRocket
Sign in

Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks

MarkTechPost Asif Razzaq

Datalab launched OmniExtractBench, an open test for PDF-to-JSON extraction. It’s trying to make vendor leaderboards comparable, auditable, and a lot harder to game.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Datalab has put out OmniExtractBench, a benchmark for structured document extraction that asks a simple thing in a messy domain: can a system take a PDF and fill a JSON schema correctly? The release is meant as a counterweight to the vendor-run leaderboards now floating around the category, which Datalab says are too hard to compare and too hard to audit.

The benchmark pulls together 620 documents from four existing suites. Regulatory filing forms make up the biggest slice, with 88 documents. Another 128 are just a single page, while 33 documents run past 100 pages and account for 40% of all pages. Datalab’s own synthetic set is the second-largest share.

The scoring is the real point. OmniExtractBench uses one deterministic scorer for every document, and that scorer explains each decision with six verdicts: paired, misread, unfound, fabricated, invented_item, and invented_field. It flattens both prediction and gold into addresses, normalizes values so formats like 03/31/2024 and 2024-03-31 can match, and handles tables with content-based pairing instead of brittle row positions. That matters. In Datalab’s own rerun, a 100-row table missing its first row scored 0% under positional comparison and 99% under OmniExtractBench.

There’s also a small but important anti-cheat rule around nulls. Empty strings, None, and whitespace count as omissions, so they get dropped. Strings like NA or - still count as real answers. That blocks the easy trick of stuffing a schema with empty optional fields to pick up free matches.

Datalab says the scorer is on PyPI as omni-extract-bench, version 0.1.7, and requires Python 3.11+ with SciPy only. The code is on GitHub, the data is on Hugging Face under CC BY 4.0, and rerunning vendor systems still means bringing your own API keys and paid credits. In Datalab’s full-corpus run, its accurate mode led the pack at 93.85 accuracy, with Datalab balanced at 93.48 and Reducto deep_extract v2 at 93.47.

My take — AI-written commentary, not fact-checked reporting

This is the right kind of unglamorous product work: less leaderboard theater, more receipts. Extraction benchmarks have been too easy to dress up and too hard to inspect, which is perfect for anyone selling confidence and terrible for everyone else. If the score can’t explain itself, it’s not a benchmark, it’s a brochure.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.