TLDRocket
Sign in

GLM 5.2 is (nearly) as accurate as a human book-keeper at less than 1% of the cost

toot-books.pages.dev

An open-weights AI model just did a UK company's quarterly VAT return almost perfectly, for under $3 in raw compute. Human book-keepers charge hundreds to thousands of pounds for the same job — and this thing landed within 7 pence of them.

Based on reporting by toot-books.pages.dev — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Vineyard Finance ran an experiment that sounds almost like a stunt until you look at the numbers: they handed GLM 5.2, an open-weights model served via Fireworks AI, the job of preparing a real UK VAT return for Q1 2026. Fifty-nine transactions, a bash-only tool harness, no vision needed since every receipt was a text PDF. The model logged into the accounting software through a CLI, worked through three months of bank feeds one by one, and came out the other end with a return that missed the human-prepared net VAT figure (Box 5) by just seven pence. Total cost: $2.73. Total time: 68 minutes across three separate monthly sessions.

That's the headline. The detail is where it gets interesting. Out of 354 individual checks — six criteria across 59 transactions — the model failed 20, and only one of those failures actually matters. It booked a founder's 10,000 GBP capital contribution into a generic "Capital Account" instead of the legally distinct "Unpaid Shares" line, which in UK company law carries real creditor-protection implications and needs specific disclosure at year-end. Wrong VAT impact: zero. Wrong in a way a real accountant would never be: yes.

The rest of the mistakes are smaller and stranger. Fourteen times the model mixed up "zero-rated" and "tax-exempt" VAT categories — two buckets that both mean no VAT changes hands but are treated differently by tax authorities for subtler reasons. Curiously, it made this exact error every single time in January and February, then got every instance right in March, a flip that suggests something closer to inconsistent reasoning than a fixed knowledge gap. A handful of remaining errors came from Wise's habit of splitting a single card payment across currency balances, which confused the model into slightly double-counting VAT on a couple of split transactions — again immaterial to the bottom line, but the kind of thing a sharp bookkeeper would spot immediately.

What the model never got wrong is arguably the more important story. It correctly untangled duplicate same-day, same-vendor, same-amount transactions. It correctly identified a transfer dressed up as a card purchase. It attached the right receipt to the right transaction every single time. These are the exact judgment calls that separate a competent bookkeeper from a cheap, unreliable one, and until very recently only frontier proprietary models or paid humans could reliably clear that bar.

Vineyard Finance's own read is blunt: bookkeeping is becoming a solved problem, and the interesting work now is building the scaffolding — the tools, the checks, the human oversight layer — that lets SMEs actually trust and deploy something like this instead of hand-waving at a black box.

My take — AI-written commentary, not fact-checked reporting

Bookkeeping was always a task defined by tedium, not genius, so it's no shock this fell first — the surprise is that an open-weights model did it, not a closed frontier one, at a price accountants would laugh at if it weren't real. The one mistake that mattered was a legal-nuance error, which is exactly the kind of thing firms will need human sign-off on for the next few years, not because the AI is unreliable overall but because a 10,000-pound misclassification is precisely where accountability has to stay human. Open weights closing this gap on a boring, high-stakes back-office task is a much bigger deal than another chatbot benchmark, and I don't think the accounting software industry has clocked it yet.

Read more about this at: toot-books.pages.dev

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.