TLDRocket
Sign in

ttok 1.0

Simon Willison’s Weblog Simon Willison ● Covered by 2 sources

ttok 1.0 is out, and its default tokenizer now points at GPT-5/GPT-6 instead of GPT-4. That was enough for Simon Willison to call it version 1.0.

Based on reporting by Simon Willison’s Weblog, Simon Willison — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Simon Willison shipped ttok 1.0 after a small but telling mistake: he upgraded from ttok 0.4, piped a file into it, and saw it still defaulting to the GPT-4 tokenizer. That sent him back to the code, and then straight to a version bump. Sometimes software only gets its first big release when it embarrasses its maker in exactly the right way.

The new default is meant to track GPT-5 and GPT-6 instead. That sounds simple, but OpenAI still hasn’t formally confirmed that GPT-6 uses the same tokenizer as the GPT-5 family. There’s even an angry issue about it, which is a very on-brand way for tokenizer questions to live out in the open.

Willison points to a commit from William Liu that backs up the practical assumption. Liu says he ran an experiment and found that all seven GPT models he tested — 5.5, 5.6 Sol/Terra/Luna, and 6 Astra/Sol/Luna — produced 44,794 tokens and matched on all 31 fixtures. On that corpus, GPT-6 adds no input-count change at all.

So ttok 1.0 is less a flashy rewrite than a correction with a version number attached. But that’s often how useful tooling improves: one wrong default, one clean fix, and suddenly the release number starts making sense.

My take — AI-written commentary, not fact-checked reporting

This is the kind of release that deserves more respect than most AI launch notes. The industry keeps acting like tokenizer details are trivia, then spends half its life stepping on rakes because the trivia was actually the product. Boring defaults are a feature, not a failure, and the people who ship them first are usually doing the grown-up work.

Read more about this at: Simon Willison’s Weblog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.