TLDRocket
Sign in

A Token Is Not a Message: Stop Storing AI Responses Like Event Logs

Ably Realtime

AI streams can vanish, fragment, or arrive cleanly depending on how they’re stored. The trick is storing the message, not the tokens, so late joiners aren’t blind.

Based on reporting by Ably Realtime — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

How a streamed AI response is stored decides what a late joiner sees. That late joiner might be a refreshed browser, a second device, or a human agent taking over a support chat halfway through. If the storage model is wrong, they get fragments, a blank screen, or a replay so messy the client has to rebuild everything itself.

The article’s test is simple: can a client that connects mid-response get the text so far, the current status, and the rest of the response without duplicates or gaps? Ephemeral delivery over SSE or WebSockets fails that test because the browser is the only place the assembled response exists. If the page refreshes, the response is gone. If the team tries to fix that by saving chunks beside the stream, they’ve already moved into chunk-log territory.

Chunk logs solve disconnects, but they hand clients a pile of fragments rather than a readable message. Every reader then needs reassembly logic, or the team has to run a projection service that rebuilds the response elsewhere. Either way, duplicates become a problem after failures, because some systems can redeliver chunks that were never confirmed. The upside is durability. The downside is that every client, or another backend service, gets to do the boring part.

Repeated record updates get closer. One row or message is rewritten as tokens arrive, so a late joiner can read the whole response so far in one shot. But that comes with a write for every token unless the team batches, and batching adds lag. The article points out the ugly arithmetic: 1,000 responses streaming at 100 tokens per second means 100,000 row updates per second. That is not a cute little backend, that is a database in a stress position.

Message appends are the cleanest option in the piece. Ably treats one AI response as one message that grows, so clients don’t rebuild the text and the database only gets one write when the response is finished. Late joiners read the accumulated message from history or rewind, see the status, and carry on. It is the rare case where the simplest rule is also the least annoying one: store the message, not the tokens.

My take — AI-written commentary, not fact-checked reporting

This is one of those engineering debates that looks abstract until support handoff turns into archaeology. Most “real-time” systems still act like the client is the only audience, which is a lovely assumption right up until a browser reloads. The future belongs to systems that treat responses as shared state, not disposable confetti.

Read more about this at: Ably Realtime

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.