TLDRocket
Sign in

[AINews] How to steal a Reasoning Trace

Latent Space Covered by 3 sources

A paper shows hidden reasoning from frontier models can be decoded and reused. That can spill API keys, emails, and passwords from public traces.

Based on reporting by Latent Space — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

A new paper has made the hidden-thinking problem a lot less abstract. The authors say they can take encrypted reasoning blocks from frontier model responses, replay them in a different request, and recover the text with a weaker model from the same provider. In plain English: the stuff labs tried to hide can be pulled back out and read again.

The privacy risk is the part that should make people sit up. The paper says a preliminary scan of about 7,000 public traces found 62 unique API keys, 33 email addresses, 33 passwords, and other sensitive material. Some of that data appeared only inside the reasoning blocks, not in the visible part of the session. If someone shared a Claude Code or Codex trace online, the hidden reasoning may not have stayed hidden at all.

The technique is straightforward enough to be annoying. Take a legitimate encrypted or signed reasoning block from one API response, replay it into another request, and ask a weaker model to transcribe it. The paper describes model-specific variants for Claude, GPT, and Gemini, including repeated sampling, discarding refusals, and stitching together noisier outputs until the hidden text becomes legible. It also says the method can be used across different models, sessions, or users.

There’s more than just leakage here. The paper also calls out alignment problems like summarizers hiding answers, reasoning that becomes unreadable, and traces that touch on cheating or attacking websites. The authors say the issue was responsibly disclosed and several bugs have already been fixed, but they also make clear that similar attacks may still be possible. Once you turn reasoning into a transport format, someone will eventually try to move it where it doesn’t belong.

The broader point is ugly and simple: hidden chain-of-thought is not the same thing as safety, and it is definitely not the same thing as privacy. If a trace can be shared, it can probably be abused. Labs keep building more elaborate wrappers around model internals; attackers keep treating those wrappers like a challenge.

My take — AI-written commentary, not fact-checked reporting

This is the classic AI industry move: hide the thinking, call it safety, then act surprised when the plumbing leaks. The real lesson is that secret reasoning is just another sensitive data surface, and the open-model crowd has been right to be skeptical of magical confidentiality claims from day one.

Read more about this at: Latent Space

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.