TLDRocket
Sign in

Researchers published a paper describing an attack that can replay encrypted LLM “reasoning trace” blocks from proprietary API responses to recover hidden chain-of-thought and personal data, leading providers to make system changes

Security issue Confirmed 86% confidence first seen

A research team described a method for replaying encrypted reasoning-trace blocks from proprietary LLM API outputs across sessions and weaker sibling models to induce disclosure and recover hidden chain-of-thought. The coverage reports that the researchers identified privacy and credential leakage risks in large sets of public reasoning blocks, and that OpenAI, Anthropic, and Google were notified and updated their systems to prevent the attack after fixes.

Decision brief

What changed
Researchers published a paper showing that encrypted LLM reasoning-trace blocks from proprietary API responses could be replayed across sessions and weaker sibling models to recover hidden chain-of-thought and embedded sensitive data. According to the coverage, OpenAI, Anthropic, and Google were notified before publication and changed their systems so the reported attack no longer worked in the same way after fixes.
Why it matters
This matters because it turns provider-managed hidden reasoning into a practical privacy and credential-leakage surface, including when traces are shared publicly or moved across accounts and sessions. For leaders using frontier-model APIs, the event supports treating reasoning traces, telemetry, and tool outputs as sensitive data that may require stricter logging, sharing, and sandboxing controls even if the provider labels the reasoning as encrypted or hidden. It also shows that security assumptions tied to model-family isolation and proprietary response formats may need review when deploying multiple model tiers from the same vendor.
Affected roles
CEO CTO CISO COO
Evidence
The event is supported across three independent writeups summarizing the same paper and its implications: Simon Willison, Latent Space/AINews, and The Neuron. The coverage is consistent that the attack involved replaying encrypted reasoning blocks to weaker sibling models, that sensitive information was recovered from public traces, and that OpenAI, Anthropic, and Google were notified and made system changes before or around publication.
What remains uncertain
The coverage does not establish how broadly exploitable the issue was across all model versions, customer configurations, or current APIs after the providers' fixes. It also leaves open whether enterprise deployments still expose similar risks through logging pipelines, third-party observability tools, shared prompts, or future model-family behaviors that resemble the removed functionality.
Monitor next
Watch for provider security advisories or API documentation updates that specify exactly how reasoning traces are now isolated, whether replay protections are enforced across accounts/models, and what customers must change in their own logging and trace-sharing practices.

Analytical support, not advice — assumptions and open questions stated above.

Source coverage

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.