TLDRocket
Sign in

The Bad Guy With An AI Named Claude

Zvi (Don't Worry About the Vase) TheZvi

Anthropic disrupted attempts to misuse Claude across seven harm areas by finding malicious activity in requests made between December 2025 and August 2026. It describes Chinese groups sending lots of user queries to Claude—such as routing requests to specific Claude models and, in at least one case, returning the outputs while labeling them as Kimi results. The result is a warning that distillation is the main threat because it can transfer Claude’s capabilities while avoiding the transfer of its safeguards, and Anthropic claims most other misuse attempts on its Claude models still failed.

Why it matters

A lot of bad guys try to use Claude to do bad things. Mostly they fail. We think. Anthropic has disrupted a bunch of them, and offers an extensive report. If Anthropic is sharing the worst cases, or anything close … Continue reading →

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.