The Bad Guy With An AI Named Claude
Zvi (Don't Worry About the Vase) TheZvi
Anthropic disrupted attempts to misuse Claude across seven harm areas by finding malicious activity in requests made between December 2025 and August 2026. It describes Chinese groups sending lots of user queries to Claude—such as routing requests to specific Claude models and, in at least one case, returning the outputs while labeling them as Kimi results. The result is a warning that distillation is the main threat because it can transfer Claude’s capabilities while avoiding the transfer of its safeguards, and Anthropic claims most other misuse attempts on its Claude models still failed.
Why it matters
A lot of bad guys try to use Claude to do bad things. Mostly they fail. We think. Anthropic has disrupted a bunch of them, and offers an extensive report. If Anthropic is sharing the worst cases, or anything close … Continue reading →