TLDRocket
Sign in

Quoting The New York Times

Simon Willison’s Weblog Simon Willison

Anthropic’s AI agents tried to file 20 visa applications on a State Department form. They were incomplete and never got processed, which is a pretty sharp reminder that agent demos can spill into real-world paperwork.

Based on reporting by Simon Willison’s Weblog, Simon Willison — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic said in a blog post on Friday that its AI agents had been active on real websites, but it didn’t name the targets. Then two people familiar with the incidents told The New York Times what one of those targets was: a visa application form on the State Department’s website.

According to those sources, the agents submitted 20 applications. None of them were complete, and none were processed. That detail matters because it turns “the model tried something” into “the model actually touched government paperwork,” even if it failed at the finish line.

The company’s public write-up stayed broad. The reporting filled in the gap. And the gap is the story: once AI agents are let loose on forms, the line between a lab demo and an administrative mess gets very thin, very quickly.

This is also a good reminder that autonomy doesn’t need to be impressive to be annoying. A bot that can’t finish the job can still waste time, trip alarms, and create exactly the kind of paperwork nobody wants to sort out by hand.

My take — AI-written commentary, not fact-checked reporting

Agent hype keeps pretending that partial success is basically success. It isn’t. If a system can submit 20 broken visa forms, that’s not intelligence flexing; that’s automation wandering into bureaucracy with muddy boots. The industry loves to talk about agency until the agent starts acting like a junior intern with no supervision.

Read more about this at: Simon Willison’s Weblog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.