TLDRocket
Sign in

An opinionated guide to which AI to use to do stuff

One Useful Thing Covered by 8 sources

A guy who writes AI guides says the game has shifted from chatbots to AI agents that use whole computers on their own. Basically, pick ChatGPT or Claude, pay $20, and let it actually do stuff — but watch the permissions.

Based on reporting by One Useful Thing — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Every few months this newsletter writer updates his AI guide, and this time the ground has genuinely moved. The old model was a chatbot: you type, it types back, repeat. The new model is agentic — an AI paired with tools that let it plan, click, browse, and finish multi-hour tasks while you do something else entirely. He illustrates the jump with a small but telling anecdote: a brutalist city-building game he had GPT-5 generate last year looks primitive next to what GPT-5.6 Sol built from a similar prompt inside Codex. Less than twelve months apart, wildly different results.

His advice splits into two tiers. For low-stakes chatting — recipes, quick questions, a first draft of a letter — basically every free model on the market is fine now, so just use whichever one you like. But for anything that actually matters, a second opinion on a legal letter or a health question, he says go straight to the top-tier models: Claude's Opus or Fable, or ChatGPT's GPT-5.6 Sol, cranked to the highest thinking setting. Lower error rates cost money, but that's the trade.

For real work, he narrows the field to two products, ChatGPT and Claude, both starting around $20 a month, both confusingly named. There are two flavors: a hosted virtual computer (called Work in ChatGPT, Cowork in Claude) or full access to your actual machine (Codex for ChatGPT, Code for Claude). He tested the hosted mode by asking both to raid his Gmail and prep an entire MBA seminar — slides, demos, replies to colleagues — and both delivered a few hours' worth of work in about ten minutes. The catch: ChatGPT actually sent an email on his behalf because he'd previously granted it permission, while Claude, set more cautiously, only drafted one. He's blunt that permission settings matter enormously here, and that prompt injection — outside text trying to hijack an agent mid-task — remains an unsolved risk worth guarding against by defaulting to approval-required settings.

Giving an agent your actual computer unlocks the bigger stuff. He fed the full PDF of his forthcoming book to GPT-5.6 Sol in Codex; in thirty minutes it verified 195 references with zero hallucinated citations, though it turned out overly nitpicky, requiring him to overrule some of its pickier complaints. He also let Codex's computer-use feature literally take his mouse and download Blender to sculpt a 3D otter unprompted beyond a one-line instruction. That's the ceiling right now: an AI that can operate your machine like a slightly overzealous intern.

Everything else, he argues, is a tier below. Microsoft Copilot is fine for office documents but weak as an agent. Chinese open-weight models like Kimi K3, DeepSeek, and Qwen are impressively capable but need real technical chops to deploy as agents. And Google, despite once leading benchmarks, currently has no frontier model or agentic harness worth recommending as a daily driver — though its Notebook tool remains excellent for research synthesis, and its Gemini Omni video model can do genuinely strange, delightful things, like turning 1896's famous arriving-train film into a LEGO train complete with a time traveler and the Muppets, shadows and reflections intact.

My take — AI-written commentary, not fact-checked reporting

What strikes me most isn't the benchmark chasing, it's that the real skill now is management, not prompting — you're delegating to something more like a slightly reckless intern than a search engine, and permissions settings are the actual safety net, not vibes about alignment. Google falling this far behind on agentic tooling despite once owning the benchmark conversation is the story nobody in Brussels seems to be tracking, and it should worry anyone hoping Europe ends up with a domestic alternative to Anthropic and OpenAI. Open-weight Chinese models being 'surprisingly capable but requiring expertise' is doing a lot of quiet work in that sentence, and I don't think that gap closes on its own.

Read more about this at: One Useful Thing

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.