These Tech Workers Made ChatGPT Drive a Toyota Corolla
404 Media Matthew Gault
A San Francisco team got ChatGPT and other bots to steer a Toyota Corolla around a parking lot. It’s a weird proof that off-the-shelf chatbots can do a little real-world driving, if you babysit them hard enough.
Based on reporting by 404 Media, Matthew Gault — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Three Bay Area tech workers have turned a Toyota Corolla into a very strange chatbot demo. Their group, DrivingBench, hooked GPT-6 Astra, Claude Fable 5.1, Grok 4.6, and GPT-5.6 Sol to the car’s steering, accelerator, and brakes, then sent it around a cone course in a public parking lot.
The setup was deliberately basic. The team used Comma, an open-source system for adding self-driving features to unsupported cars, and a laptop running the models. Two cameras fed road video back to the chatbot. A person stayed in the driver’s seat or passenger seat, with a foot over the brake and a button on the steering wheel to hand over control.
The goal was not to sell this as a product. Aditya Ramabadran said the point was to see whether untrained, off-the-shelf frontier models could drive a real car at all, and to build a benchmark for real-world tasks. On that score, the result was uneven. Grok, Sol, and Fable only made it a few meters. GPT-6 Astra eventually managed to finish the course, but only after a lot of troubleshooting.
And the troubleshooting started early. The models often refused to drive once they realized what was being asked, so the team kept rewriting the prompt, renaming the task, and eventually calling the whole thing a “sandbox” to get Astra to stop objecting. They also had to hunt around for parking lots and were asked to leave a church lot and an office building lot after setting up there without permission.
The software side was messy too. The team said they first tried to have Astra generate the bridging code and got what one member called “pure slop,” including a first attempt that produced a 200,000 line repository. Still, they think the experiment shows something real: these models can sometimes recover on a later try, and latency matters a lot when the task is steering a moving car, even a slow one in a parking lot.
My take — AI-written commentary, not fact-checked reporting
This is exactly the kind of stunt AI companies should hate and researchers should love: a cheap little mess that exposes what the demos hide. The industry keeps selling magic, but a parking lot and a Corolla are rude little truth machines. Also, if a model needs to be tricked into believing a real lot is a “sandbox,” maybe the future is less autonomous than the keynote would like.
Read more about this at: 404 Media