Three Engineers Put GPT, Claude, and Grok Behind the Wheel of a Real Car β Only One Reached the Drive-Thru
Published: October 7, 2026
In one of the more unusual autonomy experiments of 2026, three engineers strapped a set of consumer large language models β GPT, Claude, and Grok β into a real 2024 Toyota Corolla and asked them to drive it to an In-N-Out. The setup was equal parts stunt and stress test: instead of purpose-built self-driving stacks, the researchers handed raw natural-language reasoning models control of a vehicle and let them figure out the road in real time.
Why an LLM Behind the Wheel?
Large language models have spent the past few years creeping out of chat windows and into physical systems β robot arms, drones, and now cars. The appeal is obvious: an LLM can interpret messy, open-ended instructions like "get me a Double-Double" without a bespoke rule set for every scenario. The risk is equally obvious. Language models are trained to predict text, not to judge following distance at 40 mph.
To bridge that gap, the team built a translation layer that converted each model's text-based decisions into steering, throttle, and brake commands. The car retained its factory safety systems, and every model's output was filtered through hard constraints before it reached the actuators.
The Results: One Winner, Two Failures
Of the three models tested, only one completed the trip without human intervention. The other two failed in ways that reveal how differently these systems reason about the physical world:
- The successful model handled lane-keeping, stop signs, and the final drive-thru approach with only minor corrections from the safety driver.
- The second model repeatedly over-corrected, misreading parked cars as active obstacles and stalling at intersections.
- The third model produced confident but nonsensical commands β at one point attempting a maneuver that would have violated traffic law had the safety layer not overridden it.
What This Actually Tells Us
The experiment is less a proof of concept for LLM-driven cars than a demonstration of how far the technology still has to go. LLMs are excellent at interpreting intent and terrible at the millisecond-scale physics that real driving demands. That is precisely why most 2026 production autonomy programs still rely on dedicated perception and control models rather than general-purpose language models.
Still, the researchers argue the test matters. As LLMs grow more capable of grounding their reasoning in sensor data, the line between "language model" and "driving model" may blur. For now, though, getting a burger by asking an AI nicely remains a research project β not a feature you can buy.
The Bottom Line
One model drove. Two did not. And the gap between a clever chatbot and a safe driver is still measured in more than miles.
via Wired AI
