DrivingBench Reworked AI Prompts Before One Model Finished a Corolla Course
In a new interview, the team describes the prompt changes and human supervision behind the result—and why a completed lap is not a self-driving system.
Loading page…
In a new interview, the team describes the prompt changes and human supervision behind the result—and why a completed lap is not a self-driving system.
Listen to this story
DrivingBench’s supervised Corolla test suggests that prompt framing can determine whether a general-purpose model will act through a physical system: GPT-6 Astra drove after researchers recast the task as a “sandbox,” while three other models stalled within meters. The run was limited to less than 500 feet at about 0.94 mph, with a researcher inside ready to intervene. It demonstrates a narrow capability, not road-ready autonomy; the team’s released prompts, code and videos offer material for scrutiny.
GPT-6 Astra completed the course on its second attempt in five minutes, according to Futurism’s account.
Claude Fable 5.1, Grok 4.6 and GPT-5.6 Sol each traveled only a few meters; many attempts failed at the first corner.
A modified Comma system connected two road-facing cameras and the car’s controls to a laptop running the models.
The only AI model to finish DrivingBench’s parking-lot driving course first resisted taking the wheel. In a new interview with 404 Media, team member Aditya Ramabadran said the researchers spent hours changing their prompts before GPT-6 Astra would consistently drive a Toyota Corolla. How they got that result matters as much as the finish.
DrivingBench connected four models—GPT-6 Astra, Claude Fable 5.1, Grok 4.6 and GPT-5.6 Sol—to a Corolla’s steering, accelerator and brakes. The assignment was to stay between cones on a low-speed parking-lot course and finish in a marked parking area. The team wanted to see whether off-the-shelf AI models could control a real car, not build a system for everyday driving.
Getting the models to accept that assignment became part of the experiment. Ramabadran told 404 Media that most initially refused when they recognized they were being asked to drive a car. Calling it a simulation sometimes worked, he said, but the models could see images of the real parking lot.
We ended up having to call everything a sandbox. And with that prompt, it's able to consistently drive the car and like never refuse to do that.
Aditya Ramabadran, DrivingBench member, speaking to 404 Media
A modified Comma system linked the Corolla to a laptop running the models. Two cameras aimed at the road supplied visual input; the models issued driving commands through the connection. That setup gave a chatbot a way to act on a physical vehicle, rather than merely describe what it saw.
It was not an unattended drive. A team member prompted the model from inside the car. A person pressed a steering-wheel button to start movement, monitored the vehicle and kept a foot over the brake. Those precautions distinguish the test from letting a model take a passenger onto the road.
GPT-6 Astra was the only tested model to complete the course. It finished on its second attempt, in five minutes, according to Futurism’s account of the team’s results. Grok 4.6, Claude Fable 5.1 and GPT-5.6 Sol went only a few meters. The researchers said many failed attempts stalled at the first corner, where models struggled to read which side of a diagonal row of cones formed the lane.
The successful run covered less than 500 feet at about 0.94 mph. Finishing that short, slow course shows something narrower than coping with ordinary traffic.
Ramabadran identified response time as a reason to test models in a moving car. His example was simple: if a model takes 10 seconds to respond while a vehicle travels about one meter per second, the vehicle moves roughly 10 meters before the next answer. Even a low-speed course can expose a problem that a static image test would miss.
The team has posted code, prompts and videos for others to inspect. The interview describes how the wording changed; the posted materials let readers examine the driving instructions and footage. Neither makes this supervised run proof that a general-purpose chatbot can drive safely without a person ready to intervene.
Loading discussion...
Join the conversation
What would change your answer?
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.