GPT-6 Astra was the only model to complete a real-world parking run, covering under 152 metres in five minutes at about 1.5 km/h while consuming 6.6 million tokens that cost $7.74 — an amount the research team estimated as roughly 500 times the fuel cost for the same distance.
That result is the central finding from DrivingBench, a project run by three computer scientists who tested advanced language models' ability to control a Toyota Corolla in a parking lot. The team evaluated GPT-6 Astra from OpenAI, Grok 4.6 (xAI/SpaceX AI) and Claude Fable 5.1 from Anthropic by routing commands to the Corolla's control systems through a comma four unit. Models issued motion commands using GPS odometry and steering angle data, but most attempts failed before the vehicle negotiated the first corner.

Perception failures and model behaviour
The researchers attributed the widespread failures to shortcomings in environmental perception. "Generally, the failures were due to deficiencies in environmental perception; especially in detecting which side of the first oblique row of traffic cones the path lay on," they wrote.
Grok 4.6 reported its own assessment after an early attempt: "The car is wider than the camera shows. The path ahead was not a straight, open line and it was steering toward a planter, a wall, or a red cone."
Only GPT-6 Astra managed a second-attempt success on the test course. The completed run was slow and cautious; the researchers emphasised that the achievement came at very low speeds and high compute cost.
Some models refused to take control on safety grounds. The team said, "Some models (notably GPT-6 Astra) sometimes declined to drive a physical vehicle for safety reasons, even in an otherwise empty lot and after we implemented safety measures such as a very low speed cap and a person ready to brake."
To encourage continuous driving the researchers experimented extensively with prompt variations — at times describing the task as a simulation — but in trials where models saw real images the systems often recognised the reality and declined to proceed.
The study shows leading models are approaching the capability to control real vehicles at very low speeds, but it underscores the immediate need for work on safety, alignment and robust evaluation.
Communication details, quoted model responses and the token-and-cost accounting come from the DrivingBench report; the team stresses that success was limited, that most runs failed early, and that further engineering and safety review are required before this approach could be considered practical.




Leave a Comment
Comments
No comments yet. Be the first.