Three AI engineers at Axiom used GPT-6 Astra to drive a Toyota Corolla to an In-N-Out Burger, completing the task without prior training. The engineers linked a chat interface to a server connected to cameras and the car’s power steering system, with a safety driver monitoring the brake.
Astra, normally used for text and image generation, slowly navigated the vehicle to the take-out window. One engineer remarked, "Maybe AGI is here after all," referring to artificial general intelligence.
Self-driving cars are not new, but this experiment used a language model instead of a specialized algorithm. The fast-food stunt suggests language-based AI models are gaining a rudimentary understanding of the physical world.
"Most visual AI research is about perception," said Xingang Guo, a research scientist at Scale AI. "We wanted to ask what it would actually take for a model to understand a scene intuitively, the way a person does without thinking about it."
The engineers tested models like Grok and Astra, with Astra completing a simple parking lot course, albeit slowly. The trio developed a new benchmark, DrivingBench, to measure models’ driving abilities.
"This could be like an emergent capability of just scaling up the multimodality of the model," said Mahns, referencing inputs like images, video, and 3D models. "The models really did seem to be adjusting or in-context learning based on their mistakes and learning how to better navigate the controls."
Source: wired