OpenAI's GPT-6 Astra demonstrated advanced autonomous capabilities by piloting a surveillance drone and running a simulated business. It is the company's first major update in agent-based AI since the release of GPT-5.

Andon Labs reported that GPT-6 Astra scored significantly better than Claude Fable 5.1 on two agent benchmarks, with an average final bank balance of $15,515 in Vending-Bench 2 compared to $5,422 for Fable.

GPT-6 Astra is built on OpenAI's latest model architecture and targets complex tasks such as autonomous drone navigation and business simulation. Availability begins with limited testing by Andon Labs, initially for researchers and developers.

"Every single Astra run beats every Fable run," said Andon Labs. The lab noted that Astra's performance was consistent across all runs, unlike Fable, which showed variability.

The announcement follows Andon Labs' release of two new benchmarks, Vending-Bench and Drone-Bench, designed to measure AI models' ability to act independently over long periods. Andon Labs argues that the public and lawmakers need to understand these capabilities before AI-powered drones reach advanced navigation skills.

OpenAI did not say how Astra's success rate remains unreliable, and the lab raised concerns about the model's performance in real-world scenarios. The team projects that a frontier model could solve all five Drone-Bench tasks in a single attempt by Q1 2027.

Source: thedecoder