Amazon Web Services released benchmarking results for its Amazon Bedrock models, showing how different OpenAI models perform on real-world tasks. It is the company's first detailed comparison of models since the launch of the Bedrock platform in 2024.
Amazon reported a cost per correct answer of $0.0021 for gpt-5.6-luna on AIME, measured against OpenAI's gpt-5.4-mini at $0.0139. That compares with gpt-5.4-nano at $0.0072, though its pass rate was 18 percent.
The models are built on Amazon Bedrock's infrastructure and target enterprise applications requiring high accuracy and efficiency. Availability remains through the Amazon Bedrock platform, initially for developers and enterprises.
"The core benchmarks here are reproducible: the harness, openai-on-aws/benchmarks-openai, runs one identical code path (the OpenAI Responses API) against both applications," said the AWS blog. The results reflect differences in the models, provider infrastructure, and model-specific configuration.
The announcement follows the July 30, 2026, price reduction for GPT-5.6 Luna and Terra on Amazon Bedrock. Amazon did not say whether newer models will replace existing cost-optimized baselines, and the source raises the open question of how to balance cost and quality for different workloads.
Source: awsml