OpenAI has released two new speech recognition models, GPT Transcribe and GPT Live Transcribe, through its API. GPT Transcribe is designed for pre-recorded audio files, processing them approximately 34 times faster than real time. The model's word error rate, according to the AA-WER benchmark, is 3.31 percent, representing a 0.7 percentage point improvement over its predecessor, GPT-4o Transcribe. Pricing for GPT Transcribe has also dropped by 25 percent, now at $0.0045 per minute of audio. Both models support text as transcription context, keywords, and multiple input languages. The new models are part of OpenAI's broader efforts in real-time transcription and model generation, complementing its recently announced Realtime model generation initiative.

In the AA-WER benchmark rankings, OpenAI remains behind several competitors. ElevenLabs Scribe v2 leads with a 2.3 percent error rate, followed by Google's Gemini 3 Pro at 2.9 percent and Mistral's Voxtral Small at 3 percent. Mistral recently introduced Voxtral Transcribe V2, which starts at $0.003 per minute. While OpenAI's models show progress, they still lag behind in error rates compared to these competitors. The company's latest releases highlight ongoing advancements in speech recognition technology, though challenges remain in achieving industry-leading performance.

According to Artificial Analysis, which runs the AA-WER benchmark, GPT Transcribe achieves a word error rate of 3.31 percent. That's a 0.7 percentage point improvement over its year-old predecessor GPT-4o Transcribe. Pricing drops 25 percent at the same time, landing at $0.0045 per minute of audio. Both models accept text as transcription context, keywords, and multiple input languages. In the AA-WER ranking, OpenAI still sits behind several competitors. ElevenLabs Scribe v2 leads with a 2.3 percent error rate, followed by Google's Gemini 3 Pro at 2.9 percent and Mistral's Voxtral Small at 3 percent. Mistral recently undercut the market with Voxtral Transcribe V2, starting at just $0.003 per minute.

Source: thedecoder