Why this is noise: An API model update for developers. Relevant if you are actively building voice agents, not actionable for most teams today.
Key Takeaways
- OpenAI launched three new voice models in its Realtime API: GPT-Realtime-2 (reasoning), GPT-Realtime-Translate (live translation), and GPT-Realtime-Whisper (streaming transcription)
- GPT-Realtime-2 is the first voice model with GPT-5-class reasoning, scoring 15.2% higher on Big Bench Audio than its predecessor
- GPT-Realtime-Translate supports 70+ input languages and 13 output languages, with live translation that keeps pace with the speaker
- Context window expanded from 32K to 128K tokens for longer, more coherent voice sessions and complex agentic workflows
- Pricing: GPT-Realtime-2 at $32 per million audio input tokens and $64 per million output tokens; Translate and Whisper billed by the minute
OpenAI has launched three new voice models in its Realtime API, moving its developer voice offering from simple call-and-response toward agents that can reason, translate, and transcribe as a conversation unfolds.
GPT-Realtime-2 is the first voice model in the API built on GPT-5-class reasoning. It can handle parallel tool calls, recover gracefully when requests fail, and maintain context across longer sessions with an expanded 128K token context window, up from 32K.
Developers can set reasoning effort from minimal to xhigh, trading latency for intelligence depending on the use case. Early partners include Zillow, Priceline, Glean, and Intercom.
GPT-Realtime-Translate provides live speech-to-speech translation across 70+ input languages and 13 output languages, designed for customer support, cross-border sales, education, and events where speakers need to hear responses in their own language without waiting for a separately produced version.
Deutsche Telekom and BolnaAI are among early testers, as enterprise voice AI adoption accelerates across the industry. BolnaAI reported 12.5% lower word error rates across Hindi, Tamil, and Telugu compared to any other model it tested.
Together, the models move realtime audio from simple call-and-response toward voice interfaces that can actually do work: listen, reason, translate, transcribe, and take action as a conversation unfolds. — OpenAI
GPT-Realtime-Whisper is a streaming speech-to-text model built for low-latency live transcription, designed to power meeting captions, live note-taking, and voice agents that need to understand users continuously. All three models are available now in the Realtime API.
Pricing: GPT-Realtime-2 is $32 per million audio input tokens and $64 per million audio output tokens. GPT-Realtime-Translate and GPT-Realtime-Whisper are billed by the minute at $0.034 and $0.017 respectively. The Realtime API supports EU Data Residency for EU-based applications.
OpenAI notes the models include active classifiers that can halt conversations violating its harmful content guidelines, and developers are required to disclose to end users when they are interacting with AI.
The three patterns OpenAI sees driving adoption are voice-to-action (user speaks, agent reasons and acts), systems-to-voice (software turns context into spoken guidance), and voice-to-voice (live cross-language conversation).
For teams currently building customer service, healthcare, or sales voice agents, the combination of GPT-5-class reasoning with live translation and streaming transcription in a single API surface is a meaningful upgrade over what was available six months ago.
The question is whether the pricing at scale makes voice workflows economically viable compared to text-based alternatives. Teams evaluating voice AI tools can find structured comparisons on Relve, an AI tools intelligence platform.
