SyncAI.news, a Varaisys broadcasting
Advancing voice intelligence with new models in the API
ON

OpenAI News

· 1 min read

AI LabsOpenAI News

Advancing voice intelligence with new models in the API

We’re introducing three audio models in the API that unlock a new class of voice apps for developers. With these models, developers can build voice experiences that feel more natural, respond more intelligently, and take action in real time:

  • GPT‑Realtime‑2, our first voice model with GPT‑5‑class reasoning that can handle harder requests and carry the conversation forward naturally.

  • GPT‑Realtime‑Translate, a new live translation model that translates speech from 70+ input languages into 13 output languages while keeping pace with the speaker.

  • GPT‑Realtime‑Whisper, a new streaming speech-to-text that transcribes speech live as the speaker talks.

Try GPT-Realtime-2

Start the session, then talk naturally with GPT-Realtime-2.

What can I ask?

After you start the session, try saying one of these:

  • I’m hosting a last-minute dinner tonight. I have 30 minutes, two vegetarian friends, one mushroom-hater, and a tiny kitchen. Help me plan a simple menu.
  • I’m welcoming guests to a live event in Japan. Say a warm, natural welcome in Japanese — like a host kicking off something special.
  • My order number is Orbit-742Q. Repeat it back clearly so I can confirm it’s right.
  • Help me practice telling my team we hit our launch milestone. First say it with quiet confidence, then with more excitement.
  • I’m planning trivia for a road trip. Give me three trick questions that sound deceivingly simple, then explain each answer in one sentence.

This demo is time-limited. By using it, you agree to OpenAI's Terms and acknowledge our Privacy Policy.

Voice is becoming one of the most natural ways for people to use software. It lets someone ask for help while driving, change a travel plan while walking through an airport, get support in their preferred language, or move through a task without stopping to type.

Together, the models we are launching move realtime audio from simple call-and-response toward voice interfaces that can actually do work: listen, reason, translate, transcribe, and take action as a conversation unfolds.

Original source

This story was published by OpenAI News. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on openai.com

Similar News