SyncAI.news, a Varaisys broadcasting
Gemini Live audio
SW

Simon Willison's Weblog

· 1 min read

AnalysisSimon Willison's Weblog

Gemini Live audio

Tool: Gemini Live audio

Google released Gemini 3.8 Live and 3.8 Live Extended Thinking today - two new speech-to-speech models that are a similar shape to OpenAI's GPT-Live family.

I pointed GPT-6 Astra Extra High at the documentation and had it build me this web UI for trying out the new models. You can select a model and voice preset, enter an optional system prompt and then start a voice conversation through your browser, including the ability to interrupt the model while it is talking.

The implementation uses no libraries. It connects to the wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent?key=... WebSocket endpoint and uses a Web Audio API AudioContext for both capture and playback.

Here's the Gemini Live tutorial for getting started with that WebSockets API.

Original source

This story was published by Simon Willison's Weblog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on simonwillison.net

Similar News