Example: GPT Live
Example: GPT Live
For the complete documentation index, see llms.txt.
GPT Live (gpt-live-1) is OpenAI’s full-duplex voice model. It listens and speaks at the same time, so a caller can interrupt, add a detail, or acknowledge the assistant while it is still talking. GPT Live is the voice layer: it handles the call, the tone, and interruptions. Lookups, tool calls, and longer reasoning go to a separate backend model, so the audio keeps moving while that work runs. The result comes back into the same call.
OpenAI describes GPT Live as a model for live conversation: it handles interruptions, detects turns on its own, and picks up alphanumeric detail such as order numbers and confirmation codes. The voice session and the backend model are billed separately.
On Voximplant, open GPT Live with OpenAI.createLiveAPIClient. It returns an OpenAI.LiveAPIClient.
Jump to the Full VoxEngine scenario.
How a session works
session.modelis the voice model,gpt-live-1. The backend model issession.delegation.responses.model.- Put conversation style in
session.instructions. Put business rules and tool schemas on the backend, undersession.delegation.responses. - Call
sessionStart()as the first session command. Wait forOpenAI.LiveAPIEvents.SessionStartedbefore you bridge audio or send another command. - Stream call audio continuously. GPT Live decides when to speak.
responseCreate()continues delegated backend work.- Function calls arrive on
OpenAI.LiveAPIEvents.ResponseEvent. Unwrapdata.payload.event(for example an innerresponse.output_item.done) and keep the outerdelegation_id. - Return a tool result with
responseItemCreate()(function_call_output), then callresponseCreate()so the backend can continue. - When you append instructions, thinking, or commentary, always include
delegation_id. Usenullfor context that applies to the whole session. Each appended string is at most 500 tokens. - Captions arrive as
SessionInputTranscriptDeltaandSessionOutputTranscriptDelta. Append each fragment in order. - Call
sessionClose()when the conversation ends, then wait forSessionClosedbefore you close the client.
Prerequisites
- Set up an inbound entrypoint for the caller:
- Phone number
- SIP user / SIP registration
- App user - also see the Web & Mobile client SDK options
- Create a routing rule that points the destination (phone number / WhatsApp / SIP username / app user alias) to this scenario.
- Store your OpenAI API key in Voximplant Secrets under
OPENAI_API_KEY.
Create the client
Start the session
session.model selects the voice model. delegation.responses selects the backend model, its instructions, and its tools. This example uses Responses delegation: OpenAI runs the backend and passes conversation context between it and GPT Live. Your scenario still runs the function when the backend calls it.
Bridge call audio
Bridge audio only after the session has started. From the same handler, append a short greeting instruction. delegation_id: null applies the instruction to the whole session.
When that instruction is in place, tell the model to begin:
Return a tool result
Keep payload.delegation_id when you append instructions, thinking, or commentary for that specific delegation.
Detect assistant speech
AI Connector emits OpenAI.LiveAPIEvents.AgentStartedSpeaking and OpenAI.LiveAPIEvents.AgentStoppedSpeaking from the output PCM. Speech starts after 100 ms of continuous sound above a peak threshold of 20, and stops after 350 ms of silence, or after 3 seconds without a new audio chunk. The events mark when audio enters the media path, before the caller hears it. Mute or buffer playback in your scenario when you need to hold the audio.
If you use VoiceDSP, call signalAssistantTurnStarted() and signalAssistantTurnEnded() from these events. VoiceDSP does not receive them on its own.