Example: GPT Live

View as Markdown

For the complete documentation index, see llms.txt.

GPT Live (gpt-live-1) is OpenAI’s full-duplex voice model. It listens and speaks at the same time, so a caller can interrupt, add a detail, or acknowledge the assistant while it is still talking. GPT Live is the voice layer: it handles the call, the tone, and interruptions. Lookups, tool calls, and longer reasoning go to a separate backend model, so the audio keeps moving while that work runs. The result comes back into the same call.

OpenAI describes GPT Live as a model for live conversation: it handles interruptions, detects turns on its own, and picks up alphanumeric detail such as order numbers and confirmation codes. The voice session and the backend model are billed separately.

On Voximplant, open GPT Live with OpenAI.createLiveAPIClient. It returns an OpenAI.LiveAPIClient.

Jump to the Full VoxEngine scenario.

How a session works

  • session.model is the voice model, gpt-live-1. The backend model is session.delegation.responses.model.
  • Put conversation style in session.instructions. Put business rules and tool schemas on the backend, under session.delegation.responses.
  • Call sessionStart() as the first session command. Wait for OpenAI.LiveAPIEvents.SessionStarted before you bridge audio or send another command.
  • Stream call audio continuously. GPT Live decides when to speak.
  • responseCreate() continues delegated backend work.
  • Function calls arrive on OpenAI.LiveAPIEvents.ResponseEvent. Unwrap data.payload.event (for example an inner response.output_item.done) and keep the outer delegation_id.
  • Return a tool result with responseItemCreate() (function_call_output), then call responseCreate() so the backend can continue.
  • When you append instructions, thinking, or commentary, always include delegation_id. Use null for context that applies to the whole session. Each appended string is at most 500 tokens.
  • Captions arrive as SessionInputTranscriptDelta and SessionOutputTranscriptDelta. Append each fragment in order.
  • Call sessionClose() when the conversation ends, then wait for SessionClosed before you close the client.

Prerequisites

Create the client

Create a GPT Live client
const voiceAIClient = await OpenAI.createLiveAPIClient({
apiKey: VoxEngine.getSecretValue("OPENAI_API_KEY"),
});

Start the session

session.model selects the voice model. delegation.responses selects the backend model, its instructions, and its tools. This example uses Responses delegation: OpenAI runs the backend and passes conversation context between it and GPT Live. Your scenario still runs the function when the backend calls it.

Start a GPT Live session
voiceAIClient.sessionStart({
session: {
model: "gpt-live-1",
instructions: SYSTEM_PROMPT,
audio: {
output: {
voice: "marin",
},
},
delegation: {
type: "responses",
responses: {
model: "gpt-5.6-luna",
instructions: BACKEND_PROMPT,
tools: [
{
type: "function",
name: "get_order_status",
description: "Look up the shipping status of an order by its order id.",
parameters: {
type: "object",
properties: {
order_id: {
type: "string",
description: "Order id, for example A0042.",
},
},
required: ["order_id"],
},
},
],
tool_choice: "auto",
},
},
},
});

Bridge call audio

Bridge audio only after the session has started. From the same handler, append a short greeting instruction. delegation_id: null applies the instruction to the whole session.

Bridge audio and greet
voiceAIClient.addEventListener(OpenAI.LiveAPIEvents.SessionStarted, () => {
VoxEngine.sendMediaBetween(call, voiceAIClient);
voiceAIClient.sessionInstructionsAppend({
content: "Greet the caller immediately in English. After the greeting, pause and listen.",
delegation_id: null,
});
});

When that instruction is in place, tell the model to begin:

Ask the model to begin
voiceAIClient.addEventListener(OpenAI.LiveAPIEvents.SessionInstructionsAppended, () => {
voiceAIClient.sessionCommentaryAppend({
content: "Begin the conversation now, following the instructions provided.",
delegation_id: null,
});
});

Return a tool result

Return a tool result
const payload = event.data?.payload || {};
const nestedEvent = payload.event;
if (nestedEvent?.type === "response.output_item.done") {
const item = nestedEvent.item;
voiceAIClient.responseItemCreate({
item: {
type: "function_call_output",
call_id: item.call_id,
output: JSON.stringify({ status: "ok", detail: "Order A0042 shipped today." }),
},
});
voiceAIClient.responseCreate({});
}

Keep payload.delegation_id when you append instructions, thinking, or commentary for that specific delegation.

Detect assistant speech

AI Connector emits OpenAI.LiveAPIEvents.AgentStartedSpeaking and OpenAI.LiveAPIEvents.AgentStoppedSpeaking from the output PCM. Speech starts after 100 ms of continuous sound above a peak threshold of 20, and stops after 350 ms of silence, or after 3 seconds without a new audio chunk. The events mark when audio enters the media path, before the caller hears it. Mute or buffer playback in your scenario when you need to hold the audio.

If you use VoiceDSP, call signalAssistantTurnStarted() and signalAssistantTurnEnded() from these events. VoiceDSP does not receive them on its own.

Assistant speech edges
voiceAIClient.addEventListener(OpenAI.LiveAPIEvents.AgentStartedSpeaking, () => {
Logger.write("===AGENT_STARTED_SPEAKING===");
});
voiceAIClient.addEventListener(OpenAI.LiveAPIEvents.AgentStoppedSpeaking, () => {
Logger.write("===AGENT_STOPPED_SPEAKING===");
});

Close the session

Close the session
voiceAIClient.sessionClose();
voiceAIClient.addEventListener(OpenAI.LiveAPIEvents.SessionClosed, () => {
voiceAIClient.close();
});

Full VoxEngine scenario

voxeengine-openai-gpt-live.js
/**
* Voximplant + OpenAI GPT Live demo
* Scenario: answer an incoming call, bridge it to gpt-live-1,
* and delegate order and store-hours tools to a backend Responses model.
*/
require(Modules.OpenAI);
const SYSTEM_PROMPT = `
You are Voxy, a helpful voice assistant for phone callers.
Keep spoken replies brief, usually one or two sentences.
Delegate order lookups and store hours to the backend.
Do not invent an order status or opening hours.
If the caller asks about an order and did not give an order id, ask for the order id.
`;
const BACKEND_PROMPT = `
You help a live voice assistant. Transcripts can contain mistakes.
Use tools only for verified facts.
Call get_order_status when the caller asks about an order and an order id is known.
Call get_store_hours when the caller asks when the store is open.
Do not guess. Return the tool result as the answer.
`;
const ORDERS = {
A0042: {
status: "confirmed",
detail: "Order A0042 shipped today and should arrive tomorrow.",
},
B1001: {
status: "processing",
detail: "Order B1001 is still being prepared and has not shipped.",
},
};
const SESSION_CONFIG = {
session: {
model: "gpt-live-1",
instructions: SYSTEM_PROMPT,
audio: {
output: {
voice: "marin",
},
},
delegation: {
type: "responses",
responses: {
model: "gpt-5.6-luna",
instructions: BACKEND_PROMPT,
tools: [
{
type: "function",
name: "get_order_status",
description: "Look up the shipping status of an order by its order id.",
parameters: {
type: "object",
properties: {
order_id: {
type: "string",
description: "Order id, for example A0042.",
},
},
required: ["order_id"],
},
},
{
type: "function",
name: "get_store_hours",
description: "Return the store opening hours.",
parameters: {
type: "object",
properties: {},
required: [],
},
},
],
tool_choice: "auto",
parallel_tool_calls: false,
},
},
},
};
VoxEngine.addEventListener(AppEvents.CallAlerting, async ({call}) => {
let voiceAIClient;
const answeredCallIds = {};
const eventPayload = (event) => event?.data?.payload || event?.data || {};
const terminate = () => {
voiceAIClient?.close();
VoxEngine.terminate();
};
const executeTool = (name, args) => {
if (name === "get_order_status") {
const orderId = String(args?.order_id || "").trim().toUpperCase();
const order = ORDERS[orderId];
if (!order) {
return {
status: "not_found",
order_id: orderId,
detail: "No order exists with that id.",
};
}
return {
status: order.status,
order_id: orderId,
detail: order.detail,
};
}
if (name === "get_store_hours") {
return {
status: "ok",
detail: "The store is open Monday to Friday, from 9 AM to 6 PM.",
};
}
return {
status: "unsupported",
detail: "That tool is not available.",
};
};
const handleFunctionCall = (item) => {
if (
!item
|| item.type !== "function_call"
|| !item.call_id
|| (item.status && item.status !== "completed")
|| answeredCallIds[item.call_id]
) {
return;
}
answeredCallIds[item.call_id] = true;
let args = {};
try {
args = item.arguments ? JSON.parse(item.arguments) : {};
} catch (error) {
Logger.write("===TOOL_ARGUMENTS_PARSE_FAILED===");
Logger.write(error);
}
const result = executeTool(item.name, args);
Logger.write(`===TOOL_CALL=== ${item.name}`);
voiceAIClient.responseItemCreate({
item: {
type: "function_call_output",
call_id: item.call_id,
output: JSON.stringify(result),
},
});
voiceAIClient.responseCreate({}); // continue the delegated backend turn
};
call.addEventListener(CallEvents.Disconnected, terminate);
call.addEventListener(CallEvents.Failed, terminate);
try {
call.answer();
voiceAIClient = await OpenAI.createLiveAPIClient({
apiKey: VoxEngine.getSecretValue("OPENAI_API_KEY"),
onWebSocketClose: terminate,
});
voiceAIClient.addEventListener(OpenAI.LiveAPIEvents.SessionStarted, () => {
Logger.write("===SESSION_STARTED===");
VoxEngine.sendMediaBetween(call, voiceAIClient);
voiceAIClient.sessionInstructionsAppend({
content: "Greet the caller immediately in English. "
+ "Say they can ask for the status of order A0042 or B1001, or for store hours. "
+ "Then pause and listen.",
delegation_id: null,
});
});
voiceAIClient.addEventListener(OpenAI.LiveAPIEvents.SessionInstructionsAppended, () => {
voiceAIClient.sessionCommentaryAppend({
content: "Begin the conversation now, following the instructions provided.",
delegation_id: null,
});
});
voiceAIClient.addEventListener(OpenAI.LiveAPIEvents.ResponseEvent, (event) => {
const payload = eventPayload(event);
const nestedEvent = payload.event;
if (nestedEvent?.type !== "response.output_item.done") return;
Logger.write(`===DELEGATION=== ${payload.delegation_id || "none"}`);
handleFunctionCall(nestedEvent.item);
});
voiceAIClient.addEventListener(OpenAI.LiveAPIEvents.AgentStartedSpeaking, () => {
Logger.write("===AGENT_STARTED_SPEAKING===");
});
voiceAIClient.addEventListener(OpenAI.LiveAPIEvents.AgentStoppedSpeaking, () => {
Logger.write("===AGENT_STOPPED_SPEAKING===");
});
voiceAIClient.addEventListener(OpenAI.LiveAPIEvents.SessionInputTranscriptDelta, (event) => {
const text = eventPayload(event).delta || eventPayload(event).transcript;
if (text) Logger.write(`===USER=== ${text}`);
});
voiceAIClient.addEventListener(OpenAI.LiveAPIEvents.SessionOutputTranscriptDelta, (event) => {
const text = eventPayload(event).delta || eventPayload(event).transcript;
if (text) Logger.write(`===AGENT=== ${text}`);
});
[
OpenAI.LiveAPIEvents.SessionDelegationCreated,
OpenAI.LiveAPIEvents.SessionClosed,
OpenAI.LiveAPIEvents.Error,
].forEach((eventName) => {
voiceAIClient.addEventListener(eventName, (event) => {
Logger.write(`===${event.name}===`);
if (event?.data) Logger.write(JSON.stringify(event.data));
});
});
voiceAIClient.sessionStart(SESSION_CONFIG);
} catch (error) {
Logger.write("===UNHANDLED_ERROR===");
Logger.write(error);
terminate();
}
});

More information