Example: Transcription with manual audio commit
For the complete documentation index, see llms.txt.
gpt-realtime-whisper transcribes live call audio, but does not support OpenAI server VAD. The VoxEngine scenario below uses Silero to detect a pause and Pipecat to decide whether the caller has finished a turn. It then calls inputAudioBufferCommit() to finalize that turn’s transcript.
This example uses gpt-realtime-whisper to match the published VoxEngine scenario. For a new transcription integration, see OpenAI’s Realtime transcription guide, which currently recommends gpt-live-transcribe.
Prerequisites
- Set up an inbound call and a routing rule for this scenario.
- Store your OpenAI API key in Voximplant Secrets as
OPENAI_API_KEY.
How the turn is committed
The client uses OpenAI.RealtimeAPIClientType.TRANSCRIPTION and selects gpt-realtime-whisper in both the connection and transcription session. Its audio.input.turn_detection is null; this model requires manual audio commits.
The call sends audio to the OpenAI client, Silero VAD, and the Pipecat turn detector. When Silero reports speechEndAt, the scenario calls turnDetector.predict(). Only a Pipecat result with endOfTurn: true commits OpenAI’s input audio buffer. OpenAI then emits transcription delta and completed events, which the example logs.