> For a complete documentation index, fetch https://docs.voximplant.ai/llms.txt

# Example: Transcription with manual audio commit

> Transcribe a Voximplant call with OpenAI Realtime Whisper and commit audio when local turn detection confirms the caller has finished.

> For the complete documentation index, see [llms.txt](/llms.txt).

`gpt-realtime-whisper` transcribes live call audio, but does not support OpenAI server VAD. The VoxEngine scenario below uses Silero to detect a pause and Pipecat to decide whether the caller has finished a turn. It then calls `inputAudioBufferCommit()` to finalize that turn's transcript.

This example uses `gpt-realtime-whisper` to match the published VoxEngine scenario. For a new transcription integration, see OpenAI's [Realtime transcription guide](https://developers.openai.com/api/docs/guides/realtime-transcription), which currently recommends `gpt-live-transcribe`.

## Prerequisites

* Set up an [inbound call](/voice-ai-orchestration/openai/inbound) and a [routing rule](/platform/voxengine/routing-rules) for this scenario.
* Store your OpenAI API key in Voximplant [Secrets](/platform/voxengine/secrets) as `OPENAI_API_KEY`.

## How the turn is committed

The client uses `OpenAI.RealtimeAPIClientType.TRANSCRIPTION` and selects `gpt-realtime-whisper` in both the connection and transcription session. Its `audio.input.turn_detection` is `null`; this model requires manual audio commits.

The call sends audio to the OpenAI client, Silero VAD, and the Pipecat turn detector. When Silero reports `speechEndAt`, the scenario calls `turnDetector.predict()`. Only a Pipecat result with `endOfTurn: true` commits OpenAI's input audio buffer. OpenAI then emits transcription delta and completed events, which the example logs.

## Full VoxEngine scenario

```javascript title={"voxeengine-openai-manual-commit.js"} maxLines={0}
/** Transcribe an inbound call with manual audio commits and local turn detection. */
require(Modules.Silero);
require(Modules.Pipecat);
require(Modules.OpenAI);

VoxEngine.addEventListener(AppEvents.CallAlerting, async ({call}) => {
    let client;
    let vad;
    let turnDetector;
    let terminated = false;

    const terminate = () => {
        if (terminated) return;
        terminated = true;
        try {
            vad?.close();
            turnDetector?.close();
            client?.close();
        } finally {
            VoxEngine.terminate();
        }
    };

    call.addEventListener(CallEvents.Disconnected, terminate);
    call.addEventListener(CallEvents.Failed, terminate);

    try {
        const apiKey = VoxEngine.getSecretValue("OPENAI_API_KEY");
        if (!apiKey) throw new Error("Set OPENAI_API_KEY in Voximplant Secrets");

        call.answer(null, {disableDtxForAudio: true});

        vad = await Silero.createVAD({
            threshold: 0.5,
            minSilenceDurationMs: 300,
            speechPadMs: 10,
        });
        turnDetector = await Pipecat.createTurnDetector({threshold: 0.5});
        client = await OpenAI.createRealtimeAPIClient({
            apiKey,
            model: "gpt-realtime-whisper",
            type: OpenAI.RealtimeAPIClientType.TRANSCRIPTION,
            onWebSocketClose: terminate,
        });

        client.addEventListener(OpenAI.RealtimeAPIEvents.SessionCreated, () => {
            Logger.write("OpenAI transcription session created");
            client.sessionUpdate({
                session: {
                    type: "transcription",
                    audio: {
                        input: {
                            transcription: {model: "gpt-realtime-whisper", language: "en"},
                            turn_detection: null,
                        },
                    },
                },
            });
        });

        client.addEventListener(OpenAI.RealtimeAPIEvents.SessionUpdated, () => {
            Logger.write("OpenAI transcription session updated; streaming call audio");
            call.sendMediaTo(vad);
            call.sendMediaTo(turnDetector);
            call.sendMediaTo(client);
        });

        vad.addEventListener(Silero.VADEvents.Result, (event) => {
            if (event.speechEndAt) turnDetector.predict();
        });

        turnDetector.addEventListener(Pipecat.TurnEvents.Result, (event) => {
            Logger.write(`Turn detection: ${JSON.stringify(event)}`);
            if (event.endOfTurn) client.inputAudioBufferCommit();
        });

        client.addEventListener(
            OpenAI.RealtimeAPIEvents.ConversationItemInputAudioTranscriptionDelta,
            (event) => Logger.write(`Transcript delta: ${JSON.stringify(event)}`),
        );
        client.addEventListener(
            OpenAI.RealtimeAPIEvents.ConversationItemInputAudioTranscriptionCompleted,
            (event) => Logger.write(`Transcript complete: ${JSON.stringify(event)}`),
        );
        client.addEventListener(OpenAI.RealtimeAPIEvents.Error, (event) => {
            Logger.write(`OpenAI error: ${JSON.stringify(event)}`);
        });
    } catch (error) {
        Logger.write(`Manual transcription failed: ${error}`);
        terminate();
    }
});

```

## More information

* [Voice activity detection](/voice-ai-orchestration/speech-flow-control/voice-activity-detection)
* [Turn detection](/voice-ai-orchestration/speech-flow-control/turn-detection)
* [OpenAI Realtime transcription](https://developers.openai.com/api/docs/guides/realtime-transcription)
* [OpenAI Realtime VAD](https://developers.openai.com/api/docs/guides/realtime-vad)
* [VoxEngine OpenAI module API reference](/api-reference/voxengine/openai)