> For a complete documentation index, fetch https://docs.voximplant.ai/llms.txt

# Features Overview

<blockquote>
  For the complete documentation index, see <a href="/llms.txt">llms.txt</a>.
</blockquote>

# Voximplant Platform Capabilities

Voximplant Platform is a **Voice AI Orchestration Platform** and an established **Cloud Communications Platform** for building programmable voice, video, and messaging applications using serverless call control, SDKs, and APIs.

Click on any card below to see more information.

## Capabilities at a glance

#### [Voice AI orchestration](#voice-ai-orchestration)

Connect real-time AI agents, speech systems, and telephony channels with code-driven orchestration.

#### [Voice telephony](#voice-telephony)

Run inbound/outbound PSTN, SIP, WebRTC, and WhatsApp voice flows with fine-grained call control.

#### [Tools and DX](#tools-and-developer-experience)

Use cloud IDE/debugging, multi-platform SDKs, and Management API automation.

#### [Media streaming](#realtime-media-streaming)

Stream live call audio over WebSockets for real-time AI, transcription, and analysis pipelines.

#### [Reliability and deployment](#network-reliability-and-deployment)

Run globally on serverless infrastructure with multi-region coverage and uptime monitoring.

## Voice AI vendors at a glance

### Direct agent and real-time connectors

#### [OpenAI Realtime](/voice-ai-orchestration/openai/overview)

<img src="https://fdr-prod-docs-files-public.s3.us-east-1.amazonaws.com/voximplant.docs.buildwithfern.com/7800930955052cd54de5d4f4261ac4f76193011c10aa2a29b54ef5566cedd0dd/docs/assets/connectors/openai.svg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=AKIA6KXJSKKNFOCF7G4B%2F20260811%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260811T054600Z&X-Amz-Expires=604800&X-Amz-Signature=5b5e90a384b8f677a320e14fce70f480b1e0553176083ca91288e86fb2f6bf93&X-Amz-SignedHeaders=host&x-amz-checksum-mode=ENABLED&x-id=GetObject" alt="OpenAI logo" />

Realtime and agent-style voice integrations.

#### [Google Gemini Live](/voice-ai-orchestration/gemini/overview)

<img src="https://fdr-prod-docs-files-public.s3.us-east-1.amazonaws.com/voximplant.docs.buildwithfern.com/404eba6940a54e63d40edcce2d2e7cb2b3dbfec765e7a1d523662b6f4e0d6747/docs/assets/connectors/gemini.svg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=AKIA6KXJSKKNFOCF7G4B%2F20260811%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260811T054600Z&X-Amz-Expires=604800&X-Amz-Signature=6e86cc470fe4c95428213283bd2782abd3bde8a67942a582f373f68cf817109a&X-Amz-SignedHeaders=host&x-amz-checksum-mode=ENABLED&x-id=GetObject" alt="Gemini logo" />

Live speech interactions with Gemini APIs.

#### [Ultravox Realtime](/voice-ai-orchestration/ultravox/overview)

<img src="https://fdr-prod-docs-files-public.s3.us-east-1.amazonaws.com/voximplant.docs.buildwithfern.com/77c44e07e74ef847190f30db5dd0235f9f545f530454c7e9102608a1e181c848/docs/assets/connectors/ultravox.svg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=AKIA6KXJSKKNFOCF7G4B%2F20260811%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260811T054600Z&X-Amz-Expires=604800&X-Amz-Signature=7c1dd57a0e90a345b4572e2d6c4c054767b5fbed9d087dcc946e58e39c905c90&X-Amz-SignedHeaders=host&x-amz-checksum-mode=ENABLED&x-id=GetObject" alt="Ultravox logo" />

WebSocket-based speech-native connector.

#### [Deepgram Voice Agent](/voice-ai-orchestration/deepgram/overview)

<img src="https://fdr-prod-docs-files-public.s3.us-east-1.amazonaws.com/voximplant.docs.buildwithfern.com/391f7d32e87f97489a3f409fc80a38bc8ef592617070e735055bfaa2c6e6f711/docs/assets/connectors/deepgram.svg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=AKIA6KXJSKKNFOCF7G4B%2F20260811%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260811T054600Z&X-Amz-Expires=604800&X-Amz-Signature=116630c58811a55c9ac72b801111976c2a62edc2c8667ebc7c277e07a5a32ad7&X-Amz-SignedHeaders=host&x-amz-checksum-mode=ENABLED&x-id=GetObject" alt="Deepgram logo" />

Native voice-agent connector and examples.

#### [ElevenLabs ElevenAgents](/voice-ai-orchestration/elevenlabs/overview)

<img src="https://fdr-prod-docs-files-public.s3.us-east-1.amazonaws.com/voximplant.docs.buildwithfern.com/24e2fe318bc1ed54a29ceb867e4ecead952355cda4241b07402d1ff6510dc133/docs/assets/connectors/elevenlabs.svg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=AKIA6KXJSKKNFOCF7G4B%2F20260811%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260811T054600Z&X-Amz-Expires=604800&X-Amz-Signature=bc19c9038431e9912eb756c73c06955078c7df3bc0aa82813273697ce0bb661e&X-Amz-SignedHeaders=host&x-amz-checksum-mode=ENABLED&x-id=GetObject" alt="ElevenLabs logo" />

Conversational AI agent integrations.

#### [Cartesia Line Agents](/voice-ai-orchestration/cartesia/overview)

<img src="https://fdr-prod-docs-files-public.s3.us-east-1.amazonaws.com/voximplant.docs.buildwithfern.com/b1c1db2b7dd505edc9e5b049837f77d220d2ad09c536579ef644cd905ba78b21/docs/assets/connectors/cartesia.png?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=AKIA6KXJSKKNFOCF7G4B%2F20260811%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260811T054600Z&X-Amz-Expires=604800&X-Amz-Signature=72aef59291f63e9055a63f72ebed51524ead637ce8e50ce3791d8629fb28583a&X-Amz-SignedHeaders=host&x-amz-checksum-mode=ENABLED&x-id=GetObject" alt="Cartesia logo" />

Line Agents runtime with VoxEngine orchestration.

#### [xAI Grok Voice Agent](/voice-ai-orchestration/grok/overview)

<img src="https://fdr-prod-docs-files-public.s3.us-east-1.amazonaws.com/voximplant.docs.buildwithfern.com/07099f0e2eea46fe8cef6a2422dee63e1aa98c3cab07f3edf9a0b333314f758b/docs/assets/connectors/grok-x.svg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=AKIA6KXJSKKNFOCF7G4B%2F20260811%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260811T054600Z&X-Amz-Expires=604800&X-Amz-Signature=c4e6b494f93e2a6ddca997bd3768869bfc34ac77c9fe4f7d8779a78ec70a10c3&X-Amz-SignedHeaders=host&x-amz-checksum-mode=ENABLED&x-id=GetObject" alt="xAI logo" />

Grok voice-agent flow and feature support.

### Speech and realtime TTS options

#### [ElevenLabs Streaming TTS](/voice-ai-orchestration/openai/half-cascade-elevenlabs)

<img src="https://fdr-prod-docs-files-public.s3.us-east-1.amazonaws.com/voximplant.docs.buildwithfern.com/24e2fe318bc1ed54a29ceb867e4ecead952355cda4241b07402d1ff6510dc133/docs/assets/connectors/elevenlabs.svg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=AKIA6KXJSKKNFOCF7G4B%2F20260811%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260811T054600Z&X-Amz-Expires=604800&X-Amz-Signature=bc19c9038431e9912eb756c73c06955078c7df3bc0aa82813273697ce0bb661e&X-Amz-SignedHeaders=host&x-amz-checksum-mode=ENABLED&x-id=GetObject" alt="ElevenLabs logo" />

Streaming/realtime TTS option for voice AI pipelines.

#### [Cartesia Realtime TTS](/voice-ai-orchestration/openai/half-cascade-cartesia)

<img src="https://fdr-prod-docs-files-public.s3.us-east-1.amazonaws.com/voximplant.docs.buildwithfern.com/b1c1db2b7dd505edc9e5b049837f77d220d2ad09c536579ef644cd905ba78b21/docs/assets/connectors/cartesia.png?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=AKIA6KXJSKKNFOCF7G4B%2F20260811%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260811T054600Z&X-Amz-Expires=604800&X-Amz-Signature=72aef59291f63e9055a63f72ebed51524ead637ce8e50ce3791d8629fb28583a&X-Amz-SignedHeaders=host&x-amz-checksum-mode=ENABLED&x-id=GetObject" alt="Cartesia logo" />

Realtime TTS pattern for half-cascade voice pipelines.

#### [Inworld Realtime TTS](/voice-ai-orchestration/openai/half-cascade-inworld)

<img src="https://fdr-prod-docs-files-public.s3.us-east-1.amazonaws.com/voximplant.docs.buildwithfern.com/04535a1b57e279ea5bccc9861b3954b4b858305f78ecaeefa73af7516fd28e93/docs/assets/connectors/inworld.png?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=AKIA6KXJSKKNFOCF7G4B%2F20260811%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260811T054600Z&X-Amz-Expires=604800&X-Amz-Signature=1ae9b3e0fdf347190d595f788d0ad96c0d5f0edbf15a4a7ab54d428c6ee46ec3&X-Amz-SignedHeaders=host&x-amz-checksum-mode=ENABLED&x-id=GetObject" alt="Inworld logo" />

Realtime TTS option for half-cascade voice flows.

#### [Google Realtime TTS](/api-reference/voxengine/google)

<img src="https://fdr-prod-docs-files-public.s3.us-east-1.amazonaws.com/voximplant.docs.buildwithfern.com/247fb8b63bc9b2f2f34b3be2e5412041cf33316d83c809c5f278c0487e42fe76/docs/assets/connectors/google.svg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=AKIA6KXJSKKNFOCF7G4B%2F20260811%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260811T054600Z&X-Amz-Expires=604800&X-Amz-Signature=145286c604c42e192cb2f2f9c11631b8adc128d885804c3d962624adb5770559&X-Amz-SignedHeaders=host&x-amz-checksum-mode=ENABLED&x-id=GetObject" alt="Google logo" />

Google realtime TTS player for low-latency speech output.

<h2 id="detailed-capabilities">
  Detailed capabilities
</h2>

<button type="button" data-capabilities-accordion-action="expand" aria-controls="capabilities-accordion-root">
  ▾ Expand all
</button>

<button type="button" data-capabilities-accordion-action="collapse" aria-controls="capabilities-accordion-root">
  ▴ Collapse all
</button>

#### Voice AI Orchestration

Voximplant AI is a serverless runtime for **Voice AI pipelines** that connects real-time agent/LLM systems
and speech engines to **PSTN / SIP / WebRTC / mobile / WhatsApp calling**, with code-driven orchestration
and provider flexibility. See [Voximplant AI](https://voximplant.ai/) and the docs [Voice AI connectors
section](/voice-ai-orchestration/overview).

**Supported vendors (direct agent / real-time LLM connectors)**

Native/direct connectivity is positioned for:

* **OpenAI** (Realtime / agent-style integrations) - [Docs:
  OpenAI](/voice-ai-orchestration/openai/overview)
* **Google Gemini (Live)** - [Docs:
  Google](/voice-ai-orchestration/gemini/overview)
* **Deepgram Voice Agent** - [Docs:
  Deepgram](/voice-ai-orchestration/deepgram/overview)
* **ElevenLabs Agents / Conversational AI** - [Docs:
  ElevenLabs](/voice-ai-orchestration/elevenlabs/overview)
* **Ultravox (WebSocket API)** - [Docs:
  Ultravox](/voice-ai-orchestration/ultravox/overview)
* **Cartesia Line Agents** - [Docs: Cartesia Line
  Agents](/voice-ai-orchestration/cartesia/overview)
* **xAI (Grok Voice Agent)** - [Docs:
  xAI](/voice-ai-orchestration/grok/overview)

Voximplant AI also explicitly supports connecting to **another WebSocket interface** (for other real-time AI
systems) in addition to the vendors above.

**Supported vendors (speech engines: STT / TTS)**

Voximplant's platform speech layer (STT/TTS) includes built-in providers such as:

* **Speech-to-Text (STT)**: Google Speech Cloud, Microsoft Azure STT, Amazon Transcribe, Yandex Speech Cloud
* **Text-to-Speech (TTS)**: Google Speech Cloud, Amazon Polly, Yandex Speech Cloud, Microsoft Azure TTS,
  Tinkoff VoiceKit

For realtime / streaming TTS used in Voice AI scenarios, Voximplant also provides native VoxEngine modules
and guides for:

* **Google Realtime TTS** - [API refs](/api-reference/voxengine/google)
* **Cartesia Realtime TTS** - [Guide: Realtime
  TTS](/voice-ai-orchestration/openai/half-cascade-cartesia)
  and [API refs](/api-reference/voxengine/cartesia)
* **Inworld Realtime TTS** - [Guide: Realtime
  TTS](/voice-ai-orchestration/openai/half-cascade-inworld)
  and [API refs](/api-reference/voxengine/inworld)
* **ElevenLabs Streaming / realtime TTS** - [Guide: ElevenLabs
  TTS](/voice-ai-orchestration/openai/half-cascade-elevenlabs)
  and [API refs](/api-reference/voxengine/elevenlabs)

**Pipeline options (architectures you can run)**

* **Speech-to-speech**: real-time audio in and real-time audio out (agent API handles full duplex loop)
* **Speech -> LLM -> TTS**: stream audio directly into a speech LLM and use a different TTS for output
* **STT -> LLM -> TTS**: stream audio to STT, pass text to an LLM/toolchain, synthesize response audio
* **Hybrid**: combine a real-time agent API for turn-taking with separate best-of-breed STT/TTS components
  (mix and match)

**Orchestration primitives (what you control)**

* **Mix and match providers**: swap STT/TTS/LLM vendors without changing your telephony integration
* **Parallel model execution**: run multiple speech/LLM components in parallel when useful (for example,
  intent extraction + generation)
* **Failover paths**: fall back to alternate speech/LLM providers when a step errors or times out
* **Wideband audio**: higher fidelity audio path for improved user experience and model comprehension
* **Deep SIP support**: SIP trunking + registration interop so agents can operate inside PBX/SBC/carrier
  environments
* **Channel portability**: reuse the same AI pipeline across PSTN numbers, SIP, WebRTC, mobile SDKs, and
  WhatsApp calling

**Real-time media integration (streaming)**

* **WebSocket-based media streaming** for connecting calls to real-time AI systems and custom pipelines
  (audio + metadata/control messages on the same channel)
* **Media gateway abstraction**: avoid building/operating custom streaming gateways when using native
  connectors/modules

#### Voice telephony

**Connectivity and endpoints**

* **PSTN calling** (inbound/outbound) via phone numbers and programmable call handling
* **Phone numbers API**: automated procurement in **60+ countries** (availability varies by country)
* **SIP calling and trunking**: connect carriers / PBXs / SBCs using SIP interop (including
  registration-based scenarios)
* **WebRTC calling** via web/mobile SDKs (VoIP calling in apps and browsers)
* **WhatsApp calling**: inbound/outbound voice calls via WhatsApp Business API integration

**Serverless call control (VoxEngine)**

* **JavaScript call logic (no XML)** for real-time call routing and application workflows
* **Per-call-leg signaling/media control** - granular control over each leg independently

**Conferencing and bridging**

* **Single conferencing API** for voice/video; **mix PSTN, SIP, WebRTC, and native mobile endpoints**
* **Conferences up to 50 participants**

**Recording, transcription, and speech processing**

* **Call recording** via `call.record()` in scenarios (supports stereo and additional options)
* **Call transcription** via `record(transcribe=true)` and retrieval via `GetCallHistory` (transcription
  delivered asynchronously)
* **Speaker/channel labeling** in transcripts (for example, "Left"/"Right" labeling pattern described in
  docs)

**Speech-to-Text (ASR) modes and features**

* **Phrase-hint mode** (best for constrained dialogs / IVRs) and **Freeform mode** (open transcription)
* **Multiple ASR engines** (for example, Google, Amazon, Microsoft, Yandex, T-bank) with selectable profiles
* **Intermediate results** support (provider-dependent) for faster partial recognition
* **Google Speech v1p1beta1 feature passthrough** (for example, word time offsets, punctuation, diarization
  config)

**Answering machine / voicemail / beep detection**

* **AMD module** for voicemail/answering machine detection in scenarios
* **Beep detection** with specified frequency lists and timeouts (scenario-level control)
* **AMD event/callback model** available in VoxEngine references

<a id="automated-outbound-calling-call-lists--dialing-logic" />

**Automated outbound calling (call lists + dialing logic)**

* **Call Lists**: upload a **CSV call list** and process it with VoxEngine scenarios (campaign-style
  calling)
* **Management API CallLists**: programmatic call-list upload/append with delimiter support
* **Predictive Dialing System (PDS)**: uses agent/load statistics and call-list progression to place calls
  and connect answered calls to agents
* **Predictive and progressive dialing modes** with tunable parameters (for example, allowed failed call
  percentage)

#### Tools and Developer Experience

**Cloud IDE and debugging**

* **Cloud IDE + debugger** in the control panel:
* **Code verification**
* **Autocompletion**
* **Diff highlighting**
* Built-in troubleshooting workflow

**Local IDE continuous integration**

* CLI tool for CI/CD automation so you can use your own IDE
* Type library for local development with autocompletion and type checking

**SDKs and client libraries**

* SDKs: **iOS, Android, Web, React Native, Flutter, Unity**
* API clients: **curl, Node.js, Python, PHP, Go, .NET, Java**

**Management API (HTTP)**

* Control accounts/services programmatically (examples from docs include managing phone numbers, messaging,
  billing, logs, records, and user access)

#### Real-time Media Streaming (WebSockets / Media Streams)

* **Media Streams**: integrate **live audio streams** into calls via WebSockets for real-time
  transcription/analysis and AI integrations
* WebSocket programming model in VoxEngine:
* Create connections via `VoxEngine.createWebSocket(...)`
* Stream audio using `WebSocket.sendMediaTo(...)`
* Recommended audio chunk duration: **\~20ms**

#### Network, Reliability, and Deployment

* **Serverless runtime**: no infrastructure to manage for call logic
* **CI/CD-friendly deployment path**: CI tool for automated pipelines
* **Global footprint**: datacenters in **14** countries
* **Status page** for live and historical uptime of subcomponents