Add a Voice channel
A Voice channel lets callers speak to an agent in real time over WebRTC.
Creating the channel takes a name. The speech providers, models, and call behavior are configured on the channel page afterward.
Before you start
The environment needs at least one realtime speech model. Without one, Creatio AI Studio refuses to create the channel with "No voice models are configured for this environment; configure realtime models in the platform LLM (GenAI proxy) settings." Learn more: Connect a model provider.
Create the channel
-
Go to the Channels section in the navigation panel.
-
Click New. This opens the "Select Channel Type" dialog.
-
Select the Voice card. It is described as "Real-time voice calls via WebRTC (STT → Agent → TTS)."
-
Fill out the channel fields.
Field
Field value
Name
Name of the channel, for example, "Support Voice Line."
Description
Optional summary of what the channel is for.
Agent Binding
The agent that answers calls. Use the search box to find an agent. You can leave it unset and bind an agent later.
-
Click Create Channel.
-
Read the confirmation step and click Done. It explains that the channel is ready and that you can start a call from the agent's Preview panel or through the Voice API. A preview call can last up to 30 minutes, the same as a regular call.
As a result, Creatio AI Studio creates the channel with the "Active" status and opens its page. The channel is bound to an environment automatically, usually Sandbox. Check the environment in the next section.
Bind an agent and an environment
Calls to the channel run the agent version that is published and deployed to the channel's environment, so edits you have not published do not reach callers. Tool calls during a call, including the tools of Creatio integrations, use the same environment, so environment-scoped tool access applies as it does in chat.
- Open the channel from the Channels list.
- Go to the Overview tab.
- Select the agent that answers calls from this channel in the Routing block, if you did not bind one while creating the channel. The Environment field appears once an agent is bound.
- Select the environment in the Environment field. This is the environment the agent runs in. The field lists the environments the agent is deployed to. If it reads "Deploy the agent to an environment first," deploy the agent and return to this step.
As a result, incoming calls reach the agent version deployed to the selected environment. The Routing block saves each change as soon as you make it. Learn more: Publish, deploy, and promote an agent version.
A channel with no agent bound shows the "No agent bound — incoming calls go unanswered" warning and an inline agent picker. Bind an agent before you route real calls to the channel.
A voice channel created before environment binding was available has no environment selected. Select its Environment in the Routing block so that its calls run the published agent version.
Adjust how the agent speaks
This step is optional. A voice call follows the same base instructions as the agent's chat conversations, so the agent's topic, policy, and safety restrictions apply on the phone as well. The voice profile of the agent changes only how the agent speaks.
- Go to the Agents section and open the agent bound to the channel.
- Open the Voice tab on the Configure tab group. The tab appears once an active voice channel is bound to the agent.
- Enter the voice instructions in the Voice system prompt field of the Voice profile block, for example, "Speak warmly, confirm the caller's intent, and keep answers short." Use it for tone, brevity, style, and pacing.
- Fill out the Voice greeting field and turn on the Spoken behavior toggles you need, for example, Be brief or Plain speech (no markdown).
As a result, Creatio AI Studio saves the voice profile to the agent's draft. Publish the agent and deploy it to the channel's environment so that callers hear the change.
Configure the voice pipeline
-
Open the channel from the Channels list.
-
Go to the Configuration tab. The Models section holds the providers and models, and the Call behavior section holds the interruption and summary settings.
-
Select the Voice mode, if the section is shown. "Realtime" uses a single model that handles speech directly. "Pipeline" chains separate speech-to-text, language, and text-to-speech models. The Voice mode section is enabled per organization and appears only where it has been turned on. Without it, the channel uses the "Realtime" mode.
-
Fill out the fields for the selected mode.
Field
Field value
Speech-To-Text Provider, Speech-To-Text Model
The service and model that transcribe the caller's speech. Available in the "Pipeline" mode.
Text-To-Speech Provider, Text-To-Speech Model, Text-To-Speech Voice
The service, model, and voice that speak the agent's reply. Available in the "Pipeline" mode. Click the preview control next to a voice to hear it.
Conversation model
The LLM that generates the reply between transcription and speech. Available in the "Pipeline" mode.
Realtime provider, Realtime model, Realtime voice
The single model that handles speech end to end. Available in the "Realtime" mode.
Tools execution mode
How a realtime model calls tools: "Platform tools," "Native provider tools," "Hybrid (platform + native)," or "No tools." Available in the "Realtime" mode. In the "Platform tools" and "Hybrid (platform + native)" modes, the agent can also use the skills attached to its published version. Skills are not available in the "Native provider tools" and "No tools" modes.
Language
The language of the conversation.
VAD Silence Threshold (ms)
How long the caller must stay silent before the agent treats the turn as finished. Accepts a value from "100" to "2000" in steps of 50. A lower value makes the agent respond faster but increases the chance it interrupts a caller who is pausing mid-sentence.
Allow interruptions
Lets the caller talk over the agent and stop its reply.
Generate call summary
Produces a summary when the call ends.
-
Click Save.
As a result, Creatio AI Studio applies the configuration to new calls on this channel.
Enable video and an avatar
This step is optional. The Capabilities tab controls what a caller can share or experience during a voice session. Video sharing and the avatar are enabled per organization and appear only where they have been turned on. Elsewhere, the Video sharing row is marked "Not available" and the Avatar row is marked "Coming soon."
- Open the channel from the Channels list.
- Go to the Capabilities tab.
- Turn on the switch on the Video sharing row of the Voice capabilities section to let the caller share a visual source during the call.
- Select the Available source field to control which visual source the caller may share, then set the Capture profile field.
- Turn on the switch on the Avatar row to show an animated agent face during the call. Select the Avatar provider and Avatar quality, then click Choose avatar and pick one. Depending on the provider, additional fields appear, for example, Anam connection mode, Avatar speech source, and Anam voice.
- Click Save.
As a result, callers can use the enabled capabilities during a session.
The avatar needs a provider key for the channel's environment, an avatar selection, and an active agent deployment. If any of these is missing, the avatar state reports the reason, for example, "No provider key" or "Not deployed," and no face appears on the call. Creatio AI Studio blocks the save with "Select an available avatar before saving." until one is chosen, so a saved avatar is one that can render on a call. Set the provider keys in the Settings section. Learn more: Connect a model provider.
Monitor voice operations
The Voice operations section of the channel page tracks latency, reliability, and completion health for live sessions.
-
Open the channel from the Channels list.
-
Select the period in the Metrics period field: "Last hour," "Last 6 hours," "Last 24 hours," or "Last 7 days."
-
Review the metrics.
- Pipeline latency — percentiles for each stage of a voice turn: STT, LLM, TTS, and the total turn.
- Experience latency — how quickly callers hear the first response and audio output.
- Call outcomes — the recent mix of completed, failed, cancelled, and superseded calls.
- Operational alerts — threshold breaches returned by the voice metrics API.
As a result, you can see where a slow or failing call spends its time. The metrics stay empty until the channel handles traffic.