You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: fern/customization/speech-configuration.mdx
+7-5Lines changed: 7 additions & 5 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -38,15 +38,17 @@ This plan defines the parameters for when the assistant begins speaking after th
38
38
-**End-of-turn prediction** - predicting when the current speaker is likely to finish their turn.
39
39
-**Backchannel prediction** - detecting moments where a listener may provide short verbal acknowledgments like "uh-huh", "yeah", etc. to show engagement, without intending to take over the speaking turn. This is better handled by the assistant's stopSpeakingPlan.
40
40
41
-
We offer different providers that can be audio-based, text-based, or audio-text based:
41
+
Vapi supports built-in smart endpointing, endpointing supplied by compatible transcribers, and custom endpointing models:
42
42
43
-
**Audio-based providers:**
43
+
**Custom endpointing:**
44
44
45
-
-**Krisp**: Audio-based model that analyzes prosodic and acoustic features such as changes in intonation, pitch, and rhythm to detect when users finish speaking. Since it's audio-based, it always notifies when the user is done speaking, even for brief acknowledgments. Vapi offers configurable acknowledgement words and a well-configured stop speaking plan to handle this properly.
45
+
-**Custom endpointing model**: Lets your server decide how long Vapi waits before considering the customer's speech finished. Use it when you want your own service to control the endpointing timeout for each turn.
46
46
47
-
Configure Krisp with a threshold between 0 and 1 (default 0.5), where 1 means the user definitely stopped speaking and 0 means they're still speaking. Use lower values for snappier conversations and higher values for more conservative detection.
47
+
Set `startSpeakingPlan.smartEndpointingPlan.provider` to `custom-endpointing-model`. Vapi sends a `call.endpointing.request` containing the conversation history to the configured `server.url` whenever it receives a new transcript.
48
48
49
-
When interacting with an AI agent, users may genuinely want to interrupt to ask a question or shift the conversation, or they might simply be using backchannel cues like "right" or "okay" to signal they're actively listening. The core challenge lies in distinguishing meaningful interruptions from casual acknowledgments. Since the audio-based model signals end-of-turn after each word, configure the stop speaking plan with the right number of words to interrupt, interruption settings, and acknowledgement phrases to handle backchanneling properly.
49
+
Your server returns a `timeoutSeconds` value from 0 to 15. Vapi resets the timeout when it receives the next transcript and sends another request. If `server` is not configured, Vapi uses `assistant.server`, then `org.server`.
50
+
51
+
See [Custom endpointing model configuration](/customization/voice-pipeline-configuration#custom-endpointing-model-configuration) for configuration details and [Call endpointing request](/server-url/events#call-endpointing-request-custom-endpointing-server) for the request and response contract.
0 commit comments