Pipecat integration
pipecat-shunyalabsai v1.1.0Native Shunyalabs STT and TTS services for Pipecat pipelines. Drop ShunyalabsSTTService and ShunyalabsTTSService into any Pipecat pipeline and get real-time streaming ASR with 46 speakers across 23 languages - no glue code required.
pipecat-ai[daily]) for WebRTC support.
Installation
Install the package from PyPI:
pip install pipecat-shunyalabsai
To include a transport (e.g. Daily WebRTC):
pip install pipecat-shunyalabsai pipecat-ai[daily]
Authentication
You need a Shunya Labs API key. Create one in the console - sign in at console.shunyalabs.ai, open API Keys, and click Create New Key (it is shown only once). Provide it as an environment variable (recommended) or pass it directly to the service classes:
export SHUNYALABS_API_KEY="your-api-key"
stt = ShunyalabsSTTService(api_key="your-api-key")
tts = ShunyalabsTTSService(api_key="your-api-key")
app.shunyalabs.ai/api/auth/token - attaches it to every request, and refreshes it in the background before it expires. You only ever handle the API key; the token is generated, used, and rotated for you, and your raw key is never sent to the STT or TTS service.
Quick start
A minimal pipeline wiring Shunyalabs STT → OpenAI LLM → Shunyalabs TTS on a local audio transport:
import asyncio, os
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.runner import PipelineRunner
from pipecat.pipeline.task import PipelineParams, PipelineTask
from pipecat.services.openai import OpenAILLMService
from pipecat.transports.local.audio import LocalAudioTransport
from pipecat_shunyalabs import ShunyalabsSTTService, ShunyalabsTTSService
async def main():
transport = LocalAudioTransport()
stt = ShunyalabsSTTService(
api_key=os.environ["SHUNYALABS_API_KEY"],
language="en",
)
llm = OpenAILLMService(
api_key=os.environ["OPENAI_API_KEY"],
model="gpt-4o",
)
tts = ShunyalabsTTSService(
api_key=os.environ["SHUNYALABS_API_KEY"],
voice="Rajesh",
language="en",
style="<Conversational>",
)
pipeline = Pipeline([transport.input(), stt, llm, tts, transport.output()])
task = PipelineTask(pipeline, PipelineParams(allow_interruptions=True))
await PipelineRunner().run(task)
if __name__ == "__main__":
asyncio.run(main())
STT - ShunyalabsSTTService
Real-time streaming speech-to-text over a persistent WebSocket connection. Supports 23 Indian and international languages with optional automatic language detection.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | str | None | API key. Falls back to SHUNYALABS_API_KEY env var. |
language | str | "auto" | Language code (e.g. "en", "hi") or "auto" for detection. |
url | str | wss://asrv2prod.shunyalabs.ai/v1/realtime | WebSocket endpoint URL. |
sample_rate | int | 16000 | Expected audio sample rate in Hz. Must match transport input. 8000 is accepted natively, so telephony audio needs no client-side resampling. |
endpoint_silence_ms | int | None | Silence before a transcript is finalised. The dominant control over how long a speaker waits after they stop talking. Unset uses the server default of 700 ms; clamped to 200-5000. See Latency tuning. |
decode_every_ms | int | None | How often interim results are produced. Unset uses the server default of 640 ms; clamped to 320-5000. Raising it lowers the cost of serving each stream at the price of coarser interim text. Does not affect how quickly a final arrives. |
vad | str | None | "silero" selects model-based endpointing. Worth setting on telephony audio, where energy thresholding can fail to register silence at all and natural endpointing then never fires. |
codeswitch | bool | None | Request code-switch refinement: a finalised segment is re-rendered in the correct scripts and delivered as a second transcript. Subject to gateway support - check the echo (below) to confirm it was enabled. |
model | str | None | Explicit model or tier. Unset lets the gateway route on language. |
min_send_bytes | int | 3200 | Audio buffered before each send. The default is 100 ms at 16 kHz. Every buffered byte is latency in front of the gateway, so lower it proportionally for 8 kHz. |
emit_turn_frames | bool | False | Emit turn-start and turn-stop frames from the gateway's own endpointing. Off by default - see Turn-taking, which explains why it must be enabled together with Pipecat's external turn strategies. |
Frame mapping
| Shunyalabs event | Pipecat frame |
|---|---|
PARTIAL | InterimTranscriptionFrame - emitted continuously as speech is recognised |
FINAL | TranscriptionFrame - a finalised segment. On a live connection this arrives once per utterance, not once per call. |
FINAL_REFINED | TranscriptionFrame - the same segment re-rendered in correct scripts, when codeswitch is enabled. Arrives after the final, as its own transcript. |
UTTERANCE_END | UserStoppedSpeakingFrame - only when emit_turn_frames=True. See Turn-taking. |
from pipecat_shunyalabs import ShunyalabsSTTService
stt = ShunyalabsSTTService(
language="hi", # Hindi; use "auto" for detection
sample_rate=16000,
)
TTS - ShunyalabsTTSService
Streaming text-to-speech over WebSocket. The plugin maintains a persistent WebSocket session that is reused across synthesis requests, streaming audio chunks back as TTSAudioRawFrame frames. Supports 46 speakers across 23 languages - any speaker can synthesise in any language.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | str | None | API key. Falls back to SHUNYALABS_API_KEY env var. |
url | str | wss://ttsv2.shunyalabs.ai/v1/realtime | WebSocket endpoint URL. |
model | str | "zero-indic" | TTS model identifier. |
voice | str | "Rajesh" | Speaker voice name. |
style | str | None | Emotion / delivery style tag. |
language | str | "en" | Output language code. |
output_format | str | "pcm" | Kept for API compatibility. The real-time stream always delivers PCM (24 kHz, 16-bit mono); other container formats are a batch REST API feature. |
speed | float | 1.0 | Kept for API compatibility. Speed control is a batch REST API feature; the real-time stream plays at natural rate. |
quality | str | None | "low", "medium" or "high". Unset uses the server default. "low" roughly halves synthesis work and is the single largest first-audio saving available - but it drops a word in a measurable fraction of short clips, so it does not belong anywhere a number, date or reference code is read out. A session setting, not per utterance. |
clause_first | bool | None | Release the first piece of audio on a clause boundary instead of waiting for the opening sentence to end. Only helps when paired with a low min_buffer_frames - on its own it measured slower. See Latency tuning. |
min_buffer_frames | int | 12 | Audio frames buffered before the first frame is released (12 x frame_ms = 480 ms of audio at the defaults). Sized for WebRTC transports, where a starved encoder is audible. Keep the default on Daily and LiveKit. |
frame_ms | int | 40 | Size of each emitted audio frame, in milliseconds. |
Style tags
| Tag | Description |
|---|---|
<Neutral> | Clean read-speech by default |
<Happy> | Joyful, upbeat tone |
<Sad> | Somber, melancholic tone |
<Angry> | Forceful, intense tone |
<Fearful> | Anxious, trembling tone |
<Surprised> | Exclamatory, astonished tone |
<Disgust> | Repulsed, disapproving tone |
<News> | Formal news-anchor style |
<Conversational> | Casual, everyday speech - recommended for voice agents |
<Narrative> | Storytelling / audiobook delivery |
<Enthusiastic> | Energetic, passionate tone |
Text formatting
The service automatically prepends the style tag before sending to the API:
tts = ShunyalabsTTSService(voice="Rajesh", style="<Happy>")
# Input: "Welcome!"
# Sent: "<Happy> Welcome!"
Turn-taking
The gateway does its own endpointing, and can tell you when a speaker has finished rather than leaving you to infer it from a second voice-activity detector reading the same audio. Set emit_turn_frames=True to surface that as Pipecat turn frames.
emit_turn_frames=True only works when the user aggregator is configured with ExternalUserTurnStrategies, because those are the only strategies that consume these frames. Neither half works alone:
| Turn strategies | emit_turn_frames | Turn-start frames seen downstream |
|---|---|---|
| default | False | 1 - correct |
| default | True | 2 - duplicated |
| external | True | 1 - correct |
| external | False | 0 - no turn signal at all |
from pipecat.processors.aggregators.llm_response_universal import (
LLMContextAggregatorPair, LLMUserAggregatorParams,
)
from pipecat.turns.user_turn_strategies import ExternalUserTurnStrategies
from pipecat_shunyalabs import ShunyalabsSTTService
stt = ShunyalabsSTTService(
language="en",
emit_turn_frames=True, # both halves are required
)
aggregator = LLMContextAggregatorPair(
context,
user_params=LLMUserAggregatorParams(
user_turn_strategies=ExternalUserTurnStrategies(),
),
)
Two caveats worth knowing before you rely on this:
- The turn-start frame is derived from the first interim result, so it is only as prompt as
decode_every_ms(640 ms by default). That is fine for deciding whose turn it is, but too slow to drive barge-in - keep a transport voice-activity detector for that. - A speaker who never pauses produces no turn-end signal at all, because the gateway cuts the segment at its maximum length instead of endpointing. If your agent must respond regardless, pair this with your own timeout.
Latency tuning
Measured against the production services. Figures are first-audio or time-to-final on a warm connection; treat them as the shape of the trade-off rather than a promise for your network.
The one setting that matters most
endpoint_silence_ms is pure waiting: it is wall-clock delay after the speaker stops, before a transcript can be finalised. It defaults to 700 ms. A deployment running it at 3000 measured just over three seconds of perceived turn latency from this setting alone.
Lower is not simply better. At 250-300 ms a speaker who pauses mid-sentence - "a table for... four" - is cut off, and the agent answers half a sentence. Raise it if that happens; lower it if turns feel sluggish.
Taking the decision yourself
For full control, leave endpoint_silence_ms high as a backstop and call commit() when your own logic decides the speaker is done - a voice-activity detector, a push-to-talk release, a semantic turn detector. The gateway finalises immediately and keeps the connection open.
stt = ShunyalabsSTTService(
language="en",
endpoint_silence_ms=1500, # backstop only
)
# ... when your own detector says the speaker finished:
await stt.commit() # finalises now, connection stays open
Measured: 231 ms from commit() to the finalised transcript. Note this replaces the connection teardown that would otherwise happen per turn - a fresh handshake to the gateway measured around 300 ms from India, against an ASR that finalises in 39 ms.
Synthesis: what actually helps
Measured on a 130-character reply, warm connection, each setting isolated:
| Setting | First audio | Change |
|---|---|---|
| defaults | 1121 ms | - |
min_buffer_frames=3 alone | 1117 ms | -4 ms (no material effect) |
clause_first=True alone | 1371 ms | +250 ms - slower |
quality="low" alone | 974 ms | -146 ms |
clause_first + min_buffer_frames=3 | 1038 ms | -83 ms |
| all three | 925 ms | -196 ms |
Two results worth reading carefully. min_buffer_frames does little on its own, because the buffer is a threshold in audio duration and the service usually delivers faster than real time - 480 ms of audio can arrive in a few tens of milliseconds. And clause_first on its own is slower: a short opening clause does not reach the default 12-frame threshold until the next piece arrives, which is why the two settings belong together.
So quality carries most of the saving, and it is the one with a quality cost. Use it for throwaway acknowledgements, not for content a caller has to hear correctly.
Barge-in
When the caller interrupts, the service discards the in-flight utterance and keeps the session open, rather than dropping and re-establishing the connection. Measured at 88 ms to discard, with the next reply starting 246 ms later on the already-warm session. Pipecat triggers this automatically on an interruption; call it directly if you detect the interruption first:
await tts.discard_current_turn()
Telephony
For 8 kHz phone audio, pass sample_rate=8000 to the STT service. The gateway accepts it natively, so there is no resampling stage on the way in - measured at 251 ms to final, against 231 ms for the same speech at 16 kHz. Lower min_send_bytes to 1600 to keep the send buffer at 100 ms.
Full pipeline example
A complete voice agent with Shunyalabs STT and TTS, OpenAI LLM, and the Daily WebRTC transport:
import asyncio, os
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.runner import PipelineRunner
from pipecat.pipeline.task import PipelineParams, PipelineTask
from pipecat.processors.aggregators.openai_llm_context import (
OpenAILLMContext, OpenAILLMContextAggregator,
)
from pipecat.services.openai import OpenAILLMService
from pipecat.transports.services.daily import DailyParams, DailyTransport
from pipecat_shunyalabs import ShunyalabsSTTService, ShunyalabsTTSService
async def run_voice_agent(room_url: str, token: str):
transport = DailyTransport(
room_url, token, "Shunyalabs Agent",
DailyParams(audio_out_enabled=True, transcription_enabled=False),
)
stt = ShunyalabsSTTService(
api_key=os.environ["SHUNYALABS_API_KEY"],
language="auto",
sample_rate=16000,
)
llm = OpenAILLMService(
api_key=os.environ["OPENAI_API_KEY"],
model="gpt-4o",
)
messages = [{"role": "system", "content": "You are a helpful voice assistant."}]
context = OpenAILLMContext(messages)
context_aggregator = llm.create_context_aggregator(context)
tts = ShunyalabsTTSService(
api_key=os.environ["SHUNYALABS_API_KEY"],
voice="Rajesh",
language="hi",
style="<Conversational>",
)
pipeline = Pipeline([
transport.input(),
stt,
context_aggregator.user(),
llm,
tts,
transport.output(),
context_aggregator.assistant(),
])
task = PipelineTask(pipeline, PipelineParams(allow_interruptions=True, enable_metrics=True))
@transport.event_handler("on_first_participant_joined")
async def on_first_participant_joined(transport, participant):
await task.queue_frames([context_aggregator.user().get_context_frame()])
await PipelineRunner().run(task)
if __name__ == "__main__":
asyncio.run(run_voice_agent(
room_url=os.environ["DAILY_ROOM_URL"],
token=os.environ["DAILY_TOKEN"],
))
Multilingual example
# Hindi conversational bot
tts = ShunyalabsTTSService(voice="Rajesh", language="hi", style="<Conversational>")
# English news-style bot
tts = ShunyalabsTTSService(voice="Varun", language="en", style="<News>")
Custom endpoints
Both services can be repointed without changing code or upgrading the package. Resolution precedence: explicit argument → endpoint returned by the token service → environment variable → built-in default.
export SHUNYALABS_ASR_WS_URL="wss://<host>/v1/realtime"
export SHUNYALABS_TTS_WS_URL="wss://<host>/v1/realtime"
stt = ShunyalabsSTTService(url="wss://<host>/v1/realtime")
tts = ShunyalabsTTSService(url="wss://<host>/v1/realtime")
endpoints object, the SDK uses it automatically - so Shunya Labs can move an endpoint centrally, with no code change or package release on your side.
Error reference
| Exception | HTTP code | Description |
|---|---|---|
AuthenticationError | 401 | Invalid or missing API key. |
PermissionDeniedError | 403 | API key lacks permission for the resource. |
RateLimitError | 429 | Rate limit exceeded. Implement exponential backoff. |
ServerError | 5xx | Server-side error. Retried automatically. |
TimeoutError | — | Request exceeded timeout (default 60 s). |
TranscriptionError | — | ASR-specific failure (e.g. unsupported audio format). |
SynthesisError | — | TTS-specific failure (e.g. invalid voice parameter). |
Troubleshooting
| Symptom | Resolution |
|---|---|
AuthenticationError on startup | Verify SHUNYALABS_API_KEY is set and valid. |
| WebSocket connection refused | Ensure outbound WSS (port 443) is open to asrv2prod.shunyalabs.ai and ttsv2.shunyalabs.ai. |
| No transcription output | Check sample_rate matches your transport input. Verify audio source is active. |
| TTS audio silent or missing | Ensure output_format=pcm matches transport output. Verify TTSStartedFrame is received. |
| High latency on first TTS chunk | Deploy closer to the Shunyalabs gateway region (asia-south1). |
ImportError: pipecat_shunyalabs | Run pip install pipecat-shunyalabsai and confirm your virtual environment is activated. |