Shunya Labs DocsShunya Labs Docs
🌐 International
🇺🇸 English
🇯🇵 Japanese
🇨🇳 Chinese (Simplified)
🇹🇼 Chinese (Traditional)
🇸🇦 Arabic
🇩🇪 German
🇫🇷 French
🇪🇸 Spanish
🇧🇷 Portuguese
🇷🇺 Russian
🇰🇷 Korean
🇹🇷 Turkish
🇻🇳 Vietnamese
🇮🇩 Indonesian
🇮🇳 Hindi Belt
हिन्दी — Hindi
भोजपुरी — Bhojpuri
मैथिली — Maithili
राजस्थानी — Rajasthani
🇮🇳 South India
தமிழ் — Tamil
తెలుగు — Telugu
ಕನ್ನಡ — Kannada
മലയാളം — Malayalam
🇮🇳 West India
मराठी — Marathi
ગુજરાતી — Gujarati
कोंकणी — Konkani
🇮🇳 East India
বাংলা — Bengali
ଓଡ଼ିଆ — Odia
অসমীয়া — Assamese
🇮🇳 North-East India
মেইতেই — Meitei
नेपाली — Nepali
🇮🇳 North India
ਪੰਜਾਬੀ — Punjabi
اردو — Urdu
کٲشُر — Kashmiri
डोगरी — Dogri
سنڌي — Sindhi

Pipecat integration

pipecat-shunyalabsai v1.1.0

Native Shunyalabs STT and TTS services for Pipecat pipelines. Drop ShunyalabsSTTService and ShunyalabsTTSService into any Pipecat pipeline and get real-time streaming ASR with 46 speakers across 23 languages - no glue code required.

Requirements
Python 3.11+, Pipecat framework, and a valid Shunyalabs API key. Install a Pipecat transport (e.g. pipecat-ai[daily]) for WebRTC support.

Installation

Install the package from PyPI:

Terminal
pip install pipecat-shunyalabsai

To include a transport (e.g. Daily WebRTC):

Terminal
pip install pipecat-shunyalabsai pipecat-ai[daily]

Authentication

You need a Shunya Labs API key. Create one in the console - sign in at console.shunyalabs.ai, open API Keys, and click Create New Key (it is shown only once). Provide it as an environment variable (recommended) or pass it directly to the service classes:

Environment variable
export SHUNYALABS_API_KEY="your-api-key"
Python — inline
stt = ShunyalabsSTTService(api_key="your-api-key")
tts = ShunyalabsTTSService(api_key="your-api-key")
Your API key mints the token - the SDK does it for you
The ASR and TTS gateways never accept your API key directly. The plugin automatically exchanges your key for a short-lived access token - a JWT minted at app.shunyalabs.ai/api/auth/token - attaches it to every request, and refreshes it in the background before it expires. You only ever handle the API key; the token is generated, used, and rotated for you, and your raw key is never sent to the STT or TTS service.
Security
Never commit API keys to source control. Use a secrets manager (GCP Secret Manager, AWS Secrets Manager, HashiCorp Vault) in production.

Quick start

A minimal pipeline wiring Shunyalabs STT → OpenAI LLM → Shunyalabs TTS on a local audio transport:

Python
import asyncio, os
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.runner import PipelineRunner
from pipecat.pipeline.task import PipelineParams, PipelineTask
from pipecat.services.openai import OpenAILLMService
from pipecat.transports.local.audio import LocalAudioTransport
from pipecat_shunyalabs import ShunyalabsSTTService, ShunyalabsTTSService

async def main():
    transport = LocalAudioTransport()

    stt = ShunyalabsSTTService(
        api_key=os.environ["SHUNYALABS_API_KEY"],
        language="en",
    )

    llm = OpenAILLMService(
        api_key=os.environ["OPENAI_API_KEY"],
        model="gpt-4o",
    )

    tts = ShunyalabsTTSService(
        api_key=os.environ["SHUNYALABS_API_KEY"],
        voice="Rajesh",
        language="en",
        style="<Conversational>",
    )

    pipeline = Pipeline([transport.input(), stt, llm, tts, transport.output()])
    task = PipelineTask(pipeline, PipelineParams(allow_interruptions=True))
    await PipelineRunner().run(task)

if __name__ == "__main__":
    asyncio.run(main())

STT - ShunyalabsSTTService

Real-time streaming speech-to-text over a persistent WebSocket connection. Supports 23 Indian and international languages with optional automatic language detection.

Parameters

ParameterTypeDefaultDescription
api_keystrNoneAPI key. Falls back to SHUNYALABS_API_KEY env var.
languagestr"auto"Language code (e.g. "en", "hi") or "auto" for detection.
urlstrwss://asrv2prod.shunyalabs.ai/v1/realtimeWebSocket endpoint URL.
sample_rateint16000Expected audio sample rate in Hz. Must match transport input. 8000 is accepted natively, so telephony audio needs no client-side resampling.
endpoint_silence_msintNoneSilence before a transcript is finalised. The dominant control over how long a speaker waits after they stop talking. Unset uses the server default of 700 ms; clamped to 200-5000. See Latency tuning.
decode_every_msintNoneHow often interim results are produced. Unset uses the server default of 640 ms; clamped to 320-5000. Raising it lowers the cost of serving each stream at the price of coarser interim text. Does not affect how quickly a final arrives.
vadstrNone"silero" selects model-based endpointing. Worth setting on telephony audio, where energy thresholding can fail to register silence at all and natural endpointing then never fires.
codeswitchboolNoneRequest code-switch refinement: a finalised segment is re-rendered in the correct scripts and delivered as a second transcript. Subject to gateway support - check the echo (below) to confirm it was enabled.
modelstrNoneExplicit model or tier. Unset lets the gateway route on language.
min_send_bytesint3200Audio buffered before each send. The default is 100 ms at 16 kHz. Every buffered byte is latency in front of the gateway, so lower it proportionally for 8 kHz.
emit_turn_framesboolFalseEmit turn-start and turn-stop frames from the gateway's own endpointing. Off by default - see Turn-taking, which explains why it must be enabled together with Pipecat's external turn strategies.

Frame mapping

Shunyalabs eventPipecat frame
PARTIALInterimTranscriptionFrame - emitted continuously as speech is recognised
FINALTranscriptionFrame - a finalised segment. On a live connection this arrives once per utterance, not once per call.
FINAL_REFINEDTranscriptionFrame - the same segment re-rendered in correct scripts, when codeswitch is enabled. Arrives after the final, as its own transcript.
UTTERANCE_ENDUserStoppedSpeakingFrame - only when emit_turn_frames=True. See Turn-taking.
Python — STT example
from pipecat_shunyalabs import ShunyalabsSTTService

stt = ShunyalabsSTTService(
    language="hi",       # Hindi; use "auto" for detection
    sample_rate=16000,
)
Auto-reconnect
If the WebSocket connection drops during audio streaming, the service automatically reconnects and resumes sending audio.

TTS - ShunyalabsTTSService

Streaming text-to-speech over WebSocket. The plugin maintains a persistent WebSocket session that is reused across synthesis requests, streaming audio chunks back as TTSAudioRawFrame frames. Supports 46 speakers across 23 languages - any speaker can synthesise in any language.

Parameters

ParameterTypeDefaultDescription
api_keystrNoneAPI key. Falls back to SHUNYALABS_API_KEY env var.
urlstrwss://ttsv2.shunyalabs.ai/v1/realtimeWebSocket endpoint URL.
modelstr"zero-indic"TTS model identifier.
voicestr"Rajesh"Speaker voice name.
stylestrNoneEmotion / delivery style tag.
languagestr"en"Output language code.
output_formatstr"pcm"Kept for API compatibility. The real-time stream always delivers PCM (24 kHz, 16-bit mono); other container formats are a batch REST API feature.
speedfloat1.0Kept for API compatibility. Speed control is a batch REST API feature; the real-time stream plays at natural rate.
qualitystrNone"low", "medium" or "high". Unset uses the server default. "low" roughly halves synthesis work and is the single largest first-audio saving available - but it drops a word in a measurable fraction of short clips, so it does not belong anywhere a number, date or reference code is read out. A session setting, not per utterance.
clause_firstboolNoneRelease the first piece of audio on a clause boundary instead of waiting for the opening sentence to end. Only helps when paired with a low min_buffer_frames - on its own it measured slower. See Latency tuning.
min_buffer_framesint12Audio frames buffered before the first frame is released (12 x frame_ms = 480 ms of audio at the defaults). Sized for WebRTC transports, where a starved encoder is audible. Keep the default on Daily and LiveKit.
frame_msint40Size of each emitted audio frame, in milliseconds.

Style tags

TagDescription
<Neutral>Clean read-speech by default
<Happy>Joyful, upbeat tone
<Sad>Somber, melancholic tone
<Angry>Forceful, intense tone
<Fearful>Anxious, trembling tone
<Surprised>Exclamatory, astonished tone
<Disgust>Repulsed, disapproving tone
<News>Formal news-anchor style
<Conversational>Casual, everyday speech - recommended for voice agents
<Narrative>Storytelling / audiobook delivery
<Enthusiastic>Energetic, passionate tone

Text formatting

The service automatically prepends the style tag before sending to the API:

Python
tts = ShunyalabsTTSService(voice="Rajesh", style="<Happy>")
# Input:  "Welcome!"
# Sent:   "<Happy> Welcome!"

Turn-taking

The gateway does its own endpointing, and can tell you when a speaker has finished rather than leaving you to infer it from a second voice-activity detector reading the same audio. Set emit_turn_frames=True to surface that as Pipecat turn frames.

Enable it together with external turn strategies
emit_turn_frames=True only works when the user aggregator is configured with ExternalUserTurnStrategies, because those are the only strategies that consume these frames. Neither half works alone:
Turn strategiesemit_turn_framesTurn-start frames seen downstream
defaultFalse1 - correct
defaultTrue2 - duplicated
externalTrue1 - correct
externalFalse0 - no turn signal at all
Pipecat's default strategies listen for the transport's own voice-activity frame, not the public turn frame, so the aggregator raises its own turn and the service's passes through as well.
Python — ASR-driven turn-taking
from pipecat.processors.aggregators.llm_response_universal import (
    LLMContextAggregatorPair, LLMUserAggregatorParams,
)
from pipecat.turns.user_turn_strategies import ExternalUserTurnStrategies
from pipecat_shunyalabs import ShunyalabsSTTService

stt = ShunyalabsSTTService(
    language="en",
    emit_turn_frames=True,          # both halves are required
)

aggregator = LLMContextAggregatorPair(
    context,
    user_params=LLMUserAggregatorParams(
        user_turn_strategies=ExternalUserTurnStrategies(),
    ),
)

Two caveats worth knowing before you rely on this:

  • The turn-start frame is derived from the first interim result, so it is only as prompt as decode_every_ms (640 ms by default). That is fine for deciding whose turn it is, but too slow to drive barge-in - keep a transport voice-activity detector for that.
  • A speaker who never pauses produces no turn-end signal at all, because the gateway cuts the segment at its maximum length instead of endpointing. If your agent must respond regardless, pair this with your own timeout.

Latency tuning

Measured against the production services. Figures are first-audio or time-to-final on a warm connection; treat them as the shape of the trade-off rather than a promise for your network.

The one setting that matters most

endpoint_silence_ms is pure waiting: it is wall-clock delay after the speaker stops, before a transcript can be finalised. It defaults to 700 ms. A deployment running it at 3000 measured just over three seconds of perceived turn latency from this setting alone.

Lower is not simply better. At 250-300 ms a speaker who pauses mid-sentence - "a table for... four" - is cut off, and the agent answers half a sentence. Raise it if that happens; lower it if turns feel sluggish.

Taking the decision yourself

For full control, leave endpoint_silence_ms high as a backstop and call commit() when your own logic decides the speaker is done - a voice-activity detector, a push-to-talk release, a semantic turn detector. The gateway finalises immediately and keeps the connection open.

Python — commit on your own endpointing
stt = ShunyalabsSTTService(
    language="en",
    endpoint_silence_ms=1500,   # backstop only
)

# ... when your own detector says the speaker finished:
await stt.commit()              # finalises now, connection stays open

Measured: 231 ms from commit() to the finalised transcript. Note this replaces the connection teardown that would otherwise happen per turn - a fresh handshake to the gateway measured around 300 ms from India, against an ASR that finalises in 39 ms.

Synthesis: what actually helps

Measured on a 130-character reply, warm connection, each setting isolated:

SettingFirst audioChange
defaults1121 ms-
min_buffer_frames=3 alone1117 ms-4 ms (no material effect)
clause_first=True alone1371 ms+250 ms - slower
quality="low" alone974 ms-146 ms
clause_first + min_buffer_frames=31038 ms-83 ms
all three925 ms-196 ms

Two results worth reading carefully. min_buffer_frames does little on its own, because the buffer is a threshold in audio duration and the service usually delivers faster than real time - 480 ms of audio can arrive in a few tens of milliseconds. And clause_first on its own is slower: a short opening clause does not reach the default 12-frame threshold until the next piece arrives, which is why the two settings belong together.

So quality carries most of the saving, and it is the one with a quality cost. Use it for throwaway acknowledgements, not for content a caller has to hear correctly.

Barge-in

When the caller interrupts, the service discards the in-flight utterance and keeps the session open, rather than dropping and re-establishing the connection. Measured at 88 ms to discard, with the next reply starting 246 ms later on the already-warm session. Pipecat triggers this automatically on an interruption; call it directly if you detect the interruption first:

Python — stop speaking immediately
await tts.discard_current_turn()

Telephony

For 8 kHz phone audio, pass sample_rate=8000 to the STT service. The gateway accepts it natively, so there is no resampling stage on the way in - measured at 251 ms to final, against 231 ms for the same speech at 16 kHz. Lower min_send_bytes to 1600 to keep the send buffer at 100 ms.

Confirming what took effect
The gateway clamps these settings to its own supported ranges and can decline optional features, reporting the result rather than rejecting the request. The service logs a warning when a value it sent was clamped. Requests that are out of range are adjusted silently at the protocol level, so a setting that appears to have no effect is worth checking against that log line.

Full pipeline example

A complete voice agent with Shunyalabs STT and TTS, OpenAI LLM, and the Daily WebRTC transport:

Python
import asyncio, os
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.runner import PipelineRunner
from pipecat.pipeline.task import PipelineParams, PipelineTask
from pipecat.processors.aggregators.openai_llm_context import (
    OpenAILLMContext, OpenAILLMContextAggregator,
)
from pipecat.services.openai import OpenAILLMService
from pipecat.transports.services.daily import DailyParams, DailyTransport
from pipecat_shunyalabs import ShunyalabsSTTService, ShunyalabsTTSService

async def run_voice_agent(room_url: str, token: str):
    transport = DailyTransport(
        room_url, token, "Shunyalabs Agent",
        DailyParams(audio_out_enabled=True, transcription_enabled=False),
    )

    stt = ShunyalabsSTTService(
        api_key=os.environ["SHUNYALABS_API_KEY"],
        language="auto",
        sample_rate=16000,
    )

    llm = OpenAILLMService(
        api_key=os.environ["OPENAI_API_KEY"],
        model="gpt-4o",
    )

    messages = [{"role": "system", "content": "You are a helpful voice assistant."}]
    context = OpenAILLMContext(messages)
    context_aggregator = llm.create_context_aggregator(context)

    tts = ShunyalabsTTSService(
        api_key=os.environ["SHUNYALABS_API_KEY"],
        voice="Rajesh",
        language="hi",
        style="<Conversational>",
    )

    pipeline = Pipeline([
        transport.input(),
        stt,
        context_aggregator.user(),
        llm,
        tts,
        transport.output(),
        context_aggregator.assistant(),
    ])

    task = PipelineTask(pipeline, PipelineParams(allow_interruptions=True, enable_metrics=True))

    @transport.event_handler("on_first_participant_joined")
    async def on_first_participant_joined(transport, participant):
        await task.queue_frames([context_aggregator.user().get_context_frame()])

    await PipelineRunner().run(task)

if __name__ == "__main__":
    asyncio.run(run_voice_agent(
        room_url=os.environ["DAILY_ROOM_URL"],
        token=os.environ["DAILY_TOKEN"],
    ))

Multilingual example

Python
# Hindi conversational bot
tts = ShunyalabsTTSService(voice="Rajesh", language="hi", style="<Conversational>")

# English news-style bot
tts = ShunyalabsTTSService(voice="Varun", language="en", style="<News>")

Custom endpoints

Both services can be repointed without changing code or upgrading the package. Resolution precedence: explicit argument → endpoint returned by the token service → environment variable → built-in default.

Environment variables (real-time = WSS)
export SHUNYALABS_ASR_WS_URL="wss://<host>/v1/realtime"
export SHUNYALABS_TTS_WS_URL="wss://<host>/v1/realtime"
Python — per instance
stt = ShunyalabsSTTService(url="wss://<host>/v1/realtime")
tts = ShunyalabsTTSService(url="wss://<host>/v1/realtime")
Zero-downtime repointing
If the token service returns an endpoints object, the SDK uses it automatically - so Shunya Labs can move an endpoint centrally, with no code change or package release on your side.

Error reference

ExceptionHTTP codeDescription
AuthenticationError401Invalid or missing API key.
PermissionDeniedError403API key lacks permission for the resource.
RateLimitError429Rate limit exceeded. Implement exponential backoff.
ServerError5xxServer-side error. Retried automatically.
TimeoutError—Request exceeded timeout (default 60 s).
TranscriptionError—ASR-specific failure (e.g. unsupported audio format).
SynthesisError—TTS-specific failure (e.g. invalid voice parameter).

Troubleshooting

SymptomResolution
AuthenticationError on startupVerify SHUNYALABS_API_KEY is set and valid.
WebSocket connection refusedEnsure outbound WSS (port 443) is open to asrv2prod.shunyalabs.ai and ttsv2.shunyalabs.ai.
No transcription outputCheck sample_rate matches your transport input. Verify audio source is active.
TTS audio silent or missingEnsure output_format=pcm matches transport output. Verify TTSStartedFrame is received.
High latency on first TTS chunkDeploy closer to the Shunyalabs gateway region (asia-south1).
ImportError: pipecat_shunyalabsRun pip install pipecat-shunyalabsai and confirm your virtual environment is activated.
Pipecat Integration | Shunya Labs Docs