Shunya Labs DocsShunya Labs Docs
🌐 International
🇺🇸 English
🇯🇵 Japanese
🇨🇳 Chinese (Simplified)
🇹🇼 Chinese (Traditional)
🇸🇦 Arabic
🇩🇪 German
🇫🇷 French
🇪🇸 Spanish
🇧🇷 Portuguese
🇷🇺 Russian
🇰🇷 Korean
🇹🇷 Turkish
🇻🇳 Vietnamese
🇮🇩 Indonesian
🇮🇳 Hindi Belt
हिन्दी — Hindi
भोजपुरी — Bhojpuri
मैथिली — Maithili
राजस्थानी — Rajasthani
🇮🇳 South India
தமிழ் — Tamil
తెలుగు — Telugu
ಕನ್ನಡ — Kannada
മലയാളം — Malayalam
🇮🇳 West India
मराठी — Marathi
ગુજરાતી — Gujarati
कोंकणी — Konkani
🇮🇳 East India
বাংলা — Bengali
ଓଡ଼ିଆ — Odia
অসমীয়া — Assamese
🇮🇳 North-East India
মেইতেই — Meitei
नेपाली — Nepali
🇮🇳 North India
ਪੰਜਾਬੀ — Punjabi
اردو — Urdu
کٲشُر — Kashmiri
डोगरी — Dogri
سنڌي — Sindhi

Official Shunyalabs plugin for LiveKit Agents. Plug shunyalabs.STT and shunyalabs.TTS directly into a LiveKit AgentSession - supports real-time streaming transcription, batch recognition, and high-fidelity multilingual voice synthesis.

Requirements
Python 3.10+, LiveKit Agents framework, and a valid Shunyalabs API key.

Installation

Terminal
pip install livekit-plugins-shunyalabsai

Authentication

You need a Shunya Labs API key. Create one in the console - sign in at console.shunyalabs.ai, open API Keys, and click Create New Key (it is shown only once). Provide it as an environment variable or pass it directly to the plugin classes:

Environment variable
export SHUNYALABS_API_KEY="your-api-key"
Python — inline
stt = shunyalabs.STT(api_key="your-api-key")
tts = shunyalabs.TTS(api_key="your-api-key")
Your API key mints the token - the SDK does it for you
The ASR and TTS gateways never accept your API key directly. The plugin automatically exchanges your key for a short-lived access token - a JWT minted at app.shunyalabs.ai/api/auth/token - attaches it to every request, and refreshes it in the background before it expires. You only ever handle the API key; the token is generated, used, and rotated for you, and your raw key is never sent to the STT or TTS service.

Quick start

The minimal wiring - pass shunyalabs.STT and shunyalabs.TTS into an AgentSession:

Python
from livekit.agents import AgentSession
from livekit.plugins import shunyalabs, silero

session = AgentSession(
    stt=shunyalabs.STT(language="en"),
    tts=shunyalabs.TTS(voice="Rajesh", style="<Neutral>"),
    vad=silero.VAD.load(),
)

STT - shunyalabs.STT

Streaming and batch speech-to-text backed by the Shunyalabs ASR service. Audio frames from LiveKit are streamed over WebSocket; transcription events are pushed back as SpeechEvents.

Parameters

ParameterTypeDefaultDescription
api_keystrNoneAPI key. Falls back to SHUNYALABS_API_KEY env var.
languagestr"auto"BCP-47 language code or "auto" for detection.
api_urlstrhttps://asrv2prod.shunyalabs.aiREST batch endpoint base URL.
ws_urlstrwss://asrv2prod.shunyalabs.ai/v1/realtimeWebSocket streaming endpoint URL.
endpoint_silence_msintNoneSilence before a transcript is finalised. The dominant control over how long a speaker waits after they stop talking. Unset uses the server default of 700 ms; clamped to 200-5000. See Latency tuning.
decode_every_msintNoneHow often interim results are produced. Unset uses the server default of 640 ms; clamped to 320-5000. Raising it lowers the cost of serving each stream at the price of coarser interim text. Does not affect how quickly a final arrives.
vadstrNone"silero" selects model-based endpointing. Worth setting on telephony audio, where energy thresholding can fail to register silence at all and natural endpointing then never fires.
codeswitchboolNoneRequest code-switch refinement, delivered as an additional final transcript once the segment has been re-rendered in the correct scripts. Subject to gateway support.
modelstrNoneExplicit model or tier. Unset lets the gateway route on language, and STT.model reports vak-v3.

Capabilities

CapabilitySupported
Streaming (real-time)Yes
Interim resultsYes
Offline / batch recognitionYes

Streaming STT

Real-time transcription over WebSocket with event mapping to LiveKit's SpeechEventType:

Shunyalabs eventLiveKit SpeechEventType
PARTIALINTERIM_TRANSCRIPT
FINAL (with text)FINAL_TRANSCRIPT + RECOGNITION_USAGE + END_OF_SPEECH, in that order
FINAL_REFINEDFINAL_TRANSCRIPT - the same segment re-rendered in correct scripts, when codeswitch is enabled
End of speech
END_OF_SPEECH follows the final transcript for the same utterance, so a handler that reacts to it always has the text already. It is emitted exactly once per utterance that carried speech.

Note that a silent line is not idle: the service keeps endpointing while audio continues to arrive, producing empty results roughly once a second for as long as nobody speaks. Those are not turn boundaries and the plugin discards them, so they generate no events. Versions before 1.1.1 surfaced them, which produced repeat END_OF_SPEECH events - and sometimes one ahead of the transcript. Upgrade if you are pinned below 1.1.1.
Python — streaming STT
from livekit.agents import AgentSession
from livekit.plugins import shunyalabs, silero

session = AgentSession(
    stt=shunyalabs.STT(language="en"),
    vad=silero.VAD.load(),
)

@session.on("user_speech_committed")
def on_speech(ev):
    print(f"User said: {ev.transcript}")

Batch STT

Single-shot transcription of an audio buffer via POST /v1/audio/transcriptions:

Python — batch STT
from livekit.plugins import shunyalabs

stt = shunyalabs.STT(language="en")

# Inside an agent context:
event = await stt.recognize(audio_buffer)
print(event.alternatives[0].text)

TTS - shunyalabs.TTS

Streaming and chunked text-to-speech. Token-by-token streaming collects text then synthesises on flush via WebSocket; the batch API handles single-shot synthesis over HTTP.

Parameters

ParameterTypeDefaultDescription
api_keystrNoneAPI key. Falls back to SHUNYALABS_API_KEY env var.
api_urlstrhttps://ttsv2.shunyalabs.aiHTTP batch endpoint base URL.
ws_urlstrwss://ttsv2.shunyalabs.ai/v1/realtimeWebSocket streaming endpoint URL.
modelstr"zero-indic"TTS model name.
voicestr"Rajesh"Voice name for the API.
stylestrNoneEmotion style tag.
languagestr"en"Language code for transliteration.
sample_rateint24000Output audio sample rate in Hz. The gateway emits 24 kHz PCM on both the streaming and batch paths.
output_formatstr"pcm"Audio format for the batch (synthesize) path: pcm, wav, mp3, ogg_opus, flac. The real-time stream is always PCM.
speedfloat1.0Speaking speed multiplier (0.25 – 4.0).

Style tags

TagDescription
<Neutral>Neutral tone by default
<Happy>Happy / cheerful
<Sad>Sad / melancholic
<Angry>Angry / intense
<Fearful>Fearful / anxious
<Surprised>Surprised / excited
<Disgust>Disgusted
<News>News anchor style
<Conversational>Casual conversational - recommended for voice agents
<Narrative>Storytelling / narration
<Enthusiastic>Enthusiastic / energetic

Text formatting

The plugin automatically prepends the style tag before sending text to the API:

Python
tts = shunyalabs.TTS(voice="Rajesh", style="<Happy>")
# Input:  "Welcome to our platform"
# Sent:   "<Happy> Welcome to our platform"

Streaming TTS example

Python
from livekit.agents import AgentSession
from livekit.plugins import shunyalabs

session = AgentSession(
    tts=shunyalabs.TTS(
        style="<Conversational>",
        model="zero-indic",
        voice="Nisha",
    ),
)

Chunked (batch) TTS example

Python
from livekit.plugins import shunyalabs

tts = shunyalabs.TTS(voice="Varun")
stream = tts.synthesize("Hello, how can I help you today?")

Latency tuning

endpoint_silence_ms is the setting that decides how long a speaker waits after they stop talking, because it is pure wall-clock delay before a transcript can be finalised. It defaults to 700 ms. A deployment running it at 3000 measured just over three seconds of perceived turn latency from this setting alone.

Lower is not simply better. At 250-300 ms a speaker who pauses mid-sentence - "a table for... four" - is cut off, and the agent answers half a sentence. Raise it if that happens; lower it if turns feel sluggish.

Python — snappier turns
from livekit.plugins import shunyalabs

stt = shunyalabs.STT(
    language="en",
    endpoint_silence_ms=400,   # default is 700
    vad="silero",              # worth it on telephony audio
)

Two supporting settings:

  • vad="silero" selects model-based endpointing. On a noisy line, energy thresholding can fail to register silence at all - in which case natural endpointing never fires and the segment runs to its maximum length instead.
  • decode_every_ms controls how often interim transcripts are produced, and so how much work serving each stream costs. Raising it does not delay the final. Lower it only if you display interim text and want it to update more smoothly.
Values are clamped, not rejected
The service adjusts out-of-range settings to its supported limits rather than returning an error, so a value that appears to have no effect may have been clamped. The plugin logs a warning when that happens - check for it before concluding a setting is ignored.

Custom endpoints

Both services can be repointed without changing code or upgrading the package. Resolution precedence: explicit argument → endpoint returned by the token service → environment variable → built-in default.

Environment variables (batch = HTTP, streaming = WSS)
export SHUNYALABS_ASR_URL="https://<host>"                # batch
export SHUNYALABS_ASR_WS_URL="wss://<host>/v1/realtime"     # streaming
export SHUNYALABS_TTS_URL="https://<host>"
export SHUNYALABS_TTS_WS_URL="wss://<host>/v1/realtime"
Python — per instance
stt = shunyalabs.STT(api_url="https://<host>", ws_url="wss://<host>/v1/realtime")
tts = shunyalabs.TTS(api_url="https://<host>", ws_url="wss://<host>/v1/realtime")
Zero-downtime repointing
If the token service returns an endpoints object, the SDK uses it automatically - so Shunya Labs can move an endpoint centrally, with no code change or package release on your side.

Full agent example

Python
import asyncio
from livekit.agents import AgentSession, Agent, RoomInputOptions
from livekit.plugins import shunyalabs, silero

class MyAgent(Agent):
    def __init__(self):
        super().__init__(instructions="You are a helpful voice assistant.")

async def entrypoint(ctx):
    session = AgentSession(
        stt=shunyalabs.STT(language="auto"),
        tts=shunyalabs.TTS(
            model="zero-indic",
            voice="Rajesh",
            style="<Conversational>",
        ),
        vad=silero.VAD.load(),
    )
    await session.start(
        agent=MyAgent(),
        room=ctx.room,
        room_input_options=RoomInputOptions(),
    )

Multilingual example

Python
# Hindi speaker
tts_hindi = shunyalabs.TTS(
    voice="Rajesh",
    language="hi", style="<Neutral>",
)

# English speaker
tts_english = shunyalabs.TTS(
    voice="Varun",
    language="en", style="<Conversational>",
)
Pipecat
Back to
Deployment overview
← Previous
Pipecat
Back to
Deployment overview
LiveKit Integration | Shunya Labs Docs