LiveKit integration
livekit-plugins-shunyalabsai v1.1.1Official Shunyalabs plugin for LiveKit Agents. Plug shunyalabs.STT and shunyalabs.TTS directly into a LiveKit AgentSession - supports real-time streaming transcription, batch recognition, and high-fidelity multilingual voice synthesis.
Installation
pip install livekit-plugins-shunyalabsai
Authentication
You need a Shunya Labs API key. Create one in the console - sign in at console.shunyalabs.ai, open API Keys, and click Create New Key (it is shown only once). Provide it as an environment variable or pass it directly to the plugin classes:
export SHUNYALABS_API_KEY="your-api-key"
stt = shunyalabs.STT(api_key="your-api-key")
tts = shunyalabs.TTS(api_key="your-api-key")
app.shunyalabs.ai/api/auth/token - attaches it to every request, and refreshes it in the background before it expires. You only ever handle the API key; the token is generated, used, and rotated for you, and your raw key is never sent to the STT or TTS service.
Quick start
The minimal wiring - pass shunyalabs.STT and shunyalabs.TTS into an AgentSession:
from livekit.agents import AgentSession
from livekit.plugins import shunyalabs, silero
session = AgentSession(
stt=shunyalabs.STT(language="en"),
tts=shunyalabs.TTS(voice="Rajesh", style="<Neutral>"),
vad=silero.VAD.load(),
)
STT - shunyalabs.STT
Streaming and batch speech-to-text backed by the Shunyalabs ASR service. Audio frames from LiveKit are streamed over WebSocket; transcription events are pushed back as SpeechEvents.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | str | None | API key. Falls back to SHUNYALABS_API_KEY env var. |
language | str | "auto" | BCP-47 language code or "auto" for detection. |
api_url | str | https://asrv2prod.shunyalabs.ai | REST batch endpoint base URL. |
ws_url | str | wss://asrv2prod.shunyalabs.ai/v1/realtime | WebSocket streaming endpoint URL. |
endpoint_silence_ms | int | None | Silence before a transcript is finalised. The dominant control over how long a speaker waits after they stop talking. Unset uses the server default of 700 ms; clamped to 200-5000. See Latency tuning. |
decode_every_ms | int | None | How often interim results are produced. Unset uses the server default of 640 ms; clamped to 320-5000. Raising it lowers the cost of serving each stream at the price of coarser interim text. Does not affect how quickly a final arrives. |
vad | str | None | "silero" selects model-based endpointing. Worth setting on telephony audio, where energy thresholding can fail to register silence at all and natural endpointing then never fires. |
codeswitch | bool | None | Request code-switch refinement, delivered as an additional final transcript once the segment has been re-rendered in the correct scripts. Subject to gateway support. |
model | str | None | Explicit model or tier. Unset lets the gateway route on language, and STT.model reports vak-v3. |
Capabilities
| Capability | Supported |
|---|---|
| Streaming (real-time) | Yes |
| Interim results | Yes |
| Offline / batch recognition | Yes |
Streaming STT
Real-time transcription over WebSocket with event mapping to LiveKit's SpeechEventType:
| Shunyalabs event | LiveKit SpeechEventType |
|---|---|
PARTIAL | INTERIM_TRANSCRIPT |
FINAL (with text) | FINAL_TRANSCRIPT + RECOGNITION_USAGE + END_OF_SPEECH, in that order |
FINAL_REFINED | FINAL_TRANSCRIPT - the same segment re-rendered in correct scripts, when codeswitch is enabled |
END_OF_SPEECH follows the final transcript for the same utterance, so a handler that reacts to it always has the text already. It is emitted exactly once per utterance that carried speech.
Note that a silent line is not idle: the service keeps endpointing while audio continues to arrive, producing empty results roughly once a second for as long as nobody speaks. Those are not turn boundaries and the plugin discards them, so they generate no events. Versions before 1.1.1 surfaced them, which produced repeat
END_OF_SPEECH events - and sometimes one ahead of the transcript. Upgrade if you are pinned below 1.1.1.
from livekit.agents import AgentSession
from livekit.plugins import shunyalabs, silero
session = AgentSession(
stt=shunyalabs.STT(language="en"),
vad=silero.VAD.load(),
)
@session.on("user_speech_committed")
def on_speech(ev):
print(f"User said: {ev.transcript}")
Batch STT
Single-shot transcription of an audio buffer via POST /v1/audio/transcriptions:
from livekit.plugins import shunyalabs
stt = shunyalabs.STT(language="en")
# Inside an agent context:
event = await stt.recognize(audio_buffer)
print(event.alternatives[0].text)
TTS - shunyalabs.TTS
Streaming and chunked text-to-speech. Token-by-token streaming collects text then synthesises on flush via WebSocket; the batch API handles single-shot synthesis over HTTP.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | str | None | API key. Falls back to SHUNYALABS_API_KEY env var. |
api_url | str | https://ttsv2.shunyalabs.ai | HTTP batch endpoint base URL. |
ws_url | str | wss://ttsv2.shunyalabs.ai/v1/realtime | WebSocket streaming endpoint URL. |
model | str | "zero-indic" | TTS model name. |
voice | str | "Rajesh" | Voice name for the API. |
style | str | None | Emotion style tag. |
language | str | "en" | Language code for transliteration. |
sample_rate | int | 24000 | Output audio sample rate in Hz. The gateway emits 24 kHz PCM on both the streaming and batch paths. |
output_format | str | "pcm" | Audio format for the batch (synthesize) path: pcm, wav, mp3, ogg_opus, flac. The real-time stream is always PCM. |
speed | float | 1.0 | Speaking speed multiplier (0.25 – 4.0). |
Style tags
| Tag | Description |
|---|---|
<Neutral> | Neutral tone by default |
<Happy> | Happy / cheerful |
<Sad> | Sad / melancholic |
<Angry> | Angry / intense |
<Fearful> | Fearful / anxious |
<Surprised> | Surprised / excited |
<Disgust> | Disgusted |
<News> | News anchor style |
<Conversational> | Casual conversational - recommended for voice agents |
<Narrative> | Storytelling / narration |
<Enthusiastic> | Enthusiastic / energetic |
Text formatting
The plugin automatically prepends the style tag before sending text to the API:
tts = shunyalabs.TTS(voice="Rajesh", style="<Happy>")
# Input: "Welcome to our platform"
# Sent: "<Happy> Welcome to our platform"
Streaming TTS example
from livekit.agents import AgentSession
from livekit.plugins import shunyalabs
session = AgentSession(
tts=shunyalabs.TTS(
style="<Conversational>",
model="zero-indic",
voice="Nisha",
),
)
Chunked (batch) TTS example
from livekit.plugins import shunyalabs
tts = shunyalabs.TTS(voice="Varun")
stream = tts.synthesize("Hello, how can I help you today?")
Latency tuning
endpoint_silence_ms is the setting that decides how long a speaker waits after they stop talking, because it is pure wall-clock delay before a transcript can be finalised. It defaults to 700 ms. A deployment running it at 3000 measured just over three seconds of perceived turn latency from this setting alone.
Lower is not simply better. At 250-300 ms a speaker who pauses mid-sentence - "a table for... four" - is cut off, and the agent answers half a sentence. Raise it if that happens; lower it if turns feel sluggish.
from livekit.plugins import shunyalabs
stt = shunyalabs.STT(
language="en",
endpoint_silence_ms=400, # default is 700
vad="silero", # worth it on telephony audio
)
Two supporting settings:
vad="silero"selects model-based endpointing. On a noisy line, energy thresholding can fail to register silence at all - in which case natural endpointing never fires and the segment runs to its maximum length instead.decode_every_mscontrols how often interim transcripts are produced, and so how much work serving each stream costs. Raising it does not delay the final. Lower it only if you display interim text and want it to update more smoothly.
Custom endpoints
Both services can be repointed without changing code or upgrading the package. Resolution precedence: explicit argument → endpoint returned by the token service → environment variable → built-in default.
export SHUNYALABS_ASR_URL="https://<host>" # batch
export SHUNYALABS_ASR_WS_URL="wss://<host>/v1/realtime" # streaming
export SHUNYALABS_TTS_URL="https://<host>"
export SHUNYALABS_TTS_WS_URL="wss://<host>/v1/realtime"
stt = shunyalabs.STT(api_url="https://<host>", ws_url="wss://<host>/v1/realtime")
tts = shunyalabs.TTS(api_url="https://<host>", ws_url="wss://<host>/v1/realtime")
endpoints object, the SDK uses it automatically - so Shunya Labs can move an endpoint centrally, with no code change or package release on your side.
Full agent example
import asyncio
from livekit.agents import AgentSession, Agent, RoomInputOptions
from livekit.plugins import shunyalabs, silero
class MyAgent(Agent):
def __init__(self):
super().__init__(instructions="You are a helpful voice assistant.")
async def entrypoint(ctx):
session = AgentSession(
stt=shunyalabs.STT(language="auto"),
tts=shunyalabs.TTS(
model="zero-indic",
voice="Rajesh",
style="<Conversational>",
),
vad=silero.VAD.load(),
)
await session.start(
agent=MyAgent(),
room=ctx.room,
room_input_options=RoomInputOptions(),
)
Multilingual example
# Hindi speaker
tts_hindi = shunyalabs.TTS(
voice="Rajesh",
language="hi", style="<Neutral>",
)
# English speaker
tts_english = shunyalabs.TTS(
voice="Varun",
language="en", style="<Conversational>",
)