Shunya Labs DocsShunya Labs Docs
🌐 International
🇺🇸 English
🇯🇵 Japanese
🇨🇳 Chinese (Simplified)
🇹🇼 Chinese (Traditional)
🇸🇦 Arabic
🇩🇪 German
🇫🇷 French
🇪🇸 Spanish
🇧🇷 Portuguese
🇷🇺 Russian
🇰🇷 Korean
🇹🇷 Turkish
🇻🇳 Vietnamese
🇮🇩 Indonesian
🇮🇳 Hindi Belt
हिन्दी — Hindi
भोजपुरी — Bhojpuri
मैथिली — Maithili
राजस्थानी — Rajasthani
🇮🇳 South India
தமிழ் — Tamil
తెలుగు — Telugu
ಕನ್ನಡ — Kannada
മലയാളം — Malayalam
🇮🇳 West India
मराठी — Marathi
ગુજરાતી — Gujarati
कोंकणी — Konkani
🇮🇳 East India
বাংলা — Bengali
ଓଡ଼ିଆ — Odia
অসমীয়া — Assamese
🇮🇳 North-East India
মেইতেই — Meitei
नेपाली — Nepali
🇮🇳 North India
ਪੰਜਾਬੀ — Punjabi
اردو — Urdu
کٲشُر — Kashmiri
डोगरी — Dogri
سنڌي — Sindhi

Streaming ASR over WebSocket

For voice agents, IVR, and live captioning. You open a WebSocket, send audio frames as they arrive, and receive partial and final transcripts in real time.

First: get an access token

The examples below send Authorization: Bearer $ACCESS_TOKEN. The speech APIs accept only a short-lived access token — never your API key directly.

From the console (recommended). Open the console, click Generate token next to your API key, and copy it — then set it:

export ACCESS_TOKEN="eyJhbGciOiJSUzI1NiIs…paste-here"

Or mint it from your API key — the path for production, where your app refreshes the token as it nears expiry (the response carries expires_in):

export ACCESS_TOKEN=$(curl -s -X POST https://app.shunyalabs.ai/api/auth/token \
  -H "api-key: $SHUNYALABS_API_KEY" | jq -r .token)

Endpoint

wss://asrv2prod.shunyalabs.ai/v1/realtime
Writing a correct client
Six rules cover almost every integration problem on this endpoint.
  • There are four event types, plus one you only see if you ask for it: ready, partial, final, error — and final_refined when you set codeswitch in the opening frame. Build the transcript from final; partial is provisional and will be revised. Do not wait for a separate completion event before using the text.
  • Result events carry seg and elapsed_ms. seg increments per segment, so it is what you order and de-duplicate on.
  • final and partial carry confidence; final also carries audio_duration_sec. confidence is the mean per-token confidence for that decode, 0–1, and it tracks quality closely: on a real call a spurious one-word fragment scored0.055 and a repetition loop 0.412, against0.74–0.94 for correct finals. audio_duration_sec is the seconds of audio that segment decoded. Use them together — a short transcript with low confidence from a long segment is the pattern worth rejecting, and you choose the threshold. Both are omitted rather than faked when a recogniser returns no per-token data, so an absent confidence means “unavailable”, never “low”. final_refined carries no confidence: it is rewritten text, so a decoder score would not describe it.
  • One turn can emit several finals. A long utterance is split mid-stream, so a single turn may produce more than one final. Read until the server goes quiet and concatenate in seg order — code that assumes exactly one final per turn drifts a turn further behind on every turn.
  • final_refined arrives after the final, not instead of it. With codeswitch on, a segment is delivered twice: the raw final immediately, then a refined version of the same seg a moment later. Nothing is held back waiting for it, so keep rendering the final and replace that seg if the refinement turns up. Match on seg, not on arrival order.
  • Send commit between turns, and end once at the close. commit finalizes the current turn and keeps the socket open; end closes it, so using end per utterance forces a fresh handshake on every turn and adds the round trips back to your latency.
  • Send an explicit language. Omitting it (or sending auto) is accepted and will attempt detection, but a live stream has to decide from the opening seconds of audio, so treat it as best-effort. If the language is not known in advance, the batch endpoint detects over the whole file and is more reliable.

Connection lifecycle

Step 1: Open the connection & send init

The first message after connecting is a JSON config that authenticates and sets audio parameters. You can't change these mid-stream.

{
  "api_key": "$ACCESS_TOKEN",
  "model": "zero-indic",
  "language": "hi",
  "sample_rate": 8000,
  "boost_phrases": "Empagliflozin||NACH mandate||Bengaluru",
  "boost_weight": 30
}

Init fields

FieldTypeDefaultDescription
api_keystringrequiredYour short-lived access token (Bearer) — not the raw API key
language, sample_ratestringautoCustom session ID for tracking
modelstringzero-indicASR model
languagestringautoISO code or full name
sample_rateint16000Audio sample rate (Hz)
boost_phrasesstring—Terms to bias recognition toward, separated by || or newlines. Sent once, applies for the whole connection. Ten terms maximum — more are dropped, and the ready frame echoes the count actually in force, so check it if you are unsure.
boost_weightfloat10How strongly to bias, for Indian-language audio; no effect on English or auto. Leave it out — the default is also the maximum, and higher values are clamped because above it the recogniser begins inserting terms that were never spoken.
diarizeboolfalseLabel each final with a speaker. Pair with num_speakers when you know the count.
endpoint_silence_msint3000Trailing silence before a segment is finalised. This is pure waiting: it is wall-clock delay between the speaker stopping and a transcript you can act on.Voice agents should lower it — 400–700 ms is a good range, and the default of 3000 adds three seconds to every turn. Below about 300 ms a speaker who pauses mid-sentence (“a table for… four”) gets cut in two, and your agent answers half a sentence. The ready frame echoes the value in force.
decode_every_msint1280How often interim results are produced. Lower gives more responsive partial events at the cost of GPU work per stream; 400–640 ms suits an agent that shows live captions or does speculative work on partials. It does not affect how quickly a final arrives — that is endpoint_silence_ms.
suppress_empty_finalsbooltrueDrop final events whose text is empty, including those produced by commit and end of stream. On by default: a silent line re-endpoints roughly every 800 ms, so on a real call most finals carry no text, and an agent that treats every final as a completed utterance will interrupt itself. Turn-end is unaffected — utterance_end still fires for every endpoint, suppressed or not. Set it to false only if your client waits for a final as its acknowledgement of commit or end of stream; prefer keying that off utterance_end instead.

Step 2: Stream audio frames

Send raw binary frames of 16-bit signed PCM, little-endian, mono at the sample_rate you declared. That is the only encoding this endpoint reads — there is no format negotiation, so audio in any other encoding is decoded as 16-bit PCM and comes back as noise.

If your source is 8 kHz telephony (G.711 µ-law or A-law), decode it to 16-bit PCM your side and declare sample_rate: 8000. Both 8 kHz and 16 kHz are supported; the service resamples internally.

Always declare sample_rate
It is read from the init frame only. A WAV header on your first frame is not parsed — it is simply short enough to be harmless — so if the declared rate does not match your audio the transcript comes back distorted rather than failing outright.
Frame size and pacing
Roughly 20 ms to 200 ms per frame works well — smaller frames give smoother partials, larger ones cut overhead. Pace them to real time: a live call arrives no faster than it is spoken, and pushing a whole file at once measures your upload speed rather than the recogniser.

Step 3: Handle server events

Event table

EventWhenKey fields
readySession open, waiting for audio. Echoes your settings as understood.language, sample_rate, diarize, codeswitch
partialInterim transcription for the current segment. Provisional — may change. Cadence is not guaranteed: partials are skipped while a final is being produced, so a short turn may receive none at all. Use them for display only — never to drive turn-taking.seg, delta (new text since the last partial), text (segment so far), elapsed_ms
finalSegment closed; its text will not change. One turn can emit more than one — long utterances are split mid-stream. Rather than waiting for the server to go quiet, check end_of_utterance: true marks the final that ends the turn, false marks a mid-utterance fragment to concatenate.seg, text, elapsed_ms, end_of_utterance, speaker (only when diarize was set)
utterance_endAn endpoint was reached. It follows the final that carries end_of_utterance: true — but it is not once per turn, and it is not a reliable trigger for your response. Once a speaker goes quiet the gateway keeps endpointing: measured on a live session, six seconds of silence produced fourteen utterance_end events, roughly one every 430 ms, in every language. An agent that responds to this event talks to an empty line.

Trigger your response on a final that carries text instead. Empty finals are the gateway re-endpointing silence; they are suppressed by default (see suppress_empty_finals) and should be ignored if you turn them back on. Ordering between final and utterance_end is also not guaranteed — they can arrive about a millisecond apart in either order — so a handler that waits for utterance_end before reading the transcript can act before it has one.
seg
errorError occurredmessage, code

Finishing a turn

Two signals, with different effects. For a voice agent you almost always want commit.

SignalFormatEffect
Commit"commit", "flush", or {"type": "commit"}Finalizes the current utterance and keeps the socket open for the next turn.
End"end", "END", "END_OF_AUDIO", or {"type": "end"}Finalizes and closes the connection. Send once, at the end of the session.

Re-use one socket per conversation, not per utterance. Opening a new connection each turn costs a TCP + TLS + upgrade handshake — three network round trips before any audio moves, which on a long-haul connection can exceed the recognition itself. Use commit to close a turn and keep the socket open.

An empty binary frame does not finalize — it is treated as audio and the client will wait forever. Send one of the signals above.

Error codes

CodeDescription
AUTH_FAILEDInvalid API key (WS close 4001)
TIMEOUTInactivity or max connection duration exceeded
CAPACITY_FULLServer at max concurrent sessions
PROTOCOL_ERRORInvalid init config or message format
INTERNAL_ERRORUnexpected server error

Full example

node
import WebSocket from "ws";
const ws = new WebSocket("wss://asrv2prod.shunyalabs.ai/v1/realtime");

ws.on("open", () => {
  ws.send(JSON.stringify({
    api_key: process.env.ACCESS_TOKEN,   // your minted access token, not the raw API key
    language: "hi",
    sample_rate: 8000,
    model: "zero-indic",
  }));
});

ws.on("message", (raw) => {
  const msg = JSON.parse(raw.toString());
  switch (msg.type) {
    case "ready":   console.log("Session open:", msg.language); break;
    case "partial": console.log("Partial:", msg.text); break;     // msg.seg, msg.delta
    case "final":
      // Respond HERE, and only when the final carries text.
      if (msg.text) respond(msg.text);                            // msg.end_of_utterance
      break;
    case "utterance_end":                                         // endpoint reached — NOT a turn
      break;                                                      // fires repeatedly during silence
    case "final_refined":                                         // only with codeswitch: true
      console.log("Refined:", msg.seg, msg.text); break;          // replaces that seg's final
    case "error":   console.error("Error:", msg.error); break;
  }
});

// Stream audio from a source (e.g. telephony media, mic, file)
audioSource.on("data", (chunk) => ws.send(chunk));
audioSource.on("end",  () => ws.send("commit"));   // "end" instead closes the socket
python
import asyncio, json, os, websockets

async def transcribe_stream(audio_iter):
    async with websockets.connect("wss://asrv2prod.shunyalabs.ai/v1/realtime") as ws:
        await ws.send(json.dumps({
            "api_key": os.environ["ACCESS_TOKEN"],  # your minted access token, not the raw API key
            "model": "zero-indic",
            "language": "hi",
            "sample_rate": 16000,
        }))

        async def sender():
            async for frame in audio_iter:
                await ws.send(frame)
            await ws.send("commit")   # finalize this turn, keep the socket open

        async def receiver():
            async for raw in ws:
                msg = json.loads(raw)
                if msg["type"] == "partial":
                    print(" partial:", msg["text"])
                elif msg["type"] == "final":
                    # Respond HERE, and only when the final carries text.
                    if msg["text"]:
                        respond(msg["text"])          # msg["end_of_utterance"]
                elif msg["type"] == "utterance_end":  # endpoint reached -- NOT a turn
                    pass                              # fires repeatedly during silence
                elif msg["type"] == "final_refined":  # only with codeswitch: true
                    print(" refined:", msg["seg"], msg["text"])

        await asyncio.gather(sender(), receiver())
ASR streaming | Shunya Labs Docs