Streaming ASR over WebSocket
For voice agents, IVR, and live captioning. You open a WebSocket, send audio frames as they arrive, and receive partial and final transcripts in real time.
The examples below send Authorization: Bearer $ACCESS_TOKEN. The speech APIs accept only a short-lived access token — never your API key directly.
From the console (recommended). Open the console, click Generate token next to your API key, and copy it — then set it:
export ACCESS_TOKEN="eyJhbGciOiJSUzI1NiIs…paste-here"Or mint it from your API key — the path for production, where your app refreshes the token as it nears expiry (the response carries expires_in):
export ACCESS_TOKEN=$(curl -s -X POST https://app.shunyalabs.ai/api/auth/token \
-H "api-key: $SHUNYALABS_API_KEY" | jq -r .token)Endpoint
wss://asrv2prod.shunyalabs.ai/v1/realtime- There are four event types, plus one you only see if you ask for it:
ready,partial,final,error— andfinal_refinedwhen you setcodeswitchin the opening frame. Build the transcript fromfinal;partialis provisional and will be revised. Do not wait for a separate completion event before using the text. - Result events carry
segandelapsed_ms.segincrements per segment, so it is what you order and de-duplicate on. finalandpartialcarryconfidence;finalalso carriesaudio_duration_sec.confidenceis the mean per-token confidence for that decode, 0–1, and it tracks quality closely: on a real call a spurious one-word fragment scored0.055and a repetition loop0.412, against0.74–0.94for correct finals.audio_duration_secis the seconds of audio that segment decoded. Use them together — a short transcript with low confidence from a long segment is the pattern worth rejecting, and you choose the threshold. Both are omitted rather than faked when a recogniser returns no per-token data, so an absentconfidencemeans “unavailable”, never “low”.final_refinedcarries no confidence: it is rewritten text, so a decoder score would not describe it.- One turn can emit several finals. A long utterance is split mid-stream, so a single turn may produce more than one
final. Read until the server goes quiet and concatenate insegorder — code that assumes exactly one final per turn drifts a turn further behind on every turn. final_refinedarrives after thefinal, not instead of it. Withcodeswitchon, a segment is delivered twice: the rawfinalimmediately, then a refined version of the samesega moment later. Nothing is held back waiting for it, so keep rendering thefinaland replace thatsegif the refinement turns up. Match onseg, not on arrival order.- Send
commitbetween turns, andendonce at the close.commitfinalizes the current turn and keeps the socket open;endcloses it, so usingendper utterance forces a fresh handshake on every turn and adds the round trips back to your latency. - Send an explicit
language. Omitting it (or sendingauto) is accepted and will attempt detection, but a live stream has to decide from the opening seconds of audio, so treat it as best-effort. If the language is not known in advance, the batch endpoint detects over the whole file and is more reliable.
Connection lifecycle
Step 1: Open the connection & send init
The first message after connecting is a JSON config that authenticates and sets audio parameters. You can't change these mid-stream.
{
"api_key": "$ACCESS_TOKEN",
"model": "zero-indic",
"language": "hi",
"sample_rate": 8000,
"boost_phrases": "Empagliflozin||NACH mandate||Bengaluru",
"boost_weight": 30
}Init fields
| Field | Type | Default | Description |
|---|---|---|---|
api_key | string | required | Your short-lived access token (Bearer) — not the raw API key |
language, sample_rate | string | auto | Custom session ID for tracking |
model | string | zero-indic | ASR model |
language | string | auto | ISO code or full name |
sample_rate | int | 16000 | Audio sample rate (Hz) |
boost_phrases | string | — | Terms to bias recognition toward, separated by || or newlines. Sent once, applies for the whole connection. Ten terms maximum — more are dropped, and the ready frame echoes the count actually in force, so check it if you are unsure. |
boost_weight | float | 10 | How strongly to bias, for Indian-language audio; no effect on English or auto. Leave it out — the default is also the maximum, and higher values are clamped because above it the recogniser begins inserting terms that were never spoken. |
diarize | bool | false | Label each final with a speaker. Pair with num_speakers when you know the count. |
endpoint_silence_ms | int | 3000 | Trailing silence before a segment is finalised. This is pure waiting: it is wall-clock delay between the speaker stopping and a transcript you can act on.Voice agents should lower it — 400–700 ms is a good range, and the default of 3000 adds three seconds to every turn. Below about 300 ms a speaker who pauses mid-sentence (“a table for… four”) gets cut in two, and your agent answers half a sentence. The ready frame echoes the value in force. |
decode_every_ms | int | 1280 | How often interim results are produced. Lower gives more responsive partial events at the cost of GPU work per stream; 400–640 ms suits an agent that shows live captions or does speculative work on partials. It does not affect how quickly a final arrives — that is endpoint_silence_ms. |
suppress_empty_finals | bool | true | Drop final events whose text is empty, including those produced by commit and end of stream. On by default: a silent line re-endpoints roughly every 800 ms, so on a real call most finals carry no text, and an agent that treats every final as a completed utterance will interrupt itself. Turn-end is unaffected — utterance_end still fires for every endpoint, suppressed or not. Set it to false only if your client waits for a final as its acknowledgement of commit or end of stream; prefer keying that off utterance_end instead. |
Step 2: Stream audio frames
Send raw binary frames of 16-bit signed PCM, little-endian, mono at the sample_rate you declared. That is the only encoding this endpoint reads — there is no format negotiation, so audio in any other encoding is decoded as 16-bit PCM and comes back as noise.
If your source is 8 kHz telephony (G.711 µ-law or A-law), decode it to 16-bit PCM your side and declare sample_rate: 8000. Both 8 kHz and 16 kHz are supported; the service resamples internally.
sample_rateStep 3: Handle server events
Event table
| Event | When | Key fields |
|---|---|---|
ready | Session open, waiting for audio. Echoes your settings as understood. | language, sample_rate, diarize, codeswitch |
partial | Interim transcription for the current segment. Provisional — may change. Cadence is not guaranteed: partials are skipped while a final is being produced, so a short turn may receive none at all. Use them for display only — never to drive turn-taking. | seg, delta (new text since the last partial), text (segment so far), elapsed_ms |
final | Segment closed; its text will not change. One turn can emit more than one — long utterances are split mid-stream. Rather than waiting for the server to go quiet, check end_of_utterance: true marks the final that ends the turn, false marks a mid-utterance fragment to concatenate. | seg, text, elapsed_ms, end_of_utterance, speaker (only when diarize was set) |
utterance_end | An endpoint was reached. It follows the final that carries end_of_utterance: true — but it is not once per turn, and it is not a reliable trigger for your response. Once a speaker goes quiet the gateway keeps endpointing: measured on a live session, six seconds of silence produced fourteen utterance_end events, roughly one every 430 ms, in every language. An agent that responds to this event talks to an empty line.Trigger your response on a final that carries text instead. Empty finals are the gateway re-endpointing silence; they are suppressed by default (see suppress_empty_finals) and should be ignored if you turn them back on. Ordering between final and utterance_end is also not guaranteed — they can arrive about a millisecond apart in either order — so a handler that waits for utterance_end before reading the transcript can act before it has one. | seg |
error | Error occurred | message, code |
Finishing a turn
Two signals, with different effects. For a voice agent you almost always want commit.
| Signal | Format | Effect |
|---|---|---|
| Commit | "commit", "flush", or {"type": "commit"} | Finalizes the current utterance and keeps the socket open for the next turn. |
| End | "end", "END", "END_OF_AUDIO", or {"type": "end"} | Finalizes and closes the connection. Send once, at the end of the session. |
Re-use one socket per conversation, not per utterance. Opening a new connection each turn costs a TCP + TLS + upgrade handshake — three network round trips before any audio moves, which on a long-haul connection can exceed the recognition itself. Use commit to close a turn and keep the socket open.
An empty binary frame does not finalize — it is treated as audio and the client will wait forever. Send one of the signals above.
Error codes
| Code | Description |
|---|---|
AUTH_FAILED | Invalid API key (WS close 4001) |
TIMEOUT | Inactivity or max connection duration exceeded |
CAPACITY_FULL | Server at max concurrent sessions |
PROTOCOL_ERROR | Invalid init config or message format |
INTERNAL_ERROR | Unexpected server error |
Full example
import WebSocket from "ws";
const ws = new WebSocket("wss://asrv2prod.shunyalabs.ai/v1/realtime");
ws.on("open", () => {
ws.send(JSON.stringify({
api_key: process.env.ACCESS_TOKEN, // your minted access token, not the raw API key
language: "hi",
sample_rate: 8000,
model: "zero-indic",
}));
});
ws.on("message", (raw) => {
const msg = JSON.parse(raw.toString());
switch (msg.type) {
case "ready": console.log("Session open:", msg.language); break;
case "partial": console.log("Partial:", msg.text); break; // msg.seg, msg.delta
case "final":
// Respond HERE, and only when the final carries text.
if (msg.text) respond(msg.text); // msg.end_of_utterance
break;
case "utterance_end": // endpoint reached — NOT a turn
break; // fires repeatedly during silence
case "final_refined": // only with codeswitch: true
console.log("Refined:", msg.seg, msg.text); break; // replaces that seg's final
case "error": console.error("Error:", msg.error); break;
}
});
// Stream audio from a source (e.g. telephony media, mic, file)
audioSource.on("data", (chunk) => ws.send(chunk));
audioSource.on("end", () => ws.send("commit")); // "end" instead closes the socketimport asyncio, json, os, websockets
async def transcribe_stream(audio_iter):
async with websockets.connect("wss://asrv2prod.shunyalabs.ai/v1/realtime") as ws:
await ws.send(json.dumps({
"api_key": os.environ["ACCESS_TOKEN"], # your minted access token, not the raw API key
"model": "zero-indic",
"language": "hi",
"sample_rate": 16000,
}))
async def sender():
async for frame in audio_iter:
await ws.send(frame)
await ws.send("commit") # finalize this turn, keep the socket open
async def receiver():
async for raw in ws:
msg = json.loads(raw)
if msg["type"] == "partial":
print(" partial:", msg["text"])
elif msg["type"] == "final":
# Respond HERE, and only when the final carries text.
if msg["text"]:
respond(msg["text"]) # msg["end_of_utterance"]
elif msg["type"] == "utterance_end": # endpoint reached -- NOT a turn
pass # fires repeatedly during silence
elif msg["type"] == "final_refined": # only with codeswitch: true
print(" refined:", msg["seg"], msg["text"])
await asyncio.gather(sender(), receiver())