ASR API reference
Every endpoint under asrv2prod.shunyalabs.ai, their request fields, response shapes, and error codes.
The examples below send Authorization: Bearer $ACCESS_TOKEN. The speech APIs accept only a short-lived access token — never your API key directly.
From the console (recommended). Open the console, click Generate token next to your API key, and copy it — then set it:
export ACCESS_TOKEN="eyJhbGciOiJSUzI1NiIs…paste-here"Or mint it from your API key — the path for production, where your app refreshes the token as it nears expiry (the response carries expires_in):
export ACCESS_TOKEN=$(curl -s -X POST https://app.shunyalabs.ai/api/auth/token \
-H "api-key: $SHUNYALABS_API_KEY" | jq -r .token)1. Authentication
All endpoints except /health and /languages require a bearer token in the Authorization header.
Authorization: Bearer $ACCESS_TOKEN2. POST /v1/audio/transcriptions
Batch transcription of an audio file or URL. Returns the full transcript, per-segment timestamps, and any intelligence results you enabled.
Required: the audio, as file, audio_base64 or url. Content-Type: multipart/form-data. See Configuration for every parameter.
Request:
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-F "file=@call.wav" \
-F "language_code=hi" \
-F "response_format=verbose_json"Request parameters (all optional except the audio):
| Field | Type | Default | Description |
|---|---|---|---|
file / audio_base64 / url | file / string | — | The audio to transcribe — a multipart file upload, base64 audio, or a public https URL fetched server-side. Exactly one is required. wav, mp3, flac, ogg, m4a, mp4 or webm, any sample rate, mono or stereo; converted to 16 kHz mono for you. Max 500 MB. |
model | string | — | Accepted for compatibility with OpenAI-shaped clients. Recognition is selected by language_code together with the feature flags below. See Models. |
language_code | string | auto | ISO code (hi, en) or full English name (Hindi); language is accepted as an alias. auto detects the language from the audio, but naming it is roughly 10× faster because detection is skipped — send an explicit code in production. An unsupported code returns 400 listing what is supported; GET /languages returns the same list. |
response_format | string | json | Defaults to json, which returns { text } and nothing else. Ask for verbose_json to get segments, word-level timestamps, confidence, nbest, and any features you enabled. |
num_speakers | integer | — | Expected number of speakers. Setting it is what turns diarization on: segments come back labelled (SPEAKER_00…) alongside speakers and speaker_turns. Leave it blank for no diarization. |
output_script | string | — | Transliterate the transcript into another script, e.g. Latin. |
boost_phrases | string | — | Domain terms to bias recognition toward, separated by || or newlines (names, products, acronyms). Applied during decoding. Works on the batch endpoint and on the real-time connection (send it in the opening frame, where at most ten terms are used). |
boost_weight | float | 10 | How strongly to bias toward boost_phrases, for Indian-language audio. No effect on English or auto — the terms still apply, the strength is ignored. High values distort ordinary words near a boosted one. |
nlp_analysis | boolean | false | Add the whole nlp_analysis object — intent, sentiment, summary and keyterms — in one pass. Use the enable_* fields below to ask for only some of them: they cost the same single pass and return sooner. |
enable_intent_detection | boolean | false | Detect the call intent (optional intent_choices, a JSON array, constrains the labels and is worth sending whenever you branch on the value). Returned under nlp_analysis.intent. |
enable_sentiment_analysis | boolean | false | Overall sentiment, under nlp_analysis.sentiment. |
enable_summarization | boolean | false | Summary of the transcript (summary_max_length caps the length), under nlp_analysis.summary. |
enable_keyterm_normalization | boolean | false | Key terms plus a normalized_text rewrite of the transcript (spelled-out acronyms written properly, standard capitalisation), under nlp_analysis. Optional keyterm_keywords glossary — a JSON array of your own brand or place names — replaces the recogniser’s guess with your spelling. |
emotion | boolean | false | Add speech emotion: a whole-clip emotion ({ label, score }), per-segment emotion and emotion_confidence, and an emotion_summary. Segments are returned even for response_format=json. |
speaker_id | boolean | false | Match diarized speakers against registered voiceprints and add speakers_identified. Send it with num_speakers (at least 2). See Speaker APIs. |
medical | boolean | false | Correct medical terminology in the transcript. This flag is the switch; no model name is required. |
profanity | boolean | false | Mask profane words in the transcript (s***). |
hash_keywords | string | — | JSON array of your own terms to redact, each replaced with ****. Matching ignores case and tolerates a word the recogniser split, but respects word boundaries, so cat does not redact catalogue. Separate from profanity. |
codeswitch | boolean | false | A post-processing pass over the finished transcript, not a separate recogniser: English words and sentences come back in Latin script and spelled correctly (केवाईसी → KYC, ई एम आई → EMI), brand and place names are fixed, and spoken numbers become digits (तीन सौ दस रुपये → 310 रुपये). Adds a few seconds to the response. |
Response (verbose_json):
{
"success": true,
"request_id": "b3f1a2c4...",
"text": "नमस्ते मोहम्मद जी, ये एक ज़रूरी कॉल है।",
"detected_language": "hi",
"detected_language_name": "Hindi",
"segments": [
{
"start": 0.51, "end": 5.70,
"text": "नमस्ते मोहम्मद जी...",
"speaker": "SPEAKER_00",
"confidence": 0.97,
"emotion": "neu",
"emotion_confidence": 0.68
}
],
"words": [
{ "word": "नमस्ते", "start": 0.51, "end": 0.94, "confidence": 0.99 }
],
"speakers": ["SPEAKER_00"],
"speaker_turns": [
{ "start": 0.51, "end": 5.70, "speaker": "SPEAKER_00" }
],
"audio_duration": 5.7,
"inference_time_ms": 812.3,
"emotion": { "label": "neu", "score": 0.68 },
"emotion_summary": {
"dominant_emotion": "neu",
"emotion_distribution": { "neu": 100.0 },
"avg_confidence": 0.68
},
"nlp_analysis": {
"intent": "Booking confirmation request",
"sentiment": "neutral",
"summary": "Caller greets Mohammad and flags an urgent call.",
"keyterms": ["ज़रूरी कॉल"]
}
}detected_language is an ISO code; detected_language_name is the same value spelled out. The nlp_analysis, emotion/emotion_summary, and speakers/speaker_turns blocks appear only when you enable the matching parameter, and per-segment emotion/emotion_confidenceappear only with emotion=true.
Response (json: minimal, OpenAI-compatible):
{ "text": "नमस्ते मोहम्मद जी, ये एक ज़रूरी कॉल है।" }3. WebSocket /v1/realtime
Full protocol documented in Streaming. Summary of the lifecycle:
- Open
wss://asrv2prod.shunyalabs.ai/v1/realtime - Send a JSON init frame carrying your access token as
token(orapi_key), pluslanguage_code(aliaslanguage) andsample_rate. Also read:model,diarize,num_speakers,codeswitch,boost_phrases,boost_weight. Nothing else in the frame is read — fields likedtypeorchunk_size_secare ignored. - Wait for
ready, then stream binary frames of 16-bit signed little-endian mono PCM at thesample_rateyou declared. There is no format negotiation: G.711 µ-law or A-law sent raw is read as int16 and comes back as noise, so telephony callers decode to PCM their side and declaresample_rate8000. - Send the text frame
"commit"to finalize a turn and keep the socket open, or"end"to finalize and close. Keep reading until the lastfinalarrives. - Events are
ready,partial,finalanderror— plusfinal_refinedwhencodeswitchis on. A single turn can emit more than onefinal, so concatenate them rather than assuming one per commit.
Errors are an error event — {"type":"error","error":"..."} — and close code 1008. The reason is the error string; there is no separate error-code vocabulary and no code or message field to switch on.
4. GET /health
Unauthenticated. Use for deployment smoke tests.
Request:
curl https://asrv2prod.shunyalabs.ai/healthResponse:
{ "ok": true }Returns ok: true when the service is healthy. No token required — use it for liveness and deployment smoke tests.
5. GET /languages
Returns the supported language list — 145 codes — alongside the model aliases the service accepts and its feature names. Unauthenticated, like /health.
Request:
curl https://asrv2prod.shunyalabs.ai/languages6. Speaker APIs
Diarization produces SPEAKER_00-style labels. To map those to actual names, register voice profiles using these two endpoints, then send speaker_id=true with num_speakers.
6.1 POST /v1/speakers/register
Request:
curl -X POST https://asrv2prod.shunyalabs.ai/v1/speakers/register \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-F "name=Priya" \
-F "file=@priya_sample.wav" \
-F "project=support_team"Response:
{ "success": true, "speaker": "Priya", "message": "Registered successfully" }6.2 DELETE /v1/speakers/delete
Request:
curl -X DELETE https://asrv2prod.shunyalabs.ai/v1/speakers/delete \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-F "name=Priya" \
-F "project=support_team"Response:
{ "success": true }7. HTTP error codes
| Status | Meaning |
|---|---|
| 200 | Success, audio or JSON body returned. |
| 400 | Bad request, missing or malformed fields. Response body: {"detail": "..."}. |
| 401 | Unauthorized — access token invalid, expired, or missing. Mint a new one and retry. |
| 413 | Payload too large — the audio is over 500 MB. |
| 415 | Unsupported media — the audio could not be decoded. |
| 422 | Validation error — a field is present but the wrong type or value. Response body: {"detail": [{"loc": ..., "msg": ..., "type": ...}]}. |
| 429 | Rate limit exceeded. Back off and retry. |
| 500 | Internal server error, unexpected server-side failure. |
| 503 | Service unavailable, an internal dependency is temporarily down. |
| 504 | Gateway timeout, request exceeded the processing window. |
8. Retry patterns
Safe to retry: 429, 500, 502, 503, 504. Not safe: 400, 401, 422: fix the request first.
Exponential backoff (Python)
import time, requests
def transcribe_with_retry(file_path, retries=3):
for attempt in range(retries):
with open(file_path, "rb") as f:
r = requests.post(
"https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions",
headers={"Authorization": f"Bearer {ACCESS_TOKEN}"},
files={"file": f},
data={"language_code": "hi"},
timeout=120,
)
if r.status_code == 200:
return r.json()
if r.status_code in (429, 500, 502, 503, 504):
time.sleep(2 ** attempt) # 1s, 2s, 4s
continue
r.raise_for_status()
raise RuntimeError("Max retries exceeded")WebSocket reconnection
async def stream_with_reconnect(max_attempts=3):
for attempt in range(max_attempts):
try:
async with websockets.connect("wss://asrv2prod.shunyalabs.ai/v1/realtime") as ws:
await ws.send(json.dumps({...}))
async for msg in ws:
yield json.loads(msg)
return
except websockets.ConnectionClosed:
await asyncio.sleep(2 ** attempt)
raise RuntimeError("Max reconnects exceeded")9. Rate limits
| Limit | Value |
|---|---|
| Max file size | 500 MB — larger audio returns 413 |
| Max audio duration per file | 4 hours |
| Concurrent requests (default tier) | 16 |
| HTTP request timeout | Use at least 120 s for long audio |
| WebSocket inactivity timeout | Server-side, not client-configurable — inactivity_timeout and max_connection_duration in the init frame are ignored |
10. Request IDs
Successful responses carry an X-Request-Id header, and a verbose_json body repeats it as request_id (a minimal json body is only { text }, so read the header there). Log it, Shunya support uses it to trace issues.