Shunya Labs DocsShunya Labs Docs
🌐 International
🇺🇸 English
🇯🇵 Japanese
🇨🇳 Chinese (Simplified)
🇹🇼 Chinese (Traditional)
🇸🇦 Arabic
🇩🇪 German
🇫🇷 French
🇪🇸 Spanish
🇧🇷 Portuguese
🇷🇺 Russian
🇰🇷 Korean
🇹🇷 Turkish
🇻🇳 Vietnamese
🇮🇩 Indonesian
🇮🇳 Hindi Belt
हिन्दी — Hindi
भोजपुरी — Bhojpuri
मैथिली — Maithili
राजस्थानी — Rajasthani
🇮🇳 South India
தமிழ் — Tamil
తెలుగు — Telugu
ಕನ್ನಡ — Kannada
മലയാളം — Malayalam
🇮🇳 West India
मराठी — Marathi
ગુજરાતી — Gujarati
कोंकणी — Konkani
🇮🇳 East India
বাংলা — Bengali
ଓଡ଼ିଆ — Odia
অসমীয়া — Assamese
🇮🇳 North-East India
মেইতেই — Meitei
नेपाली — Nepali
🇮🇳 North India
ਪੰਜਾਬੀ — Punjabi
اردو — Urdu
کٲشُر — Kashmiri
डोगरी — Dogri
سنڌي — Sindhi

ASR API reference

Every endpoint under asrv2prod.shunyalabs.ai, their request fields, response shapes, and error codes.

First: get an access token

The examples below send Authorization: Bearer $ACCESS_TOKEN. The speech APIs accept only a short-lived access token — never your API key directly.

From the console (recommended). Open the console, click Generate token next to your API key, and copy it — then set it:

export ACCESS_TOKEN="eyJhbGciOiJSUzI1NiIs…paste-here"

Or mint it from your API key — the path for production, where your app refreshes the token as it nears expiry (the response carries expires_in):

export ACCESS_TOKEN=$(curl -s -X POST https://app.shunyalabs.ai/api/auth/token \
  -H "api-key: $SHUNYALABS_API_KEY" | jq -r .token)

1. Authentication

All endpoints except /health and /languages require a bearer token in the Authorization header.

Authorization: Bearer $ACCESS_TOKEN

2. POST /v1/audio/transcriptions

Batch transcription of an audio file or URL. Returns the full transcript, per-segment timestamps, and any intelligence results you enabled.

Required: the audio, as file, audio_base64 or url. Content-Type: multipart/form-data. See Configuration for every parameter.

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@call.wav" \
  -F "language_code=hi" \
  -F "response_format=verbose_json"

Request parameters (all optional except the audio):

FieldTypeDefaultDescription
file / audio_base64 / urlfile / string—The audio to transcribe — a multipart file upload, base64 audio, or a public https URL fetched server-side. Exactly one is required. wav, mp3, flac, ogg, m4a, mp4 or webm, any sample rate, mono or stereo; converted to 16 kHz mono for you. Max 500 MB.
modelstring—Accepted for compatibility with OpenAI-shaped clients. Recognition is selected by language_code together with the feature flags below. See Models.
language_codestringautoISO code (hi, en) or full English name (Hindi); language is accepted as an alias. auto detects the language from the audio, but naming it is roughly 10× faster because detection is skipped — send an explicit code in production. An unsupported code returns 400 listing what is supported; GET /languages returns the same list.
response_formatstringjsonDefaults to json, which returns { text } and nothing else. Ask for verbose_json to get segments, word-level timestamps, confidence, nbest, and any features you enabled.
num_speakersinteger—Expected number of speakers. Setting it is what turns diarization on: segments come back labelled (SPEAKER_00…) alongside speakers and speaker_turns. Leave it blank for no diarization.
output_scriptstring—Transliterate the transcript into another script, e.g. Latin.
boost_phrasesstring—Domain terms to bias recognition toward, separated by || or newlines (names, products, acronyms). Applied during decoding. Works on the batch endpoint and on the real-time connection (send it in the opening frame, where at most ten terms are used).
boost_weightfloat10How strongly to bias toward boost_phrases, for Indian-language audio. No effect on English or auto — the terms still apply, the strength is ignored. High values distort ordinary words near a boosted one.
nlp_analysisbooleanfalseAdd the whole nlp_analysis object — intent, sentiment, summary and keyterms — in one pass. Use the enable_* fields below to ask for only some of them: they cost the same single pass and return sooner.
enable_intent_detectionbooleanfalseDetect the call intent (optional intent_choices, a JSON array, constrains the labels and is worth sending whenever you branch on the value). Returned under nlp_analysis.intent.
enable_sentiment_analysisbooleanfalseOverall sentiment, under nlp_analysis.sentiment.
enable_summarizationbooleanfalseSummary of the transcript (summary_max_length caps the length), under nlp_analysis.summary.
enable_keyterm_normalizationbooleanfalseKey terms plus a normalized_text rewrite of the transcript (spelled-out acronyms written properly, standard capitalisation), under nlp_analysis. Optional keyterm_keywords glossary — a JSON array of your own brand or place names — replaces the recogniser’s guess with your spelling.
emotionbooleanfalseAdd speech emotion: a whole-clip emotion ({ label, score }), per-segment emotion and emotion_confidence, and an emotion_summary. Segments are returned even for response_format=json.
speaker_idbooleanfalseMatch diarized speakers against registered voiceprints and add speakers_identified. Send it with num_speakers (at least 2). See Speaker APIs.
medicalbooleanfalseCorrect medical terminology in the transcript. This flag is the switch; no model name is required.
profanitybooleanfalseMask profane words in the transcript (s***).
hash_keywordsstring—JSON array of your own terms to redact, each replaced with ****. Matching ignores case and tolerates a word the recogniser split, but respects word boundaries, so cat does not redact catalogue. Separate from profanity.
codeswitchbooleanfalseA post-processing pass over the finished transcript, not a separate recogniser: English words and sentences come back in Latin script and spelled correctly (केवाईसी → KYC, ई एम आई → EMI), brand and place names are fixed, and spoken numbers become digits (तीन सौ दस रुपये → 310 रुपये). Adds a few seconds to the response.

Response (verbose_json):

{
  "success": true,
  "request_id": "b3f1a2c4...",
  "text": "नमस्ते मोहम्मद जी, ये एक ज़रूरी कॉल है।",
  "detected_language": "hi",
  "detected_language_name": "Hindi",
  "segments": [
    {
      "start": 0.51, "end": 5.70,
      "text": "नमस्ते मोहम्मद जी...",
      "speaker": "SPEAKER_00",
      "confidence": 0.97,
      "emotion": "neu",
      "emotion_confidence": 0.68
    }
  ],
  "words": [
    { "word": "नमस्ते", "start": 0.51, "end": 0.94, "confidence": 0.99 }
  ],
  "speakers": ["SPEAKER_00"],
  "speaker_turns": [
    { "start": 0.51, "end": 5.70, "speaker": "SPEAKER_00" }
  ],
  "audio_duration": 5.7,
  "inference_time_ms": 812.3,
  "emotion": { "label": "neu", "score": 0.68 },
  "emotion_summary": {
    "dominant_emotion": "neu",
    "emotion_distribution": { "neu": 100.0 },
    "avg_confidence": 0.68
  },
  "nlp_analysis": {
    "intent": "Booking confirmation request",
    "sentiment": "neutral",
    "summary": "Caller greets Mohammad and flags an urgent call.",
    "keyterms": ["ज़रूरी कॉल"]
  }
}

detected_language is an ISO code; detected_language_name is the same value spelled out. The nlp_analysis, emotion/emotion_summary, and speakers/speaker_turns blocks appear only when you enable the matching parameter, and per-segment emotion/emotion_confidenceappear only with emotion=true.

Response (json: minimal, OpenAI-compatible):

{ "text": "नमस्ते मोहम्मद जी, ये एक ज़रूरी कॉल है।" }

3. WebSocket /v1/realtime

Full protocol documented in Streaming. Summary of the lifecycle:

  1. Open wss://asrv2prod.shunyalabs.ai/v1/realtime
  2. Send a JSON init frame carrying your access token as token (or api_key), plus language_code (alias language) and sample_rate. Also read: model, diarize, num_speakers, codeswitch, boost_phrases, boost_weight. Nothing else in the frame is read — fields like dtype or chunk_size_sec are ignored.
  3. Wait for ready, then stream binary frames of 16-bit signed little-endian mono PCM at the sample_rate you declared. There is no format negotiation: G.711 µ-law or A-law sent raw is read as int16 and comes back as noise, so telephony callers decode to PCM their side and declare sample_rate 8000.
  4. Send the text frame "commit" to finalize a turn and keep the socket open, or "end" to finalize and close. Keep reading until the last final arrives.
  5. Events are ready, partial, final and error — plus final_refined when codeswitch is on. A single turn can emit more than one final, so concatenate them rather than assuming one per commit.

Errors are an error event — {"type":"error","error":"..."} — and close code 1008. The reason is the error string; there is no separate error-code vocabulary and no code or message field to switch on.

4. GET /health

Unauthenticated. Use for deployment smoke tests.

Request:

curl https://asrv2prod.shunyalabs.ai/health

Response:

{ "ok": true }

Returns ok: true when the service is healthy. No token required — use it for liveness and deployment smoke tests.

5. GET /languages

Returns the supported language list — 145 codes — alongside the model aliases the service accepts and its feature names. Unauthenticated, like /health.

Request:

curl https://asrv2prod.shunyalabs.ai/languages

6. Speaker APIs

Diarization produces SPEAKER_00-style labels. To map those to actual names, register voice profiles using these two endpoints, then send speaker_id=true with num_speakers.

Reference clip requirements
5-15 seconds, speaker alone, no background music, no overlapping voices, 16 kHz or higher sample rate.

6.1 POST /v1/speakers/register

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/speakers/register \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "name=Priya" \
  -F "file=@priya_sample.wav" \
  -F "project=support_team"

Response:

{ "success": true, "speaker": "Priya", "message": "Registered successfully" }

6.2 DELETE /v1/speakers/delete

Request:

curl -X DELETE https://asrv2prod.shunyalabs.ai/v1/speakers/delete \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "name=Priya" \
  -F "project=support_team"

Response:

{ "success": true }

7. HTTP error codes

StatusMeaning
200Success, audio or JSON body returned.
400Bad request, missing or malformed fields. Response body: {"detail": "..."}.
401Unauthorized — access token invalid, expired, or missing. Mint a new one and retry.
413Payload too large — the audio is over 500 MB.
415Unsupported media — the audio could not be decoded.
422Validation error — a field is present but the wrong type or value. Response body: {"detail": [{"loc": ..., "msg": ..., "type": ...}]}.
429Rate limit exceeded. Back off and retry.
500Internal server error, unexpected server-side failure.
503Service unavailable, an internal dependency is temporarily down.
504Gateway timeout, request exceeded the processing window.

8. Retry patterns

Safe to retry: 429, 500, 502, 503, 504. Not safe: 400, 401, 422: fix the request first.

Exponential backoff (Python)

import time, requests

def transcribe_with_retry(file_path, retries=3):
    for attempt in range(retries):
        with open(file_path, "rb") as f:
            r = requests.post(
                "https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions",
                headers={"Authorization": f"Bearer {ACCESS_TOKEN}"},
                files={"file": f},
                data={"language_code": "hi"},
                timeout=120,
            )
        if r.status_code == 200:
            return r.json()
        if r.status_code in (429, 500, 502, 503, 504):
            time.sleep(2 ** attempt)  # 1s, 2s, 4s
            continue
        r.raise_for_status()
    raise RuntimeError("Max retries exceeded")

WebSocket reconnection

async def stream_with_reconnect(max_attempts=3):
    for attempt in range(max_attempts):
        try:
            async with websockets.connect("wss://asrv2prod.shunyalabs.ai/v1/realtime") as ws:
                await ws.send(json.dumps({...}))
                async for msg in ws:
                    yield json.loads(msg)
                return
        except websockets.ConnectionClosed:
            await asyncio.sleep(2 ** attempt)
    raise RuntimeError("Max reconnects exceeded")

9. Rate limits

LimitValue
Max file size500 MB — larger audio returns 413
Max audio duration per file4 hours
Concurrent requests (default tier)16
HTTP request timeoutUse at least 120 s for long audio
WebSocket inactivity timeoutServer-side, not client-configurable — inactivity_timeout and max_connection_duration in the init frame are ignored

10. Request IDs

Successful responses carry an X-Request-Id header, and a verbose_json body repeats it as request_id (a minimal json body is only { text }, so read the header there). Log it, Shunya support uses it to trace issues.

ASR API reference | Shunya Labs Docs