Shunya Labs DocsShunya Labs Docs
🌐 International
🇺🇸 English
🇯🇵 Japanese
🇨🇳 Chinese (Simplified)
🇹🇼 Chinese (Traditional)
🇸🇦 Arabic
🇩🇪 German
🇫🇷 French
🇪🇸 Spanish
🇧🇷 Portuguese
🇷🇺 Russian
🇰🇷 Korean
🇹🇷 Turkish
🇻🇳 Vietnamese
🇮🇩 Indonesian
🇮🇳 Hindi Belt
हिन्दी — Hindi
भोजपुरी — Bhojpuri
मैथिली — Maithili
राजस्थानी — Rajasthani
🇮🇳 South India
தமிழ் — Tamil
తెలుగు — Telugu
ಕನ್ನಡ — Kannada
മലയാളം — Malayalam
🇮🇳 West India
मराठी — Marathi
ગુજરાતી — Gujarati
कोंकणी — Konkani
🇮🇳 East India
বাংলা — Bengali
ଓଡ଼ିଆ — Odia
অসমীয়া — Assamese
🇮🇳 North-East India
মেইতেই — Meitei
नेपाली — Nepali
🇮🇳 North India
ਪੰਜਾਬੀ — Punjabi
اردو — Urdu
کٲشُر — Kashmiri
डोगरी — Dogri
سنڌي — Sindhi

ASR features

The intelligence layer on top of Zero STT. Enable each with a boolean flag, none of them are required, and they combine freely. All of these return their results in the same JSON response as the transcript.

First: get an access token

The examples below send Authorization: Bearer $ACCESS_TOKEN. The speech APIs accept only a short-lived access token — never your API key directly.

From the console (recommended). Open the console, click Generate token next to your API key, and copy it — then set it:

export ACCESS_TOKEN="eyJhbGciOiJSUzI1NiIs…paste-here"

Or mint it from your API key — the path for production, where your app refreshes the token as it nears expiry (the response carries expires_in):

export ACCESS_TOKEN=$(curl -s -X POST https://app.shunyalabs.ai/api/auth/token \
  -H "api-key: $SHUNYALABS_API_KEY" | jq -r .token)
Easier to browse?
Open the Intelligence overview for a card-based UI with jump links and collapsible request/response examples for every feature.

1. Diarization

"Who spoke when." Adds speaker: SPEAKER_XX to every segment, a top-level speakers array, and a speaker_turns array of raw turn boundaries. The top-level text stays the plain transcript — it carries no speaker tags, so read the speaker off segments or speaker_turns.

num_speakers=N turns diarization on as well, and tells it how many voices to expect.

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@meeting.wav" \
  -F "model=zero-indic" \
  -F "enable_diarization=true"

Response:

{
  "text": "नमस्ते, आप कैसे हैं? मैं ठीक हूँ, धन्यवाद।",
  "segments": [
    { "start": 0.5, "end": 3.2, "text": "नमस्ते, आप कैसे हैं?", "speaker": "SPEAKER_00", "confidence": 0.95 },
    { "start": 4.1, "end": 6.8, "text": "मैं ठीक हूँ, धन्यवाद।", "speaker": "SPEAKER_01", "confidence": 0.98 }
  ],
  "speakers": ["SPEAKER_00", "SPEAKER_01"],
  "speaker_turns": [
    { "start": 0.5, "end": 3.2, "speaker": "SPEAKER_00" },
    { "start": 4.1, "end": 6.8, "speaker": "SPEAKER_01" }
  ]
}
  • Segments are capped at 30 seconds each to maintain transcription quality.
  • Works on any number of speakers, but best with 2-6 distinct voices.
  • Speaker labels are per-request. SPEAKER_00 in one call is not the same person as SPEAKER_00 in the next.

2. Speaker identification

Adds speakers_identified, one voiceprint summary per diarized speaker. Requires diarization. It does not rename anyone: the labels in segments stay SPEAKER_XX, and a registered name does not come back in the transcription response. Registration and the transcript are separate today — the endpoint stores a voiceprint, and transcription reports the voiceprints it found, but nothing joins the two.

Step 1: register a speaker (use a 5-15 second clip of the speaker alone, no background music):

curl -X POST https://asrv2prod.shunyalabs.ai/v1/speakers/register \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "name=Priya" \
  -F "file=@priya_sample.wav" \
  -F "project=support_team"

Step 2: transcribe with identification on:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@call.wav" \
  -F "model=zero-indic" \
  -F "enable_diarization=true" \
  -F "enable_speaker_identification=true" \
  -F "project=support_team"

Response:

{
  "speakers": ["SPEAKER_00", "SPEAKER_01"],
  "speakers_identified": [
    { "speaker": "SPEAKER_00", "turns": 1, "voiceprint_dim": 256 },
    { "speaker": "SPEAKER_01", "turns": 1, "voiceprint_dim": 256 }
  ]
}

On the register and delete endpoints, project namespaces a voice library and defaults to playground. Remove a profile with DELETE /v1/speakers/delete, passing the same name and project.

3. Emotion diarization

Detects the dominant emotion in the audio. Adds a top-level emotion object and an emotion_summary, plus emotion and emotion_confidence on every segment. Labels are short codes, not words — neu, hap, sad. Works independently of speaker diarization but is commonly used together.

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@call.wav" \
  -F "model=zero-indic" \
  -F "enable_diarization=true" \
  -F "enable_emotion_diarization=true"

Response:

{
  "emotion": { "label": "hap", "score": 0.5839 },
  "emotion_summary": {
    "dominant_emotion": "neu",
    "emotion_distribution": { "neu": 50.0, "hap": 50.0 },
    "avg_confidence": 0.4773
  },
  "segments": [
    { "start": 0.08, "end": 5.263, "text": "...", "speaker": "SPEAKER_00", "emotion": "neu", "emotion_confidence": 0.4005 },
    { "start": 5.263, "end": 18.9, "text": "...", "speaker": "SPEAKER_01", "emotion": "hap", "emotion_confidence": 0.554 }
  ]
}

4. Intent detection

Classifies the overall transcript intent. Pass intent_choices to constrain to your taxonomy — the label comes back as one of yours, copied verbatim — or leave it off for a free-form phrase such as Request NACH registration. nlp_analysis.intent is a plain string. There is no confidence or reasoning field alongside it.

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@call.wav" \
  -F "model=zero-indic" \
  -F "enable_intent_detection=true" \
  -F 'intent_choices=["complaint","inquiry","service_request","compliment"]'

Response:

{
  "nlp_analysis": {
    "intent": "service_request"
  }
}

5. Sentiment analysis

Overall sentiment of the transcript. nlp_analysis.sentiment is a plain string — positive, neutral or negative. There is no numeric score and no explanation field.

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@call.wav" \
  -F "model=zero-indic" \
  -F "enable_sentiment_analysis=true"

Response:

{
  "nlp_analysis": {
    "sentiment": "negative"
  }
}

6. Summarization

Concise summary of the transcript. summary_max_length is an approximate cap in characters, not words, and it is guidance rather than a hard truncation — the summary finishes its sentence rather than being cut off. Measured: a request for 50 came back at 69 characters, one for 150 at 163.

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@call.wav" \
  -F "model=zero-indic" \
  -F "enable_summarization=true" \
  -F "summary_max_length=150"

Response:

{
  "nlp_analysis": {
    "summary": "Customer called about a vehicle breakdown. Agent confirmed the complaint was registered and promised a technician within the hour."
  }
}

7. Keyterm normalization

Cleans up domain-specific terms the ASR model might render informally, emi → EMI, nach mandate → NACH mandate. Preserves the original language. Optionally focus on a specific glossary with keyterm_keywords.

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@call.wav" \
  -F "model=zero-indic" \
  -F "enable_keyterm_normalization=true" \
  -F 'keyterm_keywords=["EMI","NACH mandate","bounce charge"]'

Response:

{
  "text": "सर आपका केवाईसी अपडेट नहीं हुआ है और ईएमआई की बाउंस चार्ज लगी है",
  "nlp_analysis": {
    "keyterms": ["KYC", "EMI", "Bounce charge"],
    "normalized_text": "सर आपका KYC अपडेट नहीं हुआ है और EMI की बाउंस चार्ज लगी है"
  }
}

This runs after recognition and writes both nlp_analysis.keyterms and nlp_analysis.normalized_text; the raw text is left as spoken. To influence what the recogniser actually hears — which is what you want for a name it has never encountered — use keyword boosting instead, or both together.

8. Translation

Translation is a separate endpoint, POST /v1/translate. Transcribe first, then send the transcript text with a target language — an ISO 639-1 code (en, hi) or full name (English).

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/translate \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"text": "नमस्ते, यह एक ज़रूरी कॉल है।", "target_language": "en"}'

Response:

{ "translation": "Hello, this is an urgent call." }

9. Profanity hashing

Masks profane words in-place in both the top-level text and each segment's text. The first letter survives and the rest becomes asterisks, so damn comes back as d***.

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@call.wav" \
  -F "model=zero-indic" \
  -F "enable_profanity_hashing=true"

Response:

{
  "text": "This a d*** s**** situation and the b****** hung up on me.",
  "segments": [
    { "start": 0.0, "end": 4.13, "text": "This a d*** s**** situation and the b****** hung up on me." }
  ]
}

10. Custom keyword redaction (hash_keywords)

Regex-based masking of specific terms or phrases. No LLM, fast, deterministic. Use for PII and domain-sensitive tokens.

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@call.wav" \
  -F "model=zero-indic" \
  -F 'hash_keywords=["account number","card number","OTP","aadhaar"]'

Response:

{
  "text": "Please confirm the **** ending 4321 and the **** we sent you yesterday.",
  "segments": [
    { "start": 0.0, "end": 6.4, "text": "Please confirm the **** ending 4321 and the **** we sent you yesterday." }
  ]
}

Matching is literal, so the terms have to be in the same language and script as the transcript. The English list above masks nothing in a Hindi transcript — redacting प्लांट takes प्लांट in the list, not plant.

Masking covers text and segments, not the timing arrays
Both maskers rewrite the top-level text and each segment's text. With response_format=verbose_json the unmasked words are still present in the top-level words array, and in each segment's nbest alternatives. If you are redacting for storage, drop those arrays or don't request them.

11. Word timestamps

Per-word timing with a confidence score. response_format=verbose_json adds one top-level words array covering the whole file — the words are not nested inside segments. Each entry is word, start, end and confidence. Segments carry their own confidence and an nbest list instead. Runs in-pipeline, with no external call and negligible latency overhead.

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@call.wav" \
  -F "model=zero-indic" \
  -F "response_format=verbose_json"

Response:

{
  "text": "टेलीविजन रिपोर्टों में प्लांट से निकलने वाला सफेद धुआं दिखाया गया है",
  "segments": [
    {
      "start": 0.0,
      "end": 5.1,
      "text": "टेलीविजन रिपोर्टों में प्लांट से निकलने वाला सफेद धुआं दिखाया गया है",
      "confidence": 0.9906
    }
  ],
  "words": [
    { "word": "टेलीविजन", "start": 0.08, "end": 0.797, "confidence": 0.9876 },
    { "word": "रिपोर्टों", "start": 0.797, "end": 1.434, "confidence": 0.9873 },
    { "word": "में", "start": 1.434, "end": 1.594, "confidence": 0.9961 }
  ]
}

12. Keyword boosting

Bias recognition toward terms the model has never seen — drug names, product names, people, places, acronyms. Unlike keyterm normalization (§7), which rewrites text after recognition, boosting acts during decoding, so it can recover a word the recogniser would otherwise never produce at all. The corrected term appears in text itself.

Send a short, specific list — the terms that matter for this call, not your whole catalogue. It works on the batch endpoint and on a real-time connection, where the terms go in the opening frame and apply for the life of the connection.

Request:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@consult.wav" \
  -F "language_code=en" \
  -F "boost_phrases=Empagliflozin||NACH mandate||Bengaluru"

Without the lexicon an unfamiliar drug name loses its opening syllable — “Mpagliflozin”. With it, the term comes back intact.

boost_weight applies to Indian languages only
For Indian-language audio, boost_weight controls how hard the decoder is pushed. The default of 10 is also the maximum, and higher values are clamped — above it the recogniser starts inserting boosted terms that were never spoken, which is worse than the mis-recognition you were fixing. For English and auto the terms still apply, but the strength setting has no effect.

Combining features

Every feature above can be enabled on the same request. Here's a full contact-centre configuration:

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "file=@call.wav" \
  -F "model=zero-indic" \
  -F "language_code=hi" \
  -F "enable_diarization=true" \
  -F "enable_speaker_identification=true" \
  -F "project=support_team" \
  -F "enable_emotion_diarization=true" \
  -F "enable_intent_detection=true" \
  -F 'intent_choices=["complaint","inquiry","service_request"]' \
  -F "enable_summarization=true" \
  -F "enable_sentiment_analysis=true" \
  -F 'hash_keywords=["account number","card number","OTP"]'
Latency trade-off
Each intelligence feature adds an extra processing pass on top of the ASR path. Enable only what you consume, don't pay for a summary you won't read.
ASR features | Shunya Labs Docs