Speech-to-Text (ASR)
Zero STT is Shunya's speech recognition family. One API surface for batch and streaming, a choice of four models tuned for different domains, and an intelligence layer that adds diarization, emotion, intent, and more on top of the transcript.
The examples below send Authorization: Bearer $ACCESS_TOKEN. The speech APIs accept only a short-lived access token — never your API key directly.
From the console (recommended). Open the console, click Generate token next to your API key, and copy it — then set it:
export ACCESS_TOKEN="eyJhbGciOiJSUzI1NiIs…paste-here"Or mint it from your API key — the path for production, where your app refreshes the token as it nears expiry (the response carries expires_in):
export ACCESS_TOKEN=$(curl -s -X POST https://app.shunyalabs.ai/api/auth/token \
-H "api-key: $SHUNYALABS_API_KEY" | jq -r .token)How it fits together
Endpoints
| Mode | Endpoint | Use for |
|---|---|---|
| Batch | POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions | Uploaded files, post-processing, async jobs. |
| Streaming | wss://asrv2prod.shunyalabs.ai/v1/realtime | Live transcription, voice agents, IVR. |
| Health | GET https://asrv2prod.shunyalabs.ai/health | Liveness checks. No auth. |
| Languages | GET https://asrv2prod.shunyalabs.ai/languages | Returns supported language names, ISO codes, and scripts. |
| Speakers | /v1/speakers/* | Register, list, identify, delete voice profiles for speaker identification. |
Batch vs Streaming
Same models, same intelligence layer, different transports. Pick batch when you have a complete audio file in hand. Pick streaming when audio is arriving live and you want partial transcripts as the speaker is still talking.
Accepts multipart/form-data. Required fields: file (or url) and model.
zero-indic, General Indian languages (Hindi, Tamil, Telugu, Kannada, Marathi, Bengali, etc.)zero-med, Medical/clinical audio, auto-applies medical terminology correction via MedGemmazero-codeswitch, Code-switched speech (Hinglish, Tanglish, etc.), auto-restores English words to Latin scriptzero-universal, 99-language Whisper model, English, European, Asian, and African languages
- Endpoint
- POST
/v1/audio/transcriptions - Host
asrv2prod.shunyalabs.ai- Content type
multipart/form-data- Auth
- Bearer
<ACCESS_TOKEN> - Required
file(orurl) andmodel- Default response
verbose_json
Real-time streaming transcription over WebSocket. Supports binary mode (raw PCM/ulaw/alaw bytes) and JSON mode (base64-encoded audio frames).
ulaw: G.711 mu-law (8-bit), Telephony (8 kHz)alaw: G.711 A-law (8-bit), Telephony (8 kHz)int16: 16-bit signed PCM, General recordingfloat32: 32-bit IEEE float, Pre-processed audio
- Endpoint
wss://asrv2prod.shunyalabs.ai/v1/realtime- Init message
- JSON config (first frame)
- Sample rate
16000Hz (default)- Chunk size
2.0s (default)- Silence threshold
0.8s (default)
Source: Shunyalabs ASR Gateway API Reference (31 March), "Base Call" and "WebSocket Streaming API" sections, reproduced verbatim.
Pick a model
Every request takes a model field. There are four to choose from:
Hindi, Tamil, Telugu, Kannada, Marathi, Bengali and 50+ Indian languages. The default for Indic content.
View model →204-language Whisper-class model. English, European, Asian, African. Auto-detects language when you don't know it.
View model →Clinical / medical speech with drug, procedure, and diagnosis vocabulary. HIPAA-cleared. Auto-applies medical terminology correction.
View model →Native handling of mixed Hindi-English (Hinglish), Tamil-English (Tanglish) and similar blends. Automatic code-switch restoration.
View model →Your first request
Minimum viable: file, model, bearer token. Add language_code whenever you already know the language — it is what decides which model actually runs.
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-F "file=@meeting.wav" \
-F "model=zero-indic" \
-F "language_code=hi"Or pass a URL instead of uploading a file:
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-F "url=https://example.com/call.wav" \
-F "model=zero-indic" \
-F "language_code=hi"hi, en) or the full English name (Hindi) and the request is routed straight to that language. Leave it out, or send auto, and the service runs language identification first and routes on what it detects. Detection is good but not free: on short, noisy or code-mixed audio it can pick the wrong language, and every downstream stage then inherits that choice. If you know the language, say so — it is the single highest-leverage field on this endpoint.