Voice agents
Conversations that run themselves
Inbound and outbound agents that answer in under half a second, pull live data mid-call and hand off to a human with full context.
Voice intelligence for every language
Sonex automatically detects and switches languages, even inside the same sentence, without changing the voice.
250+
Languages, one model
< 200 ms
Time to first audio
Zero-shot
Cross-lingual voice switching
Playground
Type naturally. Mix languages if you want. Pāṇini figures out the rest.
Full conversations, one voice
Loading agents…
SDKs, Pipecat & APIs
Send the text. Pāṇini detects the language, handles code-switching and keeps the same voice.
1import requests23response = requests.post(4 "https://api.sonexlabs.com/v1/speech/stream",5 headers={"Authorization": "Bearer YOUR_SECRET_TOKEN"},6 json={7 "text": "नमस्ते, मैं आपकी कैसे मदद कर सकती हूँ?",8 "voice_id": "lkwlapu6ab",9 },10 stream=True,11)1213# first audio bytes arrive in under 200 msAtlas of scripts
Pāṇini spans 34 writing systems and 250+ languages, including languages rarely supported by modern speech models.
Where each writing system was born
34 scripts · 250+ languages
Plate I · hover or tap a point to hold on that script
7th century
600M+ speakers reached
नमस्ते
namaste
Languages on Pāṇini
Hindi, Marathi, Nepali, Bhojpuri, Maithili, Marwari
One voice
The same voice ID carries its timbre, accent and warmth across every language, so a customer hears one person, not five.
The speaker
The language
आपकी E M I इस महीने पंद्रह तारीख को due है, समय पर payment जरूर कर दीजिएगा। अगर कोई दिक्कत हो तो हमें कभी भी call कर सकते हैं।
हिन्दी · Your EMI is due on the 15th — call us anytime if there's an issue.
Products
Run live conversations, or orchestrate them at campaign scale.
Voice agents
Inbound and outbound agents that answer in under half a second, pull live data mid-call and hand off to a human with full context.
Sonex Flows
Build multi-step calling campaigns visually, the way you would in n8n. Triggers, branches, retries, CRM writebacks, all versioned.
Enterprise
DPDP, HIPAA-ready, GDPR and CCPA, with private VPC or on-premises deployment when the data cannot leave your perimeter.
Under the hood
Use your own OpenAI, Anthropic, or Gemini keys. Pay your contracted model rates directly. We add nothing on top.
Native MCP support. Call any external API, trigger workflows, update CRMs, all mid-conversation.
Managed cloud, private VPC, on-prem, or hybrid. Your data stays in your boundary unless you say otherwise.
Word-level streaming transcription with turn detection built in. Handles interruptions, crosstalk and noisy lines.
Human escalation with full conversation context passed in-band. No re-explanation, no cold transfer.
Separate data planes per workspace. Credentials, voice models and call logs never cross tenant boundaries.
Stateless inference workers auto-scale horizontally. Tested at 10,000 simultaneous sessions without queue latency.
TLS in transit, AES-256 at rest. Tenant-scoped keys. No shared secrets, no plaintext audio storage.
FAQ
Two separate products, priced separately. The Pāṇini TTS API ($0.0313/min) is a pure text-to-speech API, you send text, you get back audio. The voice calling plan ($0.055/min) is the full end-to-end agent stack: speech-to-text, LLM reasoning, Pāṇini TTS, telephony and workflow orchestration, all bundled. Building a phone agent? Use the calling plan. Just need TTS audio in your own system? Use the API.
Pāṇini covers 250+ languages including deep Indian regional dialects: Bhojpuri, Garhwali, Marwari, Saraiki, Assamese, Odia, Sanskrit and many more. Every dialect is billed at the same rate as Hindi or English. No quality tiers, no additional setup.
Pāṇini is trained for telephony conditions, not just clean studio audio. Natural pacing, accurate intonation and proper prosody across 250+ languages. You can also clone a custom voice from a short sample. Audio is streamed rather than generated as a single block, so there's no lag waiting for a full sentence to render.
Interruptions are handled in real time, the agent stops speaking, listens, and continues from where it needs to. Background noise, poor line quality and silence are handled by the STT layer, which is trained on real telephone audio. Hold music is detected so the agent doesn't speak into it.
End-to-end latency is under 500 ms, measured from when the caller finishes speaking to when the agent starts responding. Audio is streamed as the LLM generates output, which keeps the conversation feeling continuous. Latency is consistent across all 250+ supported languages.
Both work. Sonex Labs provides managed phone numbers ready to use immediately. If you already have numbers with Exotel, Twilio or another provider, you connect them without migration. WhatsApp and SMS are available as add-ons.
Yes. The agent queries your systems during the conversation, not just at the start. We support MCP, REST API calls, webhooks and native connectors for HubSpot, Salesforce and others.
Live transfer is built in. You configure the escalation trigger, a keyword, a sentiment threshold, a specific intent, and the agent hands off instantly with full conversation context passed to the human.
Yes, no card required. The free tier includes 10,000 TTS characters and one voice clone. Unlimited agents, the Sonex Flows workflow builder, full API access, CRM integrations, MCP and tool calling, RBAC and webhooks are included at no cost.
All calls, transcripts and customer records are encrypted end-to-end in transit and at rest. Every workspace is tenant-isolated. Sonex Labs is GDPR, CCPA, HIPAA-ready and DPDP compliant, with full audit trails. Enterprise customers can deploy in a private VPC or on-premises for complete data residency control.
Pāṇini covers 250+ languages including deep Indian regional dialects, generates audio in under 500ms, and clones a voice from a 10-second sample. The API is $0.0313/min, billed per 1,000 characters, the same rate across every language with no per-language tiers.
Sonex Labs pairs Pāṇini TTS with the full Sonex Flows calling stack, speech-to-text, LLM reasoning, telephony and workflow orchestration, at $0.055/min end to end. It adds native support for Indian regional dialects (Bhojpuri, Marwari, Garhwali and more) at no extra cost, which most alternatives don't cover.
Start free with no card, or book fifteen minutes with the people who built the model.