Partnerships
xAIa's journey with Soniox: 1,000,000+ voice AI calls later
How xAIa's voice AI agents ran 1,000,000+ production calls on Soniox across the GCC: Gulf Arabic, code-switching, alphanumerics, latency, cost and TTS v2.

- Author
- Team xAIa
- Published
- 24 September 2026
- Reading time
- 10 min read
xAIa's Voice AI agents have run on Soniox speech-to-text for the past year, across telecom, utilities, distribution, real estate, healthcare and government in the UAE and the wider Gulf. Here is what we learned, why we chose Soniox, and what we are building together next.
xAIa is a Dubai-based AI consultancy and implementation firm that builds and deploys voice AI agents across every department, the AI employees that answer and place calls for enterprises and government entities across the UAE and the wider GCC. By September 2026, xAIa's AI employees had handled more than 1,000,000 production calls on Soniox speech-to-text.
Every buyer evaluating voice AI eventually asks what happens at real volume, on real traffic, at their peak. A demo answers none of that. A year of production does.
This is the part of a deployment nobody asks about in a sales meeting and everybody feels on the call. We spent that year on one speech provider, and it is worth saying in public why, and what it looked like at a million calls.
What we build, and what the Gulf demands of it
xAIa designs, builds and deploys AI employees across voice and chat for enterprises and government entities across the GCC. They answer and place calls, qualify inquiries, book appointments, place orders and resolve service requests, then write the outcome back into the systems of record and hand to a human when it matters.
Every live call passes through three models: speech-to-text to hear the caller, a reasoning model to decide what to say, and text-to-speech to reply. The caller judges the sum. And in this market, the first of those three carries a load most speech models were never built for.
Languages mix inside a single sentence.
A caller opens in Arabic, reads an account number in English, returns to Arabic to explain the problem, then drops an English product name mid-clause. This is the default shape of a business call in the UAE.
Calls are dense with structured data.
Emirates ID references, order codes, plate numbers, amounts. A transcript that is 97 percent correct but scrambles one digit has failed the call.
Consistency beats average speed.
A caller tolerates 300ms every turn. A caller does not tolerate 300ms on three turns and 1.5 seconds on the fourth, because the fourth is where they talk over the agent and the conversation collapses.
The audio is never clean.
Mobile networks, contact centre headsets, callers interrupting and restarting mid-thought.
Underneath all four sits Arabic, which most speech models handle by handling Modern Standard Arabic. Nobody uses it on a phone call. A caller in Dubai speaks Emirati or Gulf dialect, and a model trained on MSA does not sound slightly off. It sounds like it does not know who it is talking to.
Why Soniox
Soniox is the complete speech layer of the voice stack: one API for speech-to-text, text-to-speech and translation across 60+ languages, with sub-200ms latency and native-speaker accuracy. Four of those claims mapped directly onto what a Gulf deployment breaks on.
Language switching mid-sentence, with no configuration.
One unified model across all 60+ languages. Nothing to load, nothing to switch, and no seam for a code-switched Arabic and English sentence to fall through.
Sub-200ms streaming that holds turn to turn.
Fast on average is easy. What decides whether a caller talks over the agent is the gap between the median and the tail, and on Soniox that gap is small.
Alphanumerics and context, captured as spoken.
Phone numbers, reference codes and Emirates IDs come through intact. Domain vocabulary is supplied when the session opens and used while listening, rather than patched afterwards.
A cost line that survives a million calls.
Real-time speech is published at $0.12 an hour. At our volume, that single line item is the difference between a deployment that scales and one that stalls at pilot.
Soniox carries the posture a regulated deployment requires, SOC 2 Type 2, ISO 27001, HIPAA and GDPR, with audio processed in memory and never stored. Where a mandate goes further, our own deployments run in region or on-premise.
We believed in Soniox early, and a million calls later that decision has held up every month. Their models let us handle calls the way our customers actually speak, across the Gulf dialects, English, Hindi and the mix of languages we serve across the UAE, Europe and the US. With TTS v2 now embedded in our AI employees, the next stage of this work is stronger and we are looking forward to what Soniox brings next so we can carry it across the region.
A year in production
We went live on Soniox at stt-rt-v2. Twelve months later we are on stt-rt-v5, running past 600,000 calls a month, and the integration has not changed once.
Over those twelve months the platform moved to sub-200ms streaming, extended language switching to 60+ languages with no configuration, and sharpened the alphanumeric and context handling that decides whether an agent captures a reference number the first time. Our volume went from first deployment to over a million calls. Their model got faster and more accurate underneath it, without a single migration on our side. That is what a speech layer is supposed to do.
From transcription to conversation
A foundation model gives you accurate transcription of Gulf Arabic. It does not give you an agent that speaks it correctly back.
Transcription is one problem. Delivery is another. Gulf Arabic varies by speaker and by market, in vocabulary, in verb forms, and in how a person is addressed, and an agent that gets it approximately right still sounds wrong to the person on the line. That layer is built per deployment, tuned against the client's own call recordings, and regression-tested with every release. We have written about what it takes to actually speak Arabic on a live line.
Soniox makes sure the agent hears the caller. We make sure it answers like someone from here.
TTS v2
Soniox released TTS v2 this year, built for the failures that only show up once voice AI is live: phone numbers scrambled, email addresses spoken wrong, foreign names mispronounced, mixed-language text falling apart. We had hit all four.
Alphanumerics now come out exactly as written, so a verification code, an account number or a delivery address is read correctly the first time, and every voice speaks every language we deploy in, so one voice identity carries a brand across Arabic, English and Hindi without changing character mid-call.
What we are building together
A million calls of Gulf Arabic on real telephony audio, switching language mid-sentence, reading reference numbers over mobile networks, is a body of production evidence very few teams hold. We know exactly where speech breaks in this region and what holds under it, and that is why we are extending our work with Soniox. The problems that decide whether voice AI works in this market, language switching, alphanumerics, latency that stays flat under load, are the ones their models have already solved in production.
TTS v2 is now going into the AI employees we deploy for enterprises and government across the region, and it leads a series we are launching this month.
More on what we are building together soon.
FAQ
Which speech to text engine handles Gulf Arabic on live calls?
xAIa runs its voice AI agents on Soniox speech-to-text, and has done for the past year, across more than 1,000,000 production calls. The deciding factors were one unified model across 60+ languages so a caller can switch language mid-sentence, sub-200ms streaming that holds turn to turn, and alphanumerics captured as spoken.
Can speech to text handle Arabic and English in the same sentence?
Yes, when one model covers both languages instead of switching between per-language models. A business call in the UAE routinely opens in Arabic, reads an account number in English and returns to Arabic. Soniox runs one model across all 60+ languages with nothing to configure, so a code-switched sentence has no seam to fall through.
Is accurate Arabic transcription enough for a voice AI agent?
Accurate transcription is the first half of it. Soniox makes sure the agent hears a Gulf Arabic caller correctly. Speaking it back correctly is a separate layer, because Gulf Arabic varies by speaker and by market in vocabulary, in verb forms and in how a person is addressed. xAIa builds that layer per deployment, tuned against the client's own call recordings and regression-tested with every release.
How much does speech to text cost for a voice AI agent?
Soniox publishes real-time speech-to-text at $0.12 an hour. At the volume xAIa runs, above 600,000 calls a month, that single line item is the difference between a deployment that scales and one that stalls at pilot.
What speech to text latency does a voice AI agent need?
The number that decides a call is the gap between the median and the tail, because a caller talks over the agent on the slow turn. Soniox publishes sub-200ms streaming speech-to-text, and across xAIa's production deployments response latency runs at roughly 500ms end to end.
Where is call audio processed, and is it stored?
Soniox processes audio in memory and does not store it, and carries SOC 2 Type 2, ISO 27001, HIPAA and GDPR. Where a mandate goes further, xAIa deployments run in region or on-premise.
How to evaluate speech to text for a voice AI agent
This is the checklist we would hand anyone running the same evaluation, drawn from what broke on our own calls.
| What to test | How to test it |
|---|---|
| Code-switching mid-sentence | Play a recording that opens in Arabic, reads an account number in English and returns to Arabic. Ask whether the provider switches language by configuration or handles every language inside one unified model. |
| Gulf dialect on real telephony audio | Run your own call recordings, in the dialect your callers speak, over the lines they call from. Ask what the provider benchmarked on, because Modern Standard Arabic results say very little about a Dubai phone call. |
| Alphanumerics as spoken | Read Emirates ID references, order codes, plate numbers and amounts into a live session, then check the transcript digit by digit against the audio rather than against the overall accuracy score. |
| Latency at the tail | Ask for the slowest turn as well as the median, and insist both figures come from streaming audio rather than uploaded files. A batch benchmark tells you nothing about a live turn. |
| Cost per hour at your real volume | Take the published real-time rate and multiply it by the call minutes you expect in a full month. A pilot hides this line item. |
| Where the audio goes | Ask whether audio is stored or held only in memory, which certifications the provider holds, and whether the deployment can run in region. |
Speak to us and we will show you what an AI employee handles in its first week.




