Guide
ElevenLabs Voices for UAE Voice Agents: Languages, Latency, Licensing
Quick answer
ElevenLabs supplies the voice layer in most UAE AI phone agents. For live calls the relevant model is Flash v2.5: about 75ms inference across 32 languages, including Arabic and Hindi but not Malayalam or Urdu. There is no Gulf serving region, so Dubai traffic is handled from Europe or Singapore, and commercial use requires a paid plan.
When a UAE business asks what voice their AI receptionist will use, the honest answer involves three separate decisions that most vendors collapse into one. Which model synthesises the speech determines the language coverage. Where the request is served determines how long the caller waits. And which plan you are on determines whether you are allowed to use the audio commercially at all. Everything below comes from ElevenLabs' own published documentation and pricing, cited at the end.
Which ElevenLabs model should a phone agent actually use?
ElevenLabs publishes several text-to-speech models, and only one of them is designed for live conversation. Their model documentation is explicit about the split: Flash v2.5 is the fast model built for real-time use, Multilingual v2 is the higher-fidelity model built for produced content, and Eleven v3 is the expressive model with the widest language list.
For a phone agent, the character limits and emotional range that dominate marketing comparisons barely matter. A receptionist answers in sentences, not chapters. What matters is inference speed and whether the caller's language is on the list.
| Model | Languages | Stated latency | Best fit |
|---|---|---|---|
| Flash v2.5 | 32 | ~75ms inference | Live voice agents and phone calls |
| Multilingual v2 | 29 | Higher than Flash | Produced audio; better number handling |
| Eleven v3 | 70+ | Not positioned for real-time | Expressive, long-form, multi-speaker |
ElevenLabs' own model selection guide recommends Flash v2.5 for agent platforms and real-time applications, and notes it carries a slight reduction in audio quality compared with Multilingual v2. On a phone line compressed to 8kHz, that quality gap is largely inaudible. The latency gap is not.
Which languages does it cover for a Dubai caller list?
This is where the model choice stops being a technical preference and starts constraining who you can serve. Flash v2.5 supports the 29 languages of Multilingual v2 plus Hungarian, Norwegian and Vietnamese. The Multilingual v2 language list, quoted directly from the docs, includes Arabic (Saudi Arabia, UAE), Hindi, Tamil, Filipino, and most major European languages.
Read that list carefully against a Dubai caller mix and two gaps appear. Malayalam and Urdu are not in Flash v2.5 or Multilingual v2. They appear only in Eleven v3's 70+ language list, which is the expressive model rather than the low-latency one. For a business whose callers are largely Emirati, GCC Arabic-speaking, or Hindi-speaking, this is a non-issue. For a Deira trading company or a clinic whose front desk fields heavy Malayalam traffic, it is the deciding constraint, and no amount of prompt engineering works around it.
The honest options are to run Eleven v3 for those languages and accept a different latency profile, to route Malayalam and Urdu callers to a different text-to-speech provider, or to route them to a human. We cover why these specific languages matter commercially in Malayalam, Tagalog and Urdu callers in Dubai. What you should not do is assume that a vendor demo in English tells you anything about the languages your phone actually rings in.
On Arabic specifically, the docs listing "Arabic (Saudi Arabia, UAE)" is a statement about supported locales, not a promise of pitch-perfect Emirati dialect. The practical register remains Khaleeji-neutral, which is what most callers expect from a business line anyway. The dialect question is worth understanding properly before you buy, and we work through it in Emirati Arabic and what voice AI can and cannot do with dialect.
What latency does a caller in Dubai actually experience?
The widely quoted 75ms figure carries a footnote in ElevenLabs' own documentation: it excludes application and network latency, and refers to model inference time only. The number a caller in Al Quoz perceives is a different number entirely.
ElevenLabs' latency optimisation guide publishes expected time-to-first-byte by region when using Flash models over websockets: 100 to 150ms for North America, Europe and South East Asia, and 150 to 200ms for South Asia and North East Asia. There is no Middle East row in that table, and the currently used serving regions are listed as the USA, the Netherlands and Singapore. A request originating in the UAE is therefore served from Amsterdam or Singapore, not from the Gulf. You can confirm which region handled a given request by reading the x-region response header.
Two further factors move the number, both documented. Voice type affects speed, with default and instant-cloned voices faster than professional voice clones. And once your plan's concurrency limit is reached, further requests queue, which the docs say typically adds around 50ms. None of this makes ElevenLabs unsuitable for UAE deployments. It does mean that anyone quoting you 75ms end-to-end for a Dubai call is quoting the wrong number.
Tip
apply_text_normalization parameter, which is Enterprise-only for v2.5 models. Test this before launch, not after.What happens when the phone rings five times at once?
Concurrency is the constraint that surprises businesses moving from a demo to production, because a demo never has two callers. ElevenLabs caps simultaneous requests by plan tier, and once the cap is hit the rest queue. The published limits for text-to-speech are:
| Plan | Multilingual v2 | Flash |
|---|---|---|
| Free | 2 | 4 |
| Starter | 3 | 6 |
| Creator | 5 | 10 |
| Pro | 10 | 20 |
| Scale / Business | 15 | 30 |
| Enterprise | Elevated | Elevated |
A clinic on a Creator plan has ten concurrent Flash generations available, which is comfortable for a single busy line and thin for a multi-branch group during a Monday morning rush. Note also that a concurrent request is not the same as a concurrent call, since a call generates many short synthesis requests rather than one long one. The practical answer is to size from your busiest observed hour rather than your monthly average, then watch the current-concurrent-requests response header in production.
What does the licence actually allow you to do?
This is the section most buyers skip and the one most likely to cause a problem. ElevenLabs' terms of use draw a hard line: a free user "may only use the Services for non-commercial purposes", while a paid subscriber may use them commercially. Answering your customers' calls is commercial use. The commercial licence starts at the Starter plan, which ElevenLabs prices at $6 per month, with prices stated to exclude taxes, levies and duties.
On ownership, the terms are more generous than people assume: you retain rights to your input and to the generated output. What you also grant, in the same section, is a perpetual, irrevocable, worldwide and sub-licensable licence for ElevenLabs to use that content to provide and improve its services, including developing new ones. The terms state that your voice will not be commercialised on a standalone basis without permission, and that you can opt out of your content being used for training through the Data use menu in your account settings. For a UAE business handling customer conversations, that opt-out is a setting worth changing on day one rather than discovering later.
Voice cloning adds a consent obligation. The cloning documentation describes two methods: instant voice cloning, which conditions on a short sample of under two minutes, and professional voice cloning, which fine-tunes on roughly 30 minutes of high-quality audio and requires a Creator plan or above. Both include a voice-captcha verification step, and the docs are candid that verification only confirms the requester was present, not that they own the voice. The responsibility for authorised use sits with you. If you clone a staff member's voice for your reception line, get their written consent and keep it, and agree what happens to that voice if they leave.
Legal caveat
How does this fit a real UAE deployment?
Start from the caller mix rather than the voice demo. List the languages your phone actually rings in, check each one against the Flash v2.5 list, and decide explicitly what happens to the languages that fall outside it. That single exercise prevents the most common post-launch surprise, which is a caller who cannot be served by a system that was signed off in English.
Then decide who holds the account. Voice platforms generally expect you to bring your own key: Vapi's integration guide requires an ElevenLabs API subscription and your key pasted into the platform before a cloned voice library will sync. That has a governance consequence worth deciding deliberately. If the key belongs to your agency, your cloned voice, your concurrency ceiling and your billing all sit inside someone else's account. The per-minute economics of the surrounding platform are a separate question, which we work through in Vapi pricing for UAE businesses.
For the multilingual outbound agent we built for IT World Trading LLC, the language coverage and the telephony had to be settled before any voice was chosen, because a UAE deployment lives or dies on TDRA-compliant call routing and an on-site install that actually works with the client's existing lines. The voice is the last decision, not the first. If you want that sequencing handled for you, our AI receptionist for Dubai and AI appointment booking deployments cover model selection, language routing, and the licensing setup as part of the build, with the scope and pricing set out on our services page.
Sources
- ElevenLabs — Models (official docs; Flash v2.5, Multilingual v2, Eleven v3 languages, latency, concurrency limits)
- ElevenLabs — Latency optimization (official docs; regional TTFB, serving regions, voice-type impact)
- ElevenLabs — Pricing (official; plan tiers and credit allowances)
- ElevenLabs — Terms of Use (official; non-commercial free tier, content licence, PHI restriction)
- ElevenLabs — Voice cloning: how it works (official docs; IVC vs PVC, verification, plan eligibility)
- ElevenLabs — Data residency (official docs; U.S. default storage, EU/India/Singapore isolated environments)
- Vapi — ElevenLabs custom voices (official docs; API subscription and key required)
Frequently asked questions
Which ElevenLabs model should a UAE voice agent use?
Does ElevenLabs support Arabic for UAE callers?
Can I use ElevenLabs for free for my business phone line?
How much latency does ElevenLabs add to a call from Dubai?
Can I clone my receptionist's voice for the AI agent?
Anam Jalal
Founder & CEO, MAJ Leads
Anam Jalal is the founder of MAJ Leads, a Dubai-based AI voice agent company deploying TDRA-compliant AI receptionists and callers for UAE clinics, brokerages and SMEs — working hands-on across UAE telephony and CRM integrations, from SIP provisioning to TDRA compliance configuration.
Read more about Anam →Related articles
Guide
Can an AI Voice Agent Speak Emirati Arabic? What Actually Works for UAE Callers
Can an AI voice agent handle Emirati Arabic callers? We break down the dialect landscape, what Khaleeji-neutral MSA Arabic actually sounds like in practice, and what to demand in a live demo.
Industry
Malayalam, Tagalog, Urdu: The Languages Your Dubai Receptionist Is Losing Leads In
Hundreds of thousands of Malayalam, Tagalog and Urdu speakers live and spend in Dubai. If your receptionist can't serve them in their language, your competitors pick up those leads. Here's how AI changes the equation.