AI voice agents are moving quickly from experiments into customer service, lead qualification, appointment booking, and collections. The wider investment trend is clear: a 2025 Genesys survey found that 42 percent of customer-experience leaders viewed increasing AI use as a top priority, while they expected roughly one-third of CX budgets to go to AI-powered technologies over the following year.

Yet the central question is not simply how many calls an AI agent can handle. It is whose speech the system can understand reliably. In Asia-Pacific, a voice agent may sound fluent in a demonstration and still fail when a customer speaks with a regional accent, changes pace, uses a local term, or moves between English and a local language halfway through a sentence. At scale, those are not edge cases. They are the market.

The problem is not just language coverage

Many voice systems still inherit the strengths and weaknesses of training data that is richer in English, especially standardized varieties of English, than in the speech patterns common across APAC. A 2024 study of OpenAI’s Whisper found better recognition for American English than for British and Australian accents, and higher accuracy for native than non-native English speakers. The lesson is broader than one model: supporting a language in a product menu does not guarantee comparable performance for the people who speak it.

Southeast Asia makes that gap particularly visible. The EMNLP 2024 SEACrowd research identified more than 1,300 indigenous languages in the region and warned that English-dominated training data, together with limited local text, image and audio resources, can weaken model quality and cultural representation.

Even widely spoken national languages contain layers of complexity. Research on Indonesian speech recognition notes that available Bahasa Indonesia data has been dominated by read, formal and clean speech, although real calls are often spontaneous, informal and noisy. The study found speaking style had the strongest effect on model performance.

Thai adds tonal distinctions and pronunciation patterns, and a 2024 study reported that even large pretrained Whisper models faced considerable difficulty with Thai speech before local fine-tuning. Vietnamese is also tonal, but variation does not stop there. A 2024 Vietnamese multi-dialect study documented distinct pronunciation across all 63 provincial dialects and highlighted the limitations of speech-recognition approaches when handling that diversity.

Then there is code-switching. A customer in Jakarta might say, “Saya mau reschedule delivery untuk tomorrow morning,” without viewing the sentence as unusual. A 2024 study on Indonesian-English code-switching described this movement between Indonesian, English, and local dialects as part of everyday conversation, while also noting the scarcity of paired code-switched speech data needed to train recognition systems well. A model that processes each language competently in isolation can still lose meaning at the switch.

When a recognition error becomes a business error

Poor voice recognition is sometimes treated as a technical quality issue, measured mainly through word error rates. Businesses experience it differently. A misheard name can prevent a customer record from being found. An incorrectly captured address can derail a delivery. Failure to understand a product term or a polite indirect refusal can turn a promising lead into a dead end. Repeated requests to “say that again” make the customer do extra work precisely when automation is supposed to remove friction.

The commercial consequences should not be overstated, but they should not be dismissed. Qualtrics reported in 2024 that 53 percent of consumers said they would cut spending after a bad customer experience, while communication problems were among the most frequently cited pain points. Not every transcription error causes a customer to leave. However, systematic failures affecting particular accents, languages or speaking styles can depress conversion, increase escalation costs and teach customers that a brand’s service was not designed with them in mind.

Trust is especially fragile in voice. A caller cannot inspect what the system heard before it responds. When the agent answers confidently but incorrectly, the experience can feel less like a minor software fault and more like the company is not listening.

Localization starts with how people actually speak

Inclusive voice AI therefore requires more than translating prompts and selecting a synthetic local-language voice. Businesses need to define the actual speech communities they intend to serve: which regions, accents, age groups, language combinations, vocabulary and levels of formality appear in real conversations. Training and testing should use local call patterns, not only studio recordings or scripted sentences.

Acoustic conditions matter as much as linguistic ones. A model should be tested on mobile networks, inexpensive handsets, street and household noise, overlapping speech, fast speakers, hesitations and interruptions. It should also recognize when confidence is low, ask a useful clarifying question and transfer context to a human rather than forcing the caller to restart.

Evaluation must extend beyond whether the transcript looks accurate. Businesses should measure whether the correct intent was identified, whether names, numbers and addresses were captured properly, whether code-switched requests were completed and whether failure rates differ across dialects or customer groups.

The 2025 GigaSpeech 2 research offers an encouraging signal: models trained on about 30,000 hours of more varied Thai, Indonesian and Vietnamese speech reduced word error rates by 25 percent to 40 percent against Whisper large-v3 on a challenging real-world test set. Better local data and adaptation can materially change performance.

The next phase of AI calling should not be judged only by concurrency, latency or cost per interaction. It should be judged by the range of people who can complete a conversation without changing how they naturally speak. As businesses deploy voice agents across APAC, the more important design question is no longer whether the system can talk. It is who the system has been built to understand, and therefore who it is prepared to serve.


Effie Fang is Director of Business – APAC at Agora, where she helps drive regional adoption of real-time communication and engagement technologies across sectors, including telehealth, conversational AI, media, education, and the future of work. With deep insight into Asia Pacific’s diverse digital markets, she focuses on helping businesses build reliable, interactive, and context-aware communication experiences that can perform across languages, devices, and network conditions.

TNGlobal INSIDER publishes contributions relevant to entrepreneurship and innovation. You may submit your own original or published contributions subject to editorial discretion.

Featured image: Jacek Dylag on Unsplash

AI agents are joining the workforce; Inclusion must be part of the job description