Article

The Catalan Test: Is Your Voice AI Truly Multilingual?

Co-official languages are not dialects of the dominant one. ToumAI's approach to underserved regional languages points to how voice AI should treat Catalan too.

ToumAI Editorial Team

September 2026 - 3 min read

Share :
Catalan voices, real understanding.

9.2M 

People speak Catalan across Spain, France, Andorra, and Italy 

93.4% 

Of Catalonia's population understand Catalan 

Spanish autonomous communities where Catalan or Valencian is co-official 

Figures from the Catalan government's language use survey (2023) and Wikipedia/Ethnologue speaker estimates (2022). Regional figures vary by survey, year, and methodology. 

ToumAI's approach to voice AI starts from a simple premise: systems built for a handful of dominant, high-resource languages fail everyone else, not because those other languages are harder to model, but because they are rarely given the data, the dialect awareness, or the product priority to be modeled properly. 

Catalan is a useful test of whether that premise holds outside the markets ToumAI was built for. It is spoken by more than nine million people, is co-official alongside Spanish in Catalonia, the Balearic Islands, and Valencia, and it is, by any linguistic measure, its own language, not a regional accent of Castilian Spanish. Voice AI that quietly treats it as one is repeating the same mistake ToumAI's core markets have already shown does not work. 

The pattern is familiar 

A voice assistant that mishears Catalan as poorly pronounced Spanish, or simply cannot process it at all, is making a familiar category error: collapsing a distinct language into an approximation of the dominant one nearby. The customer experience cost is predictable. Self-service fails, the call escalates to a human agent, and the system has told the customer, functionally, that their language does not count as a valid input. 

This is not a hypothetical risk. Catalonia alone represents a substantial share of Spain's banking, telecom, and insurance activity. A contact centre or app that defaults to Castilian Spanish for Catalan speakers is running the same self-service failure rate seen anywhere a regional or minority language gets quietly defaulted to the dominant language next door. 

Variation inside the language, not just around it 

Catalan is not internally uniform. Central Catalan, spoken around Barcelona, differs from Valencian and from the Balearic varieties spoken across Mallorca, Menorca, and Ibiza. A system trained only on the Barcelona standard and marketed as full “Catalan support” can still fail a meaningful share of Catalan speakers, the same way training on any single regional variety of a language leaves its other varieties poorly served. 

ToumAI's dialect data work exists precisely because language coverage is never finished at the level of a single label. It has to extend to the regional and generational variation inside that label, and to the natural code-switching between Catalan and Spanish that bilingual speakers use as a matter of course, not an edge case. 

 What language-aware voice AI actually requires ? 

Adding Catalan to a language dropdown is not the same as supporting Catalan speakers. It requires voice data collected from real speakers across the language's regional varieties, annotation that captures code-switching between Catalan and Spanish rather than treating it as noise, and recognition built around Catalan's own grammar and pronunciation rules rather than a set of substitutions layered onto a Spanish model. 

That is the same technical discipline behind ToumAI's approach to underserved languages generally: real speech, not idealized speech; regional variation, not a single standard; code-switching treated as the default, not the exception. 

A pattern that extends well past one language 

Catalan is one instance of a much broader pattern. Basque, Galician, Welsh, Breton, and dozens of other co-official or minority languages face the same structural neglect: real languages with real speakers, treated by most voice AI as optional variants of whichever language sits next to them on a map. ToumAI's broader lesson, that inclusive voice AI has to be designed around how people actually speak rather than a simplified map of “one country, one language,” applies just as directly here. 

Voice AI that only works well for a market's dominant language is not inclusive by default, wherever that market is. It has to be built, deliberately, language by language and dialect by dialect, to work for the people a simplified map leaves out. 


Exploring voice AI for Catalan, another co-official language, or a regional dialect your customers actually speak?


See what it looks like on your side of the cliff.

Book a 30-minute session with our team — we'll show you exactly where AI can move the needle in your contact center, in your language.

Tags:

CatalanCo-official LanguagesLanguage CoverageDialect-Aware AI

Keep reading

Related articles

View all articles

Cookies

We use cookies to keep this site reliable, understand performance, and improve your experience.