Ask a voice assistant something in Bhojpuri. Or Maithili. Or Tulu. Or any of the dozens of languages that hundreds of millions of Indians speak as their first language and use every day of their lives. In most cases, the model either fails outright or produces something that bears a loose resemblance to what was said. The user repeats themselves, switches to Hindi, or gives up entirely.
This is treated as a minor inconvenience in most conversations about AI progress. It is not minor. It is a structural failure that compounds every time a new AI product launches, scales to hundreds of millions of users, and quietly excludes the ones whose languages were not in the training data.
I have been working in AI data collection since the vertical at ConsultBae started from a phone call asking for speakers of six Indian languages. That was the beginning. What I have seen since then is an industry that understands the problem intellectually and has not solved it in practice.
The scale of what is missing
The numbers are worth sitting with for a moment. India has 22 languages listed in the Eighth Schedule of the Constitution. Beyond those, the 2011 census recorded 121 languages spoken by 10,000 or more people. There are hundreds of dialects with meaningful speaker populations that do not appear in that count at all.
Of these, a small number have reasonable representation in global speech training data. Hindi has more than most, though even Hindi models struggle significantly with regional accents and dialectal variation. Tamil, Telugu, Bengali, and Marathi have some coverage. Beyond that, the representation drops sharply. Languages like Bhojpuri, spoken by an estimated 50 million people, Maithili, Konkani, Dogri, Santali, and dozens of others have minimal or no meaningful presence in the training datasets behind the speech models most people use every day.
This is not because these languages are rare. It is because the data was never collected.
A language spoken by 50 million people should not be invisible to a speech model. The fact that it is has nothing to do with the language and everything to do with whose data collection budgets prioritised it.
Why this happened
Speech models learn from data. The data that exists in large, accessible quantities on the internet is overwhelmingly in English, followed by a small group of other high-resource languages: Mandarin, Spanish, French, German, and a handful of others. These languages dominate because they dominate the written and recorded internet. Decades of digital content creation happened in these languages first, at scale, and it was that content that became the raw material for training.
Indian languages, with the exception of a few, did not accumulate digital footprints at the same pace or scale. The reasons are structural: lower rates of internet penetration in earlier years, content creation concentrated in English and a small number of dominant regional languages, and limited investment in building the digital infrastructure that would have created usable training data as a byproduct.
The result is that when the AI industry began building speech models at scale, it trained on what existed. What existed was English-heavy, and the models reflected that. Low-resource Indian languages were not deliberately excluded. They were simply absent from the data, and absent from the data means absent from the model.
Who pays the price
The people who pay the price are not the ones who built the models. They are the hundreds of millions of Indians for whom English is not a first, second, or even comfortable third language, and for whom AI products that work well in English deliver a fundamentally worse experience in their own language.
This includes rural populations accessing government services and healthcare information through voice interfaces. It includes first-generation smartphone users in tier-two and tier-three cities for whom voice is the most natural input method. It includes elderly speakers who have never been comfortable typing but would engage readily with a voice interface that understood them. It includes children being educated in regional medium schools for whom AI-assisted learning tools are effectively unavailable in the language they actually think in.
Each of these groups is large. Together, they represent a population that is being systematically underserved by AI systems that were not built with their languages in mind because the data required to build those systems was never collected.
Why this is a commercial problem, not just a social one
The next 300 to 500 million internet users coming online in India over the next decade will not be English-first. They will be regional-language-first. They will interact with technology through voice before they interact through text. The AI products that want to serve them, and the commercial opportunity attached to that user base is enormous, will need models that work in the languages those users actually speak.
This is not a charitable argument for building better language coverage. It is a market argument. The companies and models that solve the Indian language data problem early will have a significant structural advantage when that user base scales. The ones that do not will be building retrofit solutions for a population they should have been training for years earlier.
We are already past the point where this can be solved by scraping the internet. The written digital content in low-resource Indian languages is not sufficient to build the speech models these users need. The data has to be collected directly, from native speakers, in controlled conditions, at scale.
What collecting this data actually requires
Collecting speech data for a low-resource Indian language is not a platform problem. You cannot set up a task on a freelance marketplace and expect the right contributors to appear. It requires finding native speakers of the specific language or dialect, in sufficient numbers, with the demographic diversity that makes the data generalisable across the real speaker population.
It means managing for accent and dialect variation within a single language, because Bhojpuri spoken in eastern Uttar Pradesh sounds meaningfully different from Bhojpuri spoken in Bihar, and a model trained only on one will fail on the other. It means building contributor networks in geographies that are not well-served by standard crowd platforms. It means quality processes that can verify language authenticity, because a contributor who speaks Hindi as a first language and Bhojpuri as a second will produce data that sounds plausible and trains the model incorrectly.
None of this is unsolvable. All of it requires operational infrastructure that was built for exactly this kind of work rather than adapted from something else.
Where ConsultBae fits into this
The AI data vertical at ConsultBae started with six Indian languages. That was not a strategic entry point we planned. It was the first project that came to us. But it meant that the operational muscle we built from day one was built around collecting language data from specific, hard-to-reach speaker populations in India, and then replicating that capability across other languages and geographies.
We have collected speech data across Indian languages including Hindi, Bengali, Marathi, Telugu, Kannada, Tamil, and others. We understand the variation within languages, the quality signals that distinguish authentic native speech from non-native approximations, and the logistical reality of finding and mobilising speakers of low-resource languages at the volumes AI training requires.
The Indian language gap in global speech models is a real problem and a solvable one. The solution is not a new algorithm. It is collected data, at scale, from the right speakers. That is work we know how to do, and it is work that needs to happen before the next wave of AI products tries to serve the users those models are currently failing.
Amitt Agrawaal is the Founder of ConsultBae. He has spent six years building ConsultBae's operations across recruitment, e-learning, and AI data collection across 100+ countries.
Building a speech model for Indian languages?
ConsultBae collects speech and language data across Indian languages and dialects at scale. Let us talk about what your model needs.
Talk to us


