Ramkhelawan Yadav* grows wheat on two acres of land in Banda district, Bundelkhand. Last year, a government-backed agri-tech startup rolled out a WhatsApp chatbot in his village — promising farmers real-time crop advice, weather alerts, and guidance on subsidies. Ramkhelawan tried it once. He typed his question the way he speaks — a mix of Bundeli dialect, local idiom, and half-formed sentences. The bot replied in textbook Hindi, formal and stiff, answering a question he hadn't asked. He hasn't used it since.
* Fictional, based on real patterns.
This is not a story about one farmer and one bad chatbot. It is a story about who artificial intelligence was built for — and who it silently leaves behind.
1 · The Internet Is Not a Mirror of Humanity
Every large language model — ChatGPT, Gemini, Claude, and the rest — learns language by reading enormous quantities of text scraped from the internet. The logic seems sound: the internet is vast, global, democratic. Surely it reflects how the world actually speaks?1
It does not.
More than 55% of all web content is in English. Hindi, spoken by nearly 600 million people, accounts for less than 0.1% of indexed content. The visualisation below makes this point more starkly than any sentence can. And Hindi is the lucky one among Indian languages.
Bhojpuri has a rich oral tradition, folk music, cinema, and a living culture stretching across Bihar, eastern Uttar Pradesh, and the Terai belt of Nepal. In the training data of most leading AI models, it is virtually absent. The same is true for Chhattisgarhi, Gondi, Tulu, Kumaoni, Mewari, and Santhali.2
The disparity has nothing to do with how many people speak a language. It is about who historically had the infrastructure, literacy, and economic incentive to put their language on the internet. Rural Indians, overwhelmingly, have not.
2 · When "Hindi Support" Is Not Enough
Technologists sometimes point out that major AI platforms now support Hindi, Tamil, Bengali, and several other Indian languages. While this is an important step forward, it obscures a deeper issue: the Hindi AI understands is often not the Hindi that most native speakers actually speak.3
Dialect is not slang. It is a fully developed linguistic system shaped over centuries by geography, culture, and community. When AI treats dialect speakers as edge cases to be corrected, it is not making a technical decision — it is making a cultural one.
Standard Hindi — taught in schools, used in newspapers, spoken on national television — is a relatively uniform, sanitised register.The Hindi spoken in everyday life is far more varied and adaptive. It blends regional dialects, modifies grammar to reflect local speech patterns, and shifts naturally between formal and informal registers, creating a linguistic richness that standardized language often fails to reflect.
The stakes are far from theoretical. AI systems are now being used to deliver agricultural guidance, support healthcare decisions, provide legal assistance, and determine eligibility for government welfare programs. If rural users cannot interact with these systems in the language they actually speak, the consequences can be severe. A farmer may lose access to a deserved subsidy, while a patient may receive inaccurate dosage instructions or misunderstand critical medical advice.
3 · Voice-First India Meets Text-First AI
There is another mismatch that goes beyond vocabulary: the medium itself. Most large language models are, at their core, text models. But rural India is not a text-first culture — it is a voice-first one.4 WhatsApp voice notes travel faster than typed messages through villages. Instructions are given by phone call. Knowledge is passed down orally.
Speech recognition has improved dramatically. OpenAI's Whisper model handles a broad range of languages and accents with impressive accuracy. But "impressive accuracy" for a news anchor's Hindi is not the same as accurate transcription for a Rajasthani farmer's Marwari, or a tribal woman from Jharkhand speaking Santali.5 Accents, ambient noise, code-switching mid-sentence — these are everyday realities that most commercial speech models still struggle with.
4 · Who Is Trying to Close the Gap?
The problem is known. There are serious, committed people working on it — though rarely with the resources the scale demands.
AI4Bharat — IIT Madras
Building open-source NLP datasets and models for Indian languages. Their IndicBERT, IndicTrans, and related projects are the most rigorous efforts toward genuinely multilingual AI within India's linguistic landscape — while remaining acutely aware of how many dialects remain uncovered.
Bhashini — Government of India
An ambitious platform aiming to enable translation, transcription, and voice interfaces across 22 scheduled languages. A meaningful policy commitment — but 22 languages still leave hundreds of dialects and unscheduled tongues largely outside the frame.
Jugalbandi — Microsoft Research × Karya
A WhatsApp chatbot helping rural citizens navigate government welfare schemes in their own languages, designed for limited literacy and voice input. Early results are promising, though its scope remains narrow relative to what AI could offer.
Karya
Rather than scraping the internet harder, Karya pays rural Indians — fairly, above minimum wage — to record voices and create language data in their own dialects. The recognition that these communities should be compensated for making AI better, not merely harvested, is a model worth replicating at scale.
These initiatives are promising. They are also underfunded, understaffed, and operating against the grain of an industry that moves fast and profits from markets that already have money.
5 · Language Is Not Just Communication — It Is a Worldview
There is a deeper point beneath all the statistics. Language is not simply a code for transmitting information. It is the medium in which people think, feel, argue, plan, and make sense of their lives.
Consider the idea of jugaad—finding creative and practical solutions using whatever resources are available, a way of solving problems that is common in rural India.6 There is no single English word that fully captures its meaning. Similarly, decisions about land or water in many villages are often made collectively, based on relationships, trust, and shared responsibilities that no Western legal model cleanly captures.
When an AI model built primarily on English-language, urban, Western text tries to assist someone embedded in rural India, it is not just translating words. It is translating — imperfectly, incompletely — between two entirely different ways of being in the world. The model does not know what it does not know.
6 · What Needs to Change
This is not an argument against AI. It is an argument for AI that actually works for everyone — which requires acknowledging that "everyone" currently means something far narrower than it should.
-
01
Data collection must go into communities
The next wave of useful language data will not come from scraping the internet harder. It requires going into villages, working with community organisations, and building infrastructure for people to contribute their voices — and be compensated. Karya's model is one template.
-
02
Dialect must be treated as data, not noise
The instinct in most ML pipelines is to normalise linguistic variation. For AI to serve rural India, that instinct must be reversed. Variation is the reality. The model needs to learn it, not smooth it away.
-
03
Voice must be the primary interface
For hundreds of millions of Indians, voice is the native digital interface. AI tools built for rural contexts should be voice-first by design, with text as the fallback — not the other way around.
-
04
Language infrastructure must be treated as public infrastructure
Bhashini is a start. But India's linguistic diversity requires something closer to what the government did for roads or electricity — a long-term, adequately funded, publicly accountable commitment to building the connective tissue.
When we talk about the digital divide, we usually mean access — who has a phone, who has internet, who can afford data. But there is a second, quieter divide running alongside the first: the gap between the language AI speaks and the language most of India actually uses.
Ramkhelawan gave up on the chatbot. He went back to calling his cousin in the city when he had questions. That works for him, most of the time. But it is not what the technology promised. And for the millions without a cousin in the city, it is not enough.
AI has the potential to become one of the most powerful technologies for reducing inequality in this century. But that promise can only be fulfilled if those building AI treat linguistic inclusion not as a simple diversity goal, but as a core part of designing these systems from the very beginning.
The language gap is not a small problem at the edges of AI development. It is a large problem at the centre of it — we have just been looking away.
References
- W3Techs. "Usage Statistics of Content Languages for Websites." Web Technology Surveys, February 2026. w3techs.com/technologies/overview/content_language
-
World Data. "Dutch Language Statistics." Worlddata.info, 2025 — worlddata.info/languages/dutch.php
Census of India 2011 / UNT Digital Library. "Bhojpuri Language Resource." — digital.library.unt.edu - Grierson, G.A. Linguistic Survey of India, Vol. V, Part II. Government of India, 1903 (foundational survey of Bihari and Hindi-belt dialects, incl. Bhojpuri, Awadhi, and related registers).
- Internet and Mobile Association of India (IAMAI) & Kantar. "Internet in India Report 2024." January 2025. ibef.org — coverage of the IAMAI–Kantar report
-
Hutiri, W. T. et al. "Does ChatGPT and Whisper Make Humanoid Robots More Relatable?" arXiv preprint, 2024.
arxiv.org/pdf/2402.07095
Tripathi, K., Gothi, R. & Wasnik, P. "Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and Tokenization." Sony Research India, arXiv, 2024. arxiv.org/pdf/2412.19785 - Radjou, N., Prabhu, J. & Ahuja, S. Jugaad Innovation: Think Frugal, Be Flexible, Generate Breakthrough Growth. Jossey-Bass, 2012.