IIT Madras-Backed Bodhan AI Launches Four Open Indian-Language AI Models for Education
Bodhan AI, a Centre of Excellence in AI for Education incubated at IIT Madras, on Friday launched four foundational artificial intelligence models built specifically for Indian languages, marking a fresh push toward homegrown, multilingual AI infrastructure for the country's education system.
Four Models, One Sovereign Stack
Developed in partnership with AI4Bharat, the models span automatic speech recognition, text-to-speech, machine translation and optical character recognition. The speech recognition model supports 27 languages, while the OCR and text-to-speech models each cover 23 languages; the translation model, called Bodhan-Translate, supports 22 Indian languages. The models are being released as open-weight digital public goods, with hosted APIs made available on what the organisation describes as sovereign digital public infrastructure.
The initiative sits within the broader Bharat EduAI Stack, an effort to create shared, reusable AI infrastructure for education rather than having individual institutions or companies build overlapping foundational models from scratch. Officials at IIT Madras framed the launch as a step toward AI systems that genuinely understand India's linguistic diversity, rather than simply adapting global models for Indian use.
Built for Classrooms, Not Just Labs
Alongside the four models, Bodhan AI introduced a Student TutorBot aimed at learners in classes six through twelve, built around NCERT and state board curricula, and a Teacher Assistant Bot designed to help educators plan lessons, generate worksheets and assessments, and evaluate student work. Students can interact with the tutor bot via text or voice in 22 Indian languages, with the system using their own textbook content to generate explanations and guided problem-solving support.
The teacher-facing tool is designed to keep educators in control, with every AI-generated output positioned as a draft for review, editing or rejection rather than an automated classroom decision. The models were trained and fine-tuned using NVIDIA's Nemotron open models and the NeMo framework, with Bodhan AI adapting Nemotron's speech recognition capabilities specifically for Indian regional dialects and accents.
Why It Matters for India's AI Ambitions
The release reflects a broader trend of Indian institutions attempting to build sovereign AI capacity rather than relying solely on foreign-trained models poorly suited to India's linguistic complexity, which spans dozens of major languages and hundreds of dialects. By offering the models as open-weight releases and APIs, Bodhan AI is positioning itself as infrastructure other developers, edtech startups, researchers and government bodies can build upon, rather than as a standalone consumer product.
Officials involved in the project said data privacy and responsible deployment protocols, including anonymisation and compliance with national education-data frameworks, are built into the architecture from the outset. The organisation, a Section 8 company funded by the Ministry of Education, said further collaboration with NVIDIA on datasets and training recipes is planned as it works toward additional foundational models for Indian languages.
For an education system serving hundreds of millions of students across dramatically varied linguistic contexts, tools capable of operating fluently across Indian languages could meaningfully widen access to AI-assisted learning, particularly for students and teachers who are more comfortable working outside English.
This is an original summary based on public reporting. See our editorial policy for how we source, write, and correct our stories.
Comments
Sign in to join the discussion.