Bridging the linguistic divide: recent developments in machine translation for Indian languages
Abstract
Significant advances have been achieved in machine translation (MT) in recent times, particularly state of the art (SOTA) models for languages like English and Indian having distinct grammatical structures and limited monolingual training data. This paper analyses various recent state-of-the-art variants of large language models (LLMs) and neural machine translation (NMT) for Indian languages in comparison to statistical machine translation (SMT). It tackles key questions, such as idiomatic expressions, morphologically complex grammar or the scarceness of parallel corpora. Furthermore, it studies bytewise BPE, compares translation models in terms of BLEU scores using separate and shared-vocabulary representation with copy actions between the BPE translations, and analyses how multitask learning (Caruana (1997)) and attention mechanisms can contribute to the quality of translation. In summary, it provides directions for future work by suggesting new avenues of research including better curated datasets, more efficient approaches for lowresource languages and culturally aware translations.
Keywords
Byte pair encoding; Indian languages; Large language models; Machine translation; Neural machine translation
Full Text:
PDFDOI: http://doi.org/10.11591/ijict.v15i3.pp1272-1289
Refbacks
- There are currently no refbacks.
Copyright (c) 2026 Institute of Advanced Engineering and Science

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
The International Journal of Informatics and Communication Technology (IJ-ICT)
p-ISSN 2252-8776, e-ISSNĀ 2722-2616
This journal is published by theĀ Intelektual Pustaka Media Utama (IPMU).