GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
Uncovering Systemic Vulnerabilities in NMT: A Linguistically Motivated Stress Test for English- German Translation
Authors
Priya Goel, Prasanna Dwivedi, Pankaj Goel, Mahek Singhal, Shruti Shukla, Sakshi
Abstract
Neural Machine Translation (NMT) has achieved human-like fluency in machine translations for generic domains in recent years. However, there are notable errors in the accuracy and progress when we discuss the need to process linguistic, deeply structured inputs in more than just English. This paper provides a comparative analysis of two popular NMT architectures, DeepL and Google Translate, for the English-German Language Pair in order to account for the observed errors. This study uses a manually curated dataset that includes extreme scenarios appropriate for target-based stress testing, rather than using a premade dataset from another source. Idiomatic expressions, structural stress, lexical density, and literary tone are the four main linguistic categories in which its performance is evaluated. The COMET metric is utilized to perform machine-based analysis upon the derived data. This analysis highlights that these metrics don’t fail completely but display their own set of strengths and weaknesses. It shows that Google Translate handles intricate syntactic dependencies in a better way alongside the involved recursive structures found in many languages, whereas DeepL is better at understanding cultural nuances and practical meaning. Regarding the drawbacks, both systems struggle to accurately depict authorial voice in conjunction with highly context-dependent lexical compounds. The goal of this study is to systematically tabulate and categorise the errors in order to highlight the limitations of generalised evaluation, even though it only discusses a single language pair. As our contribution, we contend that the development of strong, reliable, and long-lasting multilingual AI systems across languages in the multilingual domain requires focused, linguistically motivated stress testing.
Pages:
1733 - 1740