GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
Voicify - A Multilingual and Expressive Text to Speech System
Authors
Sandip Shinde, Anushka Dabhade, Nirwani Adhau, Madhura Chitupe, Niharika Deshmukh
Abstract
The human voice is essential for storytelling, but creating speech that feels real and emotional instead of robotic is a significant challenge today. Many current Text-to-Speech (TTS) systems face an "uncanny valley" effect. They produce monotone voices that struggle with the details of Indic languages and fail to handle structured numbers accurately. Most existing models choose between being expressive and remaining stable. When processing long texts this often leads to issues like audio glitches or system crashes. This paper introduces Voicify, a new TTS framework developed for RPM Studios to solve these problems and improve broadcast media synthesis. Unlike standard single-engine systems, Voicify employs an Intelligent Routing mechanism. This mechanism directs English narrative text to a Parler-TTS engine for rich emotional expression while routing Hindi, Marathi, and structured data to gTTS for accurate language processing. This two-path system is supported by a real-time RAM monitor that keeps everything stable during large document conversions. By combining smart text normalization with efficient resource management, Voicify offers an audio pipeline that sounds more human-like and outperforms current solutions in both expressiveness and reliability. It is specifically designed to meet the diverse needs of the Indian broadcast industry.
Pages:
4197 - 4203