GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
Personalized Voice Cloning and Synthesis: Enhancing Human Computer Interaction through Speech Replication
Authors
P. Changamma, Bhavana Yaparla, Bharathi Mareddy, Gowtham Naik Mude, Bala Narasimhulu Korru, K. Keerthi Naidu
Abstract
This paper describes a real time voice cloning system that combines with Groq’s real time orchestration, Whisper ASR, zero shot text to speech and speech to speech synthesis and HiFi GAN neural vocoders with just 5-30 seconds of enrollment audio, the system achieves 95.6% speaker cloning accuracy, a mean opinion score of 4.7 and end to end latency under 900 milliseconds. Not like diffusion based models such as MegaTTs and NaturalSpeech 2, this united pipeline removes the need for a separate enrollment phase while maintaining prosody and emotion with 89.9% accuracy. Experimental results show interface speeds 2.8 times faster than current state-of-the-art systems with comparable quality, supporting real time applications in assistive communication, conversational AI , and accessibility.
Pages:
1026 - 1034