GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
A Comprehensive Survey on AI Hallucination Detection and Trust Evaluation in Large Language Models
Authors
Nithin S, Tejaswi K
Abstract
Today, applications are powered by large language models (LLMs) like Claude, GPT-4, LLaMA and Gemini in healthcare, education, law, finance and science. They also continue to hallucinate, making things up, giving out 'garbage data' or just being completely wrong. This is a reduction in user trust and difficult for safe deployment. This survey summarizes the research on hallucination detection and trust evaluation in the period of 2020 to 2026. We propose a five-class hierarchical taxonomy of hallucination, classify over eighty detection methods into four paradigms (internal, external, hybrid, and human-in-the-loop) and establish a trust evaluation framework with eight dimensions. In addition, over 15 benchmarks are explored, benchmarks metrics are compared, and nine open challenges and ten directions for future work are presented. Lastly, we sketch the Hallucination-Aware Trust Architecture (HATA), a modular pipeline consisting of detection, verification, scoring and explainability for more trustworthy deployment.
Pages:
5085 - 5094