Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 12 (2026), Issue 2

Enhancing Medical Visual Question and Answering through Domain-Specific LLMs & ILA Framework

Authors

Atla Hruthvika, K. Radhika

Abstract

The Medical Visual Question Answering (MVQA) has become a research necessity in the creation of intelligent clinical decision-support systems, but the vast majority of current multimodal AI systems are black-box opaque models whose diagnostic results are fluent but cannot be factually verified. These ungrounded or hallucinated reactions are dangerous and unsafe in high risk health care environments. The Image-to-Label-to-Answer (ILA) framework proposed in this paper is a formalizable, auditable MVQA system which architecturally implements the decoupling between visual measurement and linguistic reasoning. It is based on the quantitative segmentation of Swin-UNETR and structured label generation and persistence, a two-stage domain-specific LLM pipeline (FLAN-T5 + BioMistral) and a safety-critical Factual Audit mechanism. The ILA framework, assessed on five multi-modal clinical datasets, obtains a mean DSC of 0.921 in segmenting organs, a 100 percent Factual Consistency Rate post-audit and a 96.4 percent end-to-end diagnostic accuracy across all modalities, indicating that structured evidence grounding can deliver accurate data, remove all hallucination as well as offer full diagnostic traceability.