GRENZE International Journal of Engineering and Technology
Vol. 9
(2023), Issue 1
Simple Methods is All You Need for Medical VQA: An ImageCLEFs Med-VQA Task Methods Review
Authors
Jitender Singh, Surender Singh
Abstract
Visual Question Answering (VQA) is an area of AI where the model takes an image and a free-form natural language question related to the image and generates an answer as the output to the question in natural language. In the medical domain, VQA generates the natural language answer to the given question related to a medical scan such as X-Ray, CT, MRI etc. With the public availability of medical VQA datasets introduced by ImageCLEF in 2018, it has become an active area of research. In this paper, we will discuss all four years of ImageCLEF’s Med-VQA competition datasets and methods used by the participants to understand which methods perform better on the medical VQA problem. Also, we describe the limitations of the ImageCLEF dataset and a simple VQAMixUp technique that can be used with simple architectures such as VGG16 and GRU and performs equivalent to the large ensemble models of the winners of the competition.
Pages:
2292 - 2299