GRENZE International Journal of Engineering and Technology
Vol. 11
(2025), Issue 2
A Robust Deep Learning-based Feature Extraction and Matching Technique for Visual Odometry in VSLAM
Authors
Madhu Sudan M P, Sudarshan Patil Kulkarni, Aprameya C V
Abstract
Simultaneous Localization and Mapping or SLAM is a computational process that concurrently estimates an agent's location and builds a map of its unknown environment, utilizing approaches such as Light Detection and Ranging (LiDAR SLAM) and Visual SLAM (VSLAM), among which the VSLAM technique per-forms better for indoor localisation and mapping. Traditional VSLAM techniques like Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), and Oriented FAST and Rotated BRIEF (ORB) use geometric methods for feature extraction and matching. While these techniques are effective in static scenarios, they often struggle in dynamic or visually ambiguous environments due to their limitations in feature extraction and motion estimation. The current work addresses these challenges by incorporating deep learning techniques to improve VSLAM’s reliability and adaptability in dynamic environment. Our approach leverages deep learning-based feature extractors and motion estimation models, designed to enhance the system's capability to interpret complex environments, thereby resulting in more robust and precise feature extraction and matching. The deep learning-based model is designed using a custom-enhanced Densely-Connected Two- Stream Network Model (D2-Net Model) with regularisation and attention mechanisms. The conventional D2-Net architecture has been impro-vised by incorporating additional convolution and dropout layers for feature ex-traction and dilated convolution layers with optimisation blocks to enhance feature discrimination, which makes the model more resilient to scene changes and viewpoint variations, especially for dynamic scenarios. Compared to traditional models, the developed architecture consumed lesser time for feature computation per image and better overall execution time for complete dataset. The designed model has been tested on the TUM-RGB-D fr1/xyz dataset, for which the model achieved 602 good feature matches per image pair with a high inlier ratio of 97.76% based on Random Sample Consensus (RANSAC) matching criteria, which showed strong consistency for the features extracted across all frames.
Pages:
1395 - 1401