GRENZE International Journal of Engineering and Technology
Vol. 10
(2024), Issue 2
Q-Net Compressor: Deep Learning Model Compression with Quantization
Authors
Yash Bansal, Sejal Garg, Kartik Saroha, Manu Singh
Abstract
The rapid advancement of deep learning has brought about remarkable achievements in various fields, particularly by utilizing Deep Neural Networks. It offers a high rate of accuracy, extensive parameters, and intensive computational requirements that often characterize these networks. Due to high memory consumption and energy usage, this results in significant challenges when deploying DNNs on hardware-constrained devices. Researchers have extensively explored various compression techniques (Al-Qurabat 2022) to address these challenges, quantifying the merging as a promising solution. However, the extensive use of quantization in previous works necessitates a comprehensive survey that provides a complete study, analysis, and comparison of various quantization approaches. This paper introduces the “Q-Net Compressor,†a novel solution for installing substantial neural (Tao, Hou, et al. 2022) network training on constrained resources edge devices. Utilizing adaptive precision levels to maintain model functionality while reducing memory and computation demands. The Q-Net Compressor, implemented with TensorFlow and PyTorch, offers flexibility across diverse deep-learning architectures, striking a balance between significant model size reduction and the preservation of predictive power. This innovative approach contributes to the evolving deep learning model compression field with quantization, bridging the gap between computational demands and resource constraints. We hope to address this gap in this survey study by providing an integrative report that investigates the ideas and methods of quantization, (Hu and Wen 2021) with a particular emphasis on its use in image classification problems.
Pages:
3150 - 3156