International Journal of Computer Science and Technology

ISSN 2996-8223

International Journal of Computer Science and Technology | Vol. 2, No. 3, March 2021 | pp. 17–24

Research Article

Title: Optimizing Deep Neural Network Inference on Heterogeneous Edge Devices Using Adaptive Model Quantization

Names of Authors: Chen Wei¹, Li Na², and Zhang Wei³

Authors’ Affiliations:
¹Department of Computer Science and Technology, Tsinghua University, Beijing, China
²State Key Laboratory of Computer Architecture, Institute of Computing Technology, CAS, Beijing, China
³School of Electronic and Information Engineering, Xi'an Jiaotong University, Xi'an, China

Abstract: Deploying state-of-the-art deep neural networks on edge devices enables real-time, low-latency computer vision and natural language applications. However, modern models possess millions of parameters, causing severe memory footprint and power consumption challenges on resource-constrained embedded hardware. Model compression via uniform quantization reduces precision from 32-bit floating-point (FP32) to 8-bit integers (INT8), but often incurs unacceptable accuracy drops in complex network topologies. This study proposes an adaptive, mixed-precision model quantization framework designed to optimize deep learning inference on heterogeneous edge processors. The system uses a layer-wise sensitivity analysis based on Hessian trace estimation to determine the structural tolerance of individual network layers to precision loss. Layers displaying high sensitivity maintain 8-bit or 16-bit precisions, while noise-tolerant components compress to ultra-low 4-bit or 2-bit formats. A multi-objective optimization engine balances model accuracy, execution latency, and memory consumption across target hardware architectures. Hardware experiments on NVIDIA Jetson Nano and Raspberry Pi 4 platforms show that the mixed-precision framework reduces model memory size by 73.5% and accelerates inference execution speed by 2.4 times compared to baseline FP32 implementations. Crucially, top-1 accuracy degradation on the ImageNet dataset is restricted to less than 0.65%, making it highly effective for real-time edge intelligence.

Keywords: Deep learning inference, Edge computing, Model quantization, Mixed-precision, Hardware acceleration, Computer vision

Manuscript Timeline: Received: December 05, 2020; Revised: January 18, 2021; Accepted: February 15, 2021; Published: March 1, 2021

Citation: Wei, C., Na, L., & Wei, Z. (2021). Optimizing deep neural network inference on heterogeneous edge devices using adaptive model quantization. International Journal of Computer Science and Technology, 2(3), 17–24. DOI: 10.46882/2021/IJCST/000015