Optimizing CNN Inference Speed over Big Social Data through Efficient Model Parallelism for Sustainable Web of Things

Home > Research > Publications & Outputs > Optimizing CNN Inference Speed over Big Social ...

Computing and Communications

Associated organisational unit

Insight

Electronic data

Optimizing CNN inference speed over big social data through efficient model parallelism for sustainable web of things-final
Accepted author manuscript, 3.12 MB, PDF document
Available under license: CC BY: Creative Commons Attribution 4.0 International License

Text available via DOI:

https://doi.org/10.1016/j.jpdc.2024.104927
Final published version

View graph of relations

Research output: Contribution to Journal/Magazine › Journal article › peer-review

Published

Yuhao Hu
Xiaolong Xu
Muhammad Bilal
Weiyi Zhong
Yuwen Liu
Huaizhen Kou
Lingzhen Kong

More...

Article number	104927
<mark>Journal publication date</mark>	31/10/2024
<mark>Journal</mark>	Journal of Parallel and Distributed Computing
Volume	192
Publication Status	Published
Early online date	8/06/24
<mark>Original language</mark>	English

Abstract

The rapid development of artificial intelligence and networking technologies has catalyzed the popularity of intelligent services based on deep learning in recent years, which in turn fosters the advancement of Web of Things (WoT). Big social data (BSD) plays an important role during the processing of intelligent services in WoT. However, intelligent BSD services are computationally intensive and require ultra-low latency. End or edge devices with limited computing power cannot realize the extremely low response latency of those services. Distributed inference of deep neural networks (DNNs) on various devices is considered a feasible solution by allocating the computing load of a DNN to several devices. In this work, an efficient model parallelism method that couples convolution layer (Conv) split with resource allocation is proposed. First, given a random computing resource allocation strategy, the Conv split decision is made through a mathematical analysis method to realize the parallel inference of convolutional neural networks (CNNs). Next, Deep Reinforcement Learning is used to get the optimal computing resource allocation strategy to maximize the resource utilization rate and minimize the CNN inference latency. Finally, simulation results show that our approach performs better than the baselines and is applicable for BSD services in WoT with a high workload.

Research

Associated organisational unit

Electronic data

Links

Text available via DOI:

Optimizing CNN Inference Speed over Big Social Data through Efficient Model Parallelism for Sustainable Web of Things

Abstract

Quick Links

Connect With Us

Faculties & Depts

Contact Us