Asset Details

MbrlCatalogueTitleDetail

Do you wish to reserve the book?

A generic deep learning architecture optimization method for edge device based on start-up latency reduction

by Meng, Lin , Li, Qi , Li, Hengyi

in Algorithms / Artificial intelligence / Central processing units / Computer architecture / Computer Graphics / Computer Science / Constraints / CPUs / Data processing / Deep learning / Devices / Image Processing and Computer Vision / Machine learning / Methods / Multimedia Information Systems / Network latency / Optimization / Pattern Recognition / Real time / Signal,Image and Speech Processing / Sparsity

2024

Yes Please

Hey, we have placed the reservation for you!

By the way, why not check out events that you can attend while you pick your title.

Oops! Something went wrong.

Looks like we were not able to place the reservation. Kindly try again later.

Are you sure you want to remove the book from the shelf?

A generic deep learning architecture optimization method for edge device based on start-up latency reduction

by Meng, Lin , Li, Qi , Li, Hengyi

2024

Confirm

Do you wish to request the book?

A generic deep learning architecture optimization method for edge device based on start-up latency reduction

by Meng, Lin , Li, Qi , Li, Hengyi

2024

Please be aware that the book you have requested cannot be checked out. If you would like to checkout this book, you can reserve another copy

How would you like to get it?

Submit

We have requested the book for you!

Your request is successful and it will be processed during the Library working hours. Please check the status of your request in My Requests.

Oops! Something went wrong.

Looks like we were not able to place your request. Kindly try again later.

Journal Article

A generic deep learning architecture optimization method for edge device based on start-up latency reduction

Meng, Lin,

Li, Qi,

Li, Hengyi

2024

Overview

In the promising Artificial Intelligence of Things technology, deep learning algorithms are implemented on edge devices to process data locally. However, high-performance deep learning algorithms are accompanied by increased computation and parameter storage costs, leading to difficulties in implementing huge deep learning algorithms on memory and power constrained edge devices, such as smartphones and drones. Thus various compression methods are proposed, such as channel pruning. According to the analysis of low-level operations on edge devices, existing channel pruning methods have limited effect on latency optimization. Due to data processing operations, the pruned residual blocks still result in significant latency, which hinders real-time processing of CNNs on edge devices. Hence, we propose a generic deep learning architecture optimization method to achieve further acceleration on edge devices. The network is optimized in two stages, Global Constraint and Start-up Latency Reduction, and pruning of both channels and residual blocks is achieved. Optimized networks are evaluated on desktop CPU, FPGA, ARM CPU, and PULP platforms. The experimental results show that the latency is reduced by up to 70.40%, which is 13.63% higher than only applying channel pruning and achieving real-time processing in the edge device.

Share this book

Add to My Shelf

Publisher

Springer Berlin Heidelberg,Springer Nature B.V

Subject

Algorithms

/ Artificial intelligence

/ Central processing units

/ Computer architecture

/ Computer Graphics

/ Computer Science

/ Constraints