How to Reach Real-Time AI on Consumer Devices? Solutions for Programmable and Custom Architectures

Venieris, Stylianos I.; Panopoulos, Ioannis; Leontiadis, Ilias; Venieris, Iakovos S.

Computer Science > Machine Learning

arXiv:2106.15021 (cs)

[Submitted on 21 Jun 2021]

Title:How to Reach Real-Time AI on Consumer Devices? Solutions for Programmable and Custom Architectures

Authors:Stylianos I. Venieris, Ioannis Panopoulos, Ilias Leontiadis, Iakovos S. Venieris

View PDF

Abstract:The unprecedented performance of deep neural networks (DNNs) has led to large strides in various Artificial Intelligence (AI) inference tasks, such as object and speech recognition. Nevertheless, deploying such AI models across commodity devices faces significant challenges: large computational cost, multiple performance objectives, hardware heterogeneity and a common need for high accuracy, together pose critical problems to the deployment of DNNs across the various embedded and mobile devices in the wild. As such, we have yet to witness the mainstream usage of state-of-the-art deep learning algorithms across consumer devices. In this paper, we provide preliminary answers to this potentially game-changing question by presenting an array of design techniques for efficient AI systems. We start by examining the major roadblocks when targeting both programmable processors and custom accelerators. Then, we present diverse methods for achieving real-time performance following a cross-stack approach. These span model-, system- and hardware-level techniques, and their combination. Our findings provide illustrative examples of AI systems that do not overburden mobile hardware, while also indicating how they can improve inference accuracy. Moreover, we showcase how custom ASIC- and FPGA-based accelerators can be an enabling factor for next-generation AI applications, such as multi-DNN systems. Collectively, these results highlight the critical need for further exploration as to how the various cross-stack solutions can be best combined in order to bring the latest advances in deep learning close to users, in a robust and efficient manner.

Comments:	Invited paper at the 32nd IEEE International Conference on Application-Specific Systems, Architectures and Processors (ASAP), 2021
Subjects:	Machine Learning (cs.LG); Hardware Architecture (cs.AR); Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2106.15021 [cs.LG]
	(or arXiv:2106.15021v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2106.15021

Submission history

From: Stylianos Venieris [view email]
[v1] Mon, 21 Jun 2021 11:23:12 UTC (945 KB)

Computer Science > Machine Learning

Title:How to Reach Real-Time AI on Consumer Devices? Solutions for Programmable and Custom Architectures

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:How to Reach Real-Time AI on Consumer Devices? Solutions for Programmable and Custom Architectures

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators