Abstract
In this paper, we present a Network-on-Chip-centric (NoC-centric) design technique for edge AI accelerator architectures. The technique enables NoC with compute capability, eliminates the need of using extra cores or frequent access to memory for neural network computing and eventually improves accuracy and energy efficiency. We demonstrate the dedicated, software-configurable NoCs with on-device inference and training tasks and show the architectures can achieve 2.8x and 2.1x lower energy per classification compared to state-of-the-art baselines, respectively.