The rapid growth of Edge AI is driving the adoption of on-device accelerator systems to deliver high throughput under a limited power budget, making energy efficiency a paramount constraint. Moreover, the increasing heterogeneity of System-on-Chip (SoC) motivates the need for standardized and reproducible AI evaluation methodologies for fair system comparison, such as the industry-standard Linux-centered MLPerf. This work focuses on AMD's Ryzen AI SoC as a representative system integrating a Neural Processing Unit (NPU), an integrated GPU (iGPU), and a CPU, with a Windows-based software stack. First, we bridge the MLPerf inference framework with the Ryzen AI platform by introducing a dedicated bridging layer. Using the MLPerf Inference benchmark, we conduct a comprehensive comparative analysis of NPU, iGPU, and CPU across standardized image classification workloads. However, the lack of, or limited availability of, fine-grained power telemetry hinders in-depth power and performance-per-watt analysis. Therefore, this work proposes a power-centric benchmarking methodology to provide a deeper understanding of the NPU power contributions to the SoC. Overall, the NPU consistently outperforms both the CPU and the iGPU in both performance and energy efficiency. When maximizing performance, the NPU achieves 12.2× and 2.1× speedups over CPU and iGPU, respectively. Instead, when targeting energy efficiency, the NPU delivers 44.3× and 6.8× improvements over CPU and iGPU, respectively. Finally, under tighter latency constraints, CPU-GPU throughput drops by 62.5% when moving from a 100 ms to a 50 ms server latency constraint, whereas CPU-NPU degrades by only 13.9 %, enabling predictable latency-bound inference at the edge.
A Power-Centric Methodology to Characterize Edge AI SoC with Limited Telemetry Capabilities
Mantovi, Giulio;Paltrinieri, Davide;Sorrentino, Giuseppe;Pilato, Christian;Conficconi, Davide
2026-01-01
Abstract
The rapid growth of Edge AI is driving the adoption of on-device accelerator systems to deliver high throughput under a limited power budget, making energy efficiency a paramount constraint. Moreover, the increasing heterogeneity of System-on-Chip (SoC) motivates the need for standardized and reproducible AI evaluation methodologies for fair system comparison, such as the industry-standard Linux-centered MLPerf. This work focuses on AMD's Ryzen AI SoC as a representative system integrating a Neural Processing Unit (NPU), an integrated GPU (iGPU), and a CPU, with a Windows-based software stack. First, we bridge the MLPerf inference framework with the Ryzen AI platform by introducing a dedicated bridging layer. Using the MLPerf Inference benchmark, we conduct a comprehensive comparative analysis of NPU, iGPU, and CPU across standardized image classification workloads. However, the lack of, or limited availability of, fine-grained power telemetry hinders in-depth power and performance-per-watt analysis. Therefore, this work proposes a power-centric benchmarking methodology to provide a deeper understanding of the NPU power contributions to the SoC. Overall, the NPU consistently outperforms both the CPU and the iGPU in both performance and energy efficiency. When maximizing performance, the NPU achieves 12.2× and 2.1× speedups over CPU and iGPU, respectively. Instead, when targeting energy efficiency, the NPU delivers 44.3× and 6.8× improvements over CPU and iGPU, respectively. Finally, under tighter latency constraints, CPU-GPU throughput drops by 62.5% when moving from a 100 ms to a 50 ms server latency constraint, whereas CPU-NPU degrades by only 13.9 %, enabling predictable latency-bound inference at the edge.| File | Dimensione | Formato | |
|---|---|---|---|
|
mlperfnpu_RAW.pdf
accesso aperto
Dimensione
589.61 kB
Formato
Adobe PDF
|
589.61 kB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



