Recent advances have enabled deep object detection on low-power edge devices, such as smart glasses, yet these systems remain static after deployment, unable to learn new objects in the field. We present Cerberus, an on-device continuous learning architecture tailored for resource-constrained processors, enabling efficient model updates without full model retraining. Cerberus employs a lightweight multi-head structure, where each head specializes in a single object class, mitigating catastrophic forgetting while maintaining minimal computational and memory overhead. Evaluated on the PAS-CAL Visual Object Classes (VOC) benchmark, our method demonstrates superior performance in data-scarce scenarios, matching or exceeding unconstrained off-device baselines even when trained with only 1% of available labeled data. For example, when learning 5 new classes incrementally after initial training on 15 classes, Cerberus achieves 48.6 [email protected] compared to 29.2 for a replay-based baseline and 34.8 for a fine-tuning baseline, while reducing sample storage requirements by over 20×. We demonstrate end-to-end deployment on a novel low-power multi-core RISC-V processor, achieving an inference latency of only 66.14ms with an energy cost of 5.49mJ. Notably, learning a new class requires only 9 KB of additional memory, and just 29.65mJ per training update, showcasing the feasibility of sustained on-device adaptive perception in battery-powered edge systems, including wearable applications, with an estimated battery life of approximately 22 hours.
Cerberus: Advancing Continuous On-Device Learning with Multi-Head Object Detection
Shalby, Hazem Hesham Yousef;Roveri, Manuel
2026-01-01
Abstract
Recent advances have enabled deep object detection on low-power edge devices, such as smart glasses, yet these systems remain static after deployment, unable to learn new objects in the field. We present Cerberus, an on-device continuous learning architecture tailored for resource-constrained processors, enabling efficient model updates without full model retraining. Cerberus employs a lightweight multi-head structure, where each head specializes in a single object class, mitigating catastrophic forgetting while maintaining minimal computational and memory overhead. Evaluated on the PAS-CAL Visual Object Classes (VOC) benchmark, our method demonstrates superior performance in data-scarce scenarios, matching or exceeding unconstrained off-device baselines even when trained with only 1% of available labeled data. For example, when learning 5 new classes incrementally after initial training on 15 classes, Cerberus achieves 48.6 [email protected] compared to 29.2 for a replay-based baseline and 34.8 for a fine-tuning baseline, while reducing sample storage requirements by over 20×. We demonstrate end-to-end deployment on a novel low-power multi-core RISC-V processor, achieving an inference latency of only 66.14ms with an energy cost of 5.49mJ. Notably, learning a new class requires only 9 KB of additional memory, and just 29.65mJ per training update, showcasing the feasibility of sustained on-device adaptive perception in battery-powered edge systems, including wearable applications, with an estimated battery life of approximately 22 hours.| File | Dimensione | Formato | |
|---|---|---|---|
|
cerberus_sensors-26.pdf
Accesso riservato
Dimensione
10.57 MB
Formato
Adobe PDF
|
10.57 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



