CANS: 협력적 자기 학습 신경외과 알고리즘을 통한 다중 사용자 공동 에지 추론 가속화
CANS: Accelerating Multiuser Collaborative Edge Inference via Cooperative Autodidactic NeuroSurgeon
최근, 모바일 엣지 컴퓨팅(MEC) 기반의 협업 심층 신경망(DNN) 추론은 자원 제약적인 모바일 기기에 지능형 서비스를 제공하는 유망한 기술로 부상하고 있습니다. 대표적인 예시가 다중 사용자 공동 에지 추론으로, 각 장치가 독립적으로 DNN 모델을 분할하고 무선 네트워크를 통해 공통 에지 서버에 백엔드 연산을 오프로딩합니다. 그러나, 알려지지 않고 시간에 따라 변하는 시스템 조건, 특히 불안정한 무선 연결 및 다양한 장치 성능으로 인해 각 장치에 대한 최적의 DNN 분할을 결정하는 것은 어렵습니다. 이러한 문제를 해결하기 위해, 우리는 Cooperative Autodidactic NeuroSurgeon (CANS)라는 협업 에지 추론 프레임워크를 제안합니다. CANS는 장치가 온라인 추론 과정에서 유용한 피드백을 공유하며 최적의 DNN 분할을 적응적으로 학습하도록 지원합니다. 장치 간 이질성을 해결하고 오프라인 추론 경험을 효과적으로 활용하기 위해, 동일 유형의 장치를 그룹화하고 로컬 오프라인 조기 종료 추론 경험을 사용하여 온라인 탐색을 초기화하는 새로운 FedLinUCB-DW 알고리즘을 통합했습니다. 또한, 우리는 FedLinUCB-DW에 대한 이론적 보장을 제공하며 후회(regret) 상한을 도출합니다. 제안된 방법은 시뮬레이션 환경과 하드웨어 프로토타입 시스템 모두에서 검증되었습니다. 실험 결과는 CANS가 최첨단 기준 모델보다 낮은 추론 지연 시간을 달성한다는 것을 보여줍니다. 특히, 두 개의 에지 장치를 사용한 프로토타입 실험에서, 제안된 CANS는 비협력 기준 모델과 비교하여 평균 추론 지연 시간을 최대 50%까지 줄였습니다.
Recently, mobile edge computing (MEC)-enabled collaborative deep neural network (DNN) inference has emerged as a promising approach for delivering intelligent services to resource-constrained mobile devices. A representative scenario is multi-user collaborative edge inference, where distinct devices independently partition their DNN models and offload backend computation to a common edge server over wireless networks. However, determining the optimal DNN partition for each device is challenging due to unknown and time-varying system conditions, including fluctuating wireless links and diverse device capabilities. To address this problem, we propose Cooperative Autodidactic NeuroSurgeon (CANS), a collaborative edge inference framework that enables devices to adaptively learn optimal DNN partitions by sharing informative feedback during online inference. To handle the challenge of device heterogeneity and better leverage offline inference experience, we integrate a novel FedLinUCB-DW algorithm that groups devices of the same type and warm-starts online exploration using local offline early-exit inference experience. Furthermore, we provide theoretical guarantees for FedLinUCB-DW by deriving the regret upper bound. We also validate our method on both a simulated environment and a hardware prototype system. Empirical evaluations demonstrate that CANS achieves lower inference latency compared to state-of-the-art baselines. Especially, in prototype experiments on two edge devices, the proposed CANS reduced average inference latency by up to 50% compared to the non-cooperative baseline.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.