2606.29961v1 Jun 29, 2026 cs.LG

DuoMem: 이중 공간 증류를 통한 강력한 온디바이스 메모리 에이전트 개발

DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation

Ondrej Bohdal
Ondrej Bohdal
Citations: 548
h-index: 11
Mete Ozay
Mete Ozay
Citations: 245
h-index: 8
T. Ceritli
T. Ceritli
Citations: 142
h-index: 6
Peyman Hosseini
Peyman Hosseini
Citations: 35
h-index: 2
Ahmed Alajrami
Ahmed Alajrami
The University of Sheffield
Citations: 44
h-index: 3
Andrea Maracani
Andrea Maracani
Citations: 5
h-index: 1
Ignacio Castro
Ignacio Castro
Citations: 31
h-index: 1
Matthew Purver
Matthew Purver
Citations: 54
h-index: 4
Savas Ozkan
Savas Ozkan
Citations: 1
h-index: 1

대규모 언어 모델(LLM) 기반 에이전트는 여러 단계를 거쳐 환경과 상호 작용하면서 복잡한 절차적 작업을 해결할 수 있지만, 이러한 능력은 일반적으로 대규모 모델, 긴 컨텍스트 및 반복적인 추론 호출에 의존합니다. 이는 리소스가 제한된 장치에 고급 메모리 강화 에이전트를 배포하기 어렵게 만듭니다. 본 논문에서는 DuoMem이라는 이중 공간 증류 프레임워크를 소개하며, 이를 통해 대규모 교사 모델의 절차적 문제 해결 능력을 소형 학생 모델로 이전합니다. DuoMem은 두 가지 상호 보완적인 방식으로 증류를 수행합니다: (1) 컨텍스트 공간 증류는 학생이 생성한 메모리를 제거하고 대신 더 높은 품질의 교사가 생성한 절차적 메모리를 학생 입력 앞에 추가하며, (2) 매개변수 공간 증류는 성공적인 교사 모델의 경로를 기반으로 경량 LoRA 어댑터를 미세 조정합니다. 어려운 로봇 의사 결정 벤치마크인 ALFWorld에서 평가 결과, DuoMem은 40억 개의 파라미터를 가진 모델의 작업 성공률을 4.3%에서 77.9%로 향상시켜 720억 개의 파라미터를 가진 교사 모델(87.1%)과의 격차를 크게 줄였습니다. 또한, DuoMem은 1천만 개 미만의 학습 가능한 파라미터와 몇 메가바이트의 미리 계산된 교사 모델 메모리만을 추가합니다. 더욱이, DuoMem을 적용한 40억 개의 파라미터를 가진 모델은 실제 시간 기준으로 720억 개의 파라미터를 가진 교사 모델보다 3배 이상 빠르게 작업을 완료하므로 실시간 에지 배포에 적합합니다. 20억 개에서 720억 개의 파라미터를 갖는 여덟 가지 모델에 대한 광범위한 실험 결과, 두 가지 증류 방식이 상호 보완적으로 기여한다는 것을 확인했습니다.

Original Abstract

Large Language Model (LLM)-based agents can solve complex procedural tasks by interacting with environments over multiple turns, but this ability typically depends on large models, long contexts, and repeated inference calls. This makes advanced memory-augmented agents difficult to deploy on resource-constrained devices. We introduce DuoMem, a dual-space distillation framework that transfers procedural problem-solving ability from a large teacher model to compact student models. DuoMem distils in two complementary spaces: (1)context-space distillation, which replaces student-generated memories with higher-quality teacher-generated procedural memories prepended to the student's input, and (2)parameter-space distillation, which fine-tunes lightweight LoRA adapters on successful teacher trajectories. Evaluated on ALFWorld, a challenging embodied decision-making benchmark, DuoMem boosts a 4B-parameter model from 4.3% to 77.9% task success rate, closing most of the gap to a 72B teacher model (87.1%), while adding fewer than 10M trainable parameters and only a few megabytes of pre-computed teacher memories. Moreover, the DuoMem-enhanced 4B model completes tasks over 3x faster than the 72B teacher in wall-clock time, making it viable for real-time edge deployment, which would be challenging for the teacher.Extensive ablations across eight models spanning 2B-72B parameters reveal that both distillation axes contribute complementary

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!