지속적 학습 벤치마크: 실제 환경에서의 상태 기반 시스템 평가
Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments
지속적 학습은 AI 시스템이 순차적인 경험을 통해 성능을 향상시키는 능력으로, 많은 관심을 받고 있지만 이를 평가할 수 있는 고품질의 벤치마크는 존재하지 않습니다. 본 연구에서는 LLM 기반 시스템이 실제로 경험을 통해 개선되는지를 측정하기 위해 설계된 최초의 난이도 높은 전문가 검증 벤치마크인 '지속적 학습 벤치마크(CL-Bench)'를 소개합니다. CL-Bench는 소프트웨어 공학, 신호 처리, 질병 발생 예측, 데이터베이스 쿼리, 전략 게임, 수요 예측 등 6가지 다양한 분야를 포괄하며, 각 분야는 전문가의 검증을 거쳐 설계되었으며, 상태 기반 시스템은 온라인으로 학습할 수 있는 잠재적 구조(코드 베이스 레이아웃, 질병 발생 동역학, 상대방 전략)를 발견할 수 있지만, 무상태 시스템은 이를 파악하기 어렵도록 구성되었습니다. 본 연구에서는 다양한 에이전트 아키텍처(naive in-context learning (ICL)부터 전용 메모리 시스템까지)에 걸쳐 최첨단 모델을 평가하고, 학습 효과를 기존 능력과 분리하기 위해 '수익(gain)' 지표를 도입했습니다. 분석 결과, 이러한 시스템은 지속적 학습 성능 향상을 위한 여지가 있음을 보여줍니다. 에이전트는 종종 즉각적인 관찰에 과적합되거나, 다양한 인스턴스에서 지식을 재사용하지 못하는 경우가 있으며, 전용 메모리 시스템으로도 이러한 문제가 해결되지 않습니다. 오히려 naive ICL 방식이 메모리 관리 시스템보다 더 나은 성능을 보이는 경우도 있었습니다. CL-Bench는 전문가 검증된 작업과 함께 실제 환경의 다양한 분야에서 지속적 학습을 평가하고, 모델 자체의 능력과 온라인 학습 효과를 분리하는 최초의 벤치마크이며, 이는 더 나은 지속적 학습 시스템 개발의 필요성을 보여줍니다.
Continual learning, the ability of AI systems to improve through sequential experience, has attracted substantial interest, but no high-quality benchmark exists to evaluate it. We introduce Continual Learning Bench (CL-Bench), the first difficult, expert-validated benchmark designed to measure whether LLM-based systems genuinely improve with experience. CL-Bench spans six diverse domains (software engineering, signal processing, disease outbreak forecasting, database querying, strategic game-playing, and demand forecasting), each validated by domain experts and designed so that tasks share a learnable latent structure (codebase layout, disease outbreak dynamics, opponent strategies) that a stateful system can discover online but a stateless one cannot. We evaluate frontier models across several agent architectures, from naive in-context learning (ICL) to dedicated memory systems, introducing a gain metric to isolate learning from prior capabilities. We find that these systems leave headroom for improved continual learning: agents frequently overfit to immediate observations or fail to reuse knowledge across instances, and dedicated memory systems do not fix this -- in fact, naive ICL outperforms systems dedicated to memory management. CL-Bench is the first benchmark to evaluate continual learning across diverse real-world domains with expert-validated tasks and isolate online learning from underlying model capability, showing a need for better continual learning systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.