AgentHPOBench: LLM 에이전트를 순차적인 하이퍼파라미터 최적화기로 평가하기 위한 벤치마크
AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers
LLM(대규모 언어 모델)이 코드 자동 완성 시스템에서 자율적인 과학 연구 에이전트로 진화함에 따라, 실험 수행 능력을 평가하는 것이 점점 더 중요해지고 있습니다. 기존의 벤치마크는 주로 정적 코드 생성, 논문 재현 또는 최종 답변의 정확성에 초점을 맞추지만, 에이전트가 실험 결과를 해석하고 이를 바탕으로 후속적인 하이퍼파라미터 결정을 내리는 능력을 직접적으로 평가하지 못합니다. 이러한 격차를 해소하기 위해, 7가지 연구 분야에 걸쳐 30개의 실행 가능한 머신러닝 작업으로 구성된 순차적 벤치마크인 AgentHPOBench를 소개합니다. 각 작업은 검증된 초기 설정으로 시작되며, 이후 에이전트는 여러 단계를 거쳐 순차적인 방식으로 실험을 진행합니다. 각 단계에서 에이전트는 누적된 설정, 지표 및 로그를 관찰한 후 다음 유효한 설정을 제안합니다. 우리는 12개의 널리 사용되는 에이전트와 기존의 HPO(하이퍼파라미터 최적화) 기준을 통일된 프로토콜 하에서 평가했습니다. 결과는 현재 에이전트가 다양한 분야에서 측정 가능한 실험 최적화 능력을 보여주지만, 지속적인 반복 개선, 복잡한 로그 분석 및 보고된 참조 성능에 대한 일관된 발전에는 여전히 명확한 한계가 있음을 나타냅니다.
As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper replication, or final answer correctness, but do not directly assess whether agents can interpret experimental evidence and use it to guide subsequent hyperparameter decisions. To address this gap, we introduce AgentHPOBench, a sequential benchmark comprising 30 executable machine learning tasks across seven research categories. Each task begins with a validated baseline run, after which an agent performs several sequential interventions. At each step, the agent observes the accumulated configurations, metrics, and logs before proposing the next valid configuration. We evaluate 12 widely used agents and conventional HPO baselines under a unified protocol. The results show that current agents exhibit measurable experimental optimization ability across domains, but still face clear limitations in sustained iterative refinement, complex log diagnosis, and consistent progress toward reported reference performance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.