2608.05573v1 Aug 06, 2026 cs.AI

SkillTV-Bench: 기술 기반 에이전트 실행의 성능 평가

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

Zihan Guo
Zihan Guo
Citations: 42
h-index: 2
Zhien Han
Zhien Han
Citations: 2
h-index: 1
C. Zeng
C. Zeng
Citations: 4
h-index: 1
Mingyu Zhou
Mingyu Zhou
Citations: 0
h-index: 0
Yang Li
Yang Li
Citations: 0
h-index: 0
Liuhaichen Yang
Liuhaichen Yang
Citations: 0
h-index: 0

LLM(Large Language Model) 에이전트는 도구 사용 및 환경 상호 작용을 통해 점차 장기적인 작업을 수행하고 있으며, 이러한 과정에서 평가는 최종 답변의 정확도를 측정하는 것에서 전체 실행 과정을 검증하는 것으로 변화하고 있습니다. 특히 기술 기반 에이전트의 경우, 작업 시간 동안 습득되는 절차적 지식을 고려해야 합니다. 이 지식은 어떤 증거를 살펴야 하는지, 그리고 어떤 실패가 중요한 문제인지 판단하는 데 필수적입니다. 그러나 기존 평가 벤치마크는 종종 최종 답변이나 정적인 실행 경로만을 제시하며, 작업 시간 동안의 기술과 직접 검토할 수 있는 결과물 및 환경을 결합하는 경우는 드뭅니다. 이에 SkillTV-Bench를 제안합니다. SkillTV-Bench는 50개의 작업에 걸쳐 11개 분야에서 얻은 실제 에이전트 실행 경로 681건으로 구성된 벤치마크이며, LLM 기반 평가 및 에이전트 기반 평가 방법을 모두 사용하여 기술 인지 실행 경로 검증을 평가하는 데 사용됩니다. 또한, SkillTV-Evolve를 제안합니다. 이는 검증 지식을 재사용 가능한 JudgeSkill로 외부화하여, 에이전트 평가자가 목표 검사를 계획하고 증거에 근거한 판단을 내릴 수 있도록 돕습니다. 별도의 개발 데이터셋에서 자동화된 진화 루프를 통해 잘못 판단된 사례들을 사용하여 JudgeSkill을 지속적으로 개선합니다. SkillTV-Bench에서 개선된 JudgeSkill은 동일한 에이전트 평가기의 정확도를 14.8% 향상시켰습니다. 또한, 오프라인 롤아웃 풀 선택 시, 롤아웃 횟수가 1회일 때 성공적인 실행 경로를 22.9% 선택하는 반면, 롤아웃 횟수가 10회일 때는 45.5%로 선택하는 성능을 보입니다. 코드와 데이터는 https://github.com/HanZhi306/SkillTV-Bench 에서 확인할 수 있습니다.

Original Abstract

LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally requires the procedural knowledge encoded in task-time skills, because this knowledge indicates what evidence to inspect and which failures are task-critical. However, existing judge benchmarks often expose final responses or static trajectories, and rarely combine task-time skills with directly inspectable artifacts and environments. We therefore introduce SkillTV-Bench, a 681-case benchmark of real agent trajectories from 50 tasks across eleven domains, designed to evaluate skill-aware trajectory verification for both LLM-as-a-Judge and Agent-as-a-Judge methods. Additionally, we propose SkillTV-Evolve, which externalizes verification knowledge as a reusable JudgeSkill that guides an agent judge to plan targeted inspections and issue evidence-grounded verdicts. On a disjoint development pool, an automated evolution loop further refines the JudgeSkill using misjudged cases. On SkillTV-Bench, the refined skill increases the same agent judge's accuracy by 14.8 percentage points. In offline rollout-pool selection, it increases selected-trajectory success from 22.9% with one rollout to 45.5% with ten rollouts. The code and data are available at https://github.com/HanZhi306/SkillTV-Bench

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!