오픈 평가 에이전트: 효율적이고 프롬프트 기반의 시각 생성 모델 평가
Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models
최근 시각 생성 모델의 발전으로 고품질 이미지 및 비디오 생성이 가능해졌지만, 이러한 모델을 평가하는 데는 수백 또는 수천 개의 이미지나 비디오를 샘플링해야 하는 경우가 많아 계산 비용이 많이 듭니다. 기존 평가 방법은 또한 경직된 파이프라인에 의존하며, 특정 사용자 요구 사항을 간과하고 명확한 설명 없이 숫자 결과만 제공합니다. 인간이 몇 가지 샘플만으로도 모델의 성능에 대한 인상을 빠르게 형성하는 방식을 모방하여, 우리는 효율적이고 동적인 다단계 평가를 위한 Evaluation Agent 프레임워크를 제안합니다. 이 프레임워크는 자세하고 사용자 맞춤형 분석을 제공합니다. 자연어 기반 평가 요청이 주어지면, 에이전트는 이를 하위 측면으로 분해하고, 대상 프롬프트를 생성하며, 평가 대상 모델에서 이미지 또는 비디오 샘플을 추출하고, 적절한 평가 도구를 호출하며, 관찰된 증거를 바탕으로 계획을 반복적으로 업데이트합니다. 이를 통해 미리 정의된 벤치마크 차원과 사용자 우려 사항 모두를 포괄합니다. 따라서 이 프레임워크는 효율적이고, 프롬프트 기반이며, 설명 가능하고, 모델 및 도구 전반에 걸쳐 확장 가능합니다. 실험 결과, Evaluation Agent는 기존 방법의 평가 시간을 10%까지 줄이면서도 동등한 결과를 제공하는 것으로 나타났습니다. 또한, 우리는 다단계 평가 과정을 통해 얻은 역사 기반 단계별 지시 조정 기록으로 구성된 데이터셋인 EA-CoT-10K를 구축하고 Qwen2.5-3B-Instruct를 기반으로 EA-3B 모델을 학습하여 Open Evaluation Agent (Open-EA)를 소개했습니다. 이 모델은 API 기반 에이전트의 구조화된 추론, 도구 호출 및 요약 프로토콜을 유지하면서 독점적인 백본에 대한 의존도를 줄이는 로컬 계획 백본 역할을 합니다. 실험 결과, API 기반 에이전트는 기존 T2I/T2V 벤치마크 및 개방형 질문에서 검증되었으며, Open-EA는 네 가지 동일 영역 T2V 생성 모델 패밀리와 세 가지 이질 영역 T2V 생성 모델 패밀리에 대해 평가되었습니다. 그 결과, 학습된 정책이 일부 패밀리 간에 전이되는 것을 확인했습니다.
Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands sampling hundreds or thousands of images or videos, which is computationally expensive. Existing evaluation methods also rely on rigid pipelines that overlook specific user needs and provide numerical results without clear explanations. Mimicking how humans quickly form impressions of a model's capabilities from only a few samples, we propose the Evaluation Agent framework, which employs human-like strategies for efficient, dynamic, multi-round evaluations, offering detailed, user-tailored analyses. Given a natural-language evaluation request, the agent decomposes it into sub-aspects, generates targeted prompts, samples images or videos from the evaluated model, invokes suitable evaluation tools, and iteratively updates its plan from the observed evidence, covering both predefined benchmark dimensions and open-ended user concerns. The framework is thus efficient, promptable, explainable, and scalable across models and tools. Experiments show that Evaluation Agent reduces evaluation time to 10% of traditional methods while delivering comparable results. We further introduce Open Evaluation Agent (Open-EA) by constructing EA-CoT-10K, a corpus of history-conditioned step-level instruction-tuning records derived from multi-round evaluation rollouts, and training EA-3B from Qwen2.5-3B-Instruct as a local planning backbone that preserves the structured reasoning, tool invocation, and summary protocol of the API-based agent while reducing dependence on proprietary backbones. Experiments validate the API-based agent on established T2I/T2V benchmarks and open-ended queries, and evaluate Open-EA on four in-domain and three out-of-domain T2V generator families, showing partial cross-family transfer of the learned policy.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.