MetaGAI: 생성 AI 모델 및 데이터 카드 생성을 위한 대규모, 고품질 벤치마크
MetaGAI: A Large-Scale and High-Quality Benchmark for Generative AI Model and Data Card Generation
생성 AI의 급속한 확산은 투명성과 거버넌스를 위한 엄격한 문서화 기준을 요구합니다. 그러나 모델 및 데이터 카드의 수동 생성은 확장성이 떨어지며, 자동화된 접근 방식은 체계적인 평가를 위한 대규모, 고품질 벤치마크가 부족합니다. 본 논문에서는 학술 논문, GitHub 저장소, Hugging Face 아티팩트를 활용한 의미론적 삼각 측량을 통해 구성된 2,541개의 검증된 문서 세트를 포함하는 포괄적인 벤치마크인 MetaGAI를 소개합니다. 기존의 단일 소스 데이터셋과 달리, MetaGAI는 특화된 검색(Retriever), 생성(Generator), 편집(Editor) 에이전트를 사용하는 다중 에이전트 프레임워크를 채택하며, 편집자가 개선한 정답에 대한 인간 평가를 포함한 4차원 인간-루프 평가를 통해 검증되었습니다. 우리는 자동화된 지표와 검증된 LLM-as-a-Judge 프레임워크를 결합한 견고한 평가 프로토콜을 구축했습니다. 광범위한 분석 결과, 희소한 Mixture-of-Experts 아키텍처가 우수한 비용-품질 효율성을 달성하는 것으로 나타났으며, 정확성과 완전성 간에는 근본적인 상충 관계가 존재합니다. MetaGAI는 대규모 자동 모델 및 데이터 카드 생성 방법의 벤치마킹, 학습 및 분석을 위한 기본적인 테스트 환경을 제공합니다. 저희의 데이터 및 코드는 다음 주소에서 이용 가능합니다: https://github.com/haoxuan-unt2024/MetaGAI-Benchmark.
The rapid proliferation of Generative AI necessitates rigorous documentation standards for transparency and governance. However, manual creation of Model and Data Cards is not scalable, while automated approaches lack large-scale, high-fidelity benchmarks for systematic evaluation. We introduce MetaGAI, a comprehensive benchmark comprising 2,541 verified document triplets constructed through semantic triangulation of academic papers, GitHub repositories, and Hugging Face artifacts. Unlike prior single-source datasets, MetaGAI employs a multi-agent framework with specialized Retriever, Generator, and Editor agents, validated through four-dimensional human-in-the-loop assessment, including human evaluation of editor-refined ground truth. We establish a robust evaluation protocol combining automated metrics with validated LLM-as-a-Judge frameworks. Extensive analysis reveals that sparse Mixture-of-Experts architectures achieve superior cost-quality efficiency, while a fundamental trade-off exists between faithfulness and completeness. MetaGAI provides a foundational testbed for benchmarking, training, and analyzing automated Model and Data Card generation methods at scale. Our data and code are available at: https://github.com/haoxuan-unt2024/MetaGAI-Benchmark.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.