STELLAR-E: 합성, 맞춤형, 엔드투엔드 LLM 애플리케이션 평가 시스템
STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator
다양한 분야에서 대규모 언어 모델(LLM)의 활용이 증가함에 따라, 특정 도메인 및 언어에 특화된 견고한 평가 데이터셋의 필요성이 강조되고 있습니다. 그러나 개인 정보 보호 문제, 규제 제한, 그리고 수동 생성에 필요한 시간 등의 어려움으로 인해 그러한 데이터셋을 구축하는 것은 쉽지 않습니다. 기존의 자동화된 벤치마킹 방법은 종종 기존 데이터에 의존하고, 확장성이 낮으며, 특정 도메인에만 초점을 맞추고, 다국어 지원이 부족한 한계를 가지고 있습니다. 본 논문에서는 최소한의 인간 개입으로 기존 데이터셋에 의존하지 않고, 원하는 크기의 고품질 합성 데이터셋을 자동으로 생성하는 시스템인 STELLAR-E를 소개합니다. 이 시스템은 두 단계로 구성됩니다. (1) TGRT Self-Instruct 프레임워크를 수정하여 제어 가능한 맞춤형 합성 데이터셋 생성 기능을 제공하는 합성 데이터 엔진을 구축하고, (2) 통계적 지표와 LLM 기반 지표를 통합하여 LLM 기반 애플리케이션 평가에 합성 데이터셋의 적용 가능성을 평가하는 평가 파이프라인을 구현했습니다. 실험 결과, STELLAR-E에서 생성된 합성 데이터셋은 기존의 언어별 벤치마크에 비해 평균 +5.7%의 LLM 평가 점수 차이를 보여, 대규모 및 소규모 LLM의 종합적인 평가를 위한 동등한 수준의 품질을 제공합니다. 실제 데이터셋이 특히 소규모 모델의 경우 LLM에게 약간 더 어려운 과제를 제시하지만, 본 연구는 LLM 애플리케이션의 공정한 평가를 지원하는 확장 가능하고 도메인 적응 가능한 벤치마킹 프레임워크를 제시하며, 수동 방식에 비해 빠르고 효율적인 자동화된 품질 보증 시스템을 가능하게 합니다.
The increasing reliance on Large Language Models (LLMs) across diverse sectors highlights the need for robust domain-specific and language-specific evaluation datasets; however, the collection of such datasets is challenging due to privacy concerns, regulatory restrictions, and the time cost for manual creation. Existing automated benchmarking methods are often limited by relying on pre-existing data, poor scalability, single-domain focus, and lack of multilingual support. We present STELLAR-E - a fully automated system to generate high-quality synthetic datasets of custom size, using minimal human inputs without depending on existing datasets. The system is structured in two stages: (1) We modify the TGRT Self-Instruct framework to create a synthetic data engine that enables controllable, custom synthetic dataset generation, and (2) an evaluation pipeline incorporating statistical and LLM-based metrics to assess the applicability of the synthetic dataset for LLM-based application evaluations. The synthetic datasets reach an average difference of +5.7% in terms of LLM-as-a-judge scores against existing language-specific benchmarks, demonstrating comparable quality for comprehensive assessment of big and small LLMs. While real datasets remain slightly more challenging for LLMs especially for smaller models, this work establishes a scalable and domain-adaptable benchmarking framework that supports fair evaluation of LLM applications, offering a faster alternative to manual approaches and enabling high-efficiency automated quality assurance cycles.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.