CDR-Bench: 합성적이고 순서에 민감한 데이터 정제 레시피의 충실성 실행 평가
CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes
데이터 정제는 다단계 레시피를 사용하여 텍스트 상태를 변경하는 과정이며, 여기서 처리 연산자의 조합과 실행 순서가 결과에 영향을 미칩니다. 기존 벤치마크들은 텍스트 편집을 독립적으로 다루거나 코드 및 도구 실행과 함께 수행하지만, LLM이 이러한 합성적이고 순서에 민감한 데이터 정제 레시피를 직접적으로 그리고 충실하게 실행할 수 있는지 여부는 불분명합니다. 이 격차를 해소하기 위해, 우리는 3,462개의 고품질 작업으로 구성된 종합적인 벤치마크인 CDR-Bench를 소개하며, 이는 네 가지 실제 데이터 정제 영역과 29가지의 서로 다른 연산자를 포함합니다. 우리의 벤치마크는 원자적, 순서 무관 및 순서 민감 설정을 모두 평가하며, 결정론적인 참조 출력을 활용하여 정확한 평가를 가능하게 합니다. 10개 이상의 최첨단 LLM에 대한 실험 결과, 일관된 실패 패턴이 나타났습니다. 즉, 합성적 설정에서는 성능이 급격히 저하되며, 순서 민감 레시피의 성공률은 현저히 낮아집니다. 이러한 결과는 현재 LLM이 신뢰할 수 있는 합성 데이터 정제를 위해 필요한 절차적 충실성이 부족하다는 점을 강조합니다.
Data refinement involves executing multi-step recipes over evolving text states, where both composition and execution order of processing operators determine the outcome. While existing benchmarks either isolate text editing or entangle it with code and tool execution, it remains unclear whether LLMs can directly and faithfully execute these compositional, order-sensitive data refinement recipes. To fill this gap, we introduce CDR-Bench, a comprehensive benchmark featuring 3,462 high-quality tasks spanning four real-world data refinement domains and 29 distinct operators. Our benchmark evaluates models across atomic, order-agnostic, and order-sensitive settings, leveraging deterministic reference outputs to enable exact evaluation. Experiments on 10+ state-of-the-art LLMs reveal consistent failure patterns: performance degrades sharply in compositional settings, and order-sensitive recipe success collapses. These findings underline that current LLMs lack the procedural faithfulness required for reliable compositional data refinement.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.