LATTEArena: LLM 기반 표 형식 데이터 특징 엔지니어링 평가 프레임워크 (확장 버전)
LATTEArena: An Evaluation Framework for LLM-powered Tabular Feature Engineering (Extended Version)
특징 엔지니어링은 표 형식 데이터 분석에 필수적이며, 대규모 언어 모델(LLM)은 이 프로세스를 자동화하는 유망한 패러다임으로 부상하여 LLM 기반 자동 표 형식 특징 엔지니어링(LATTE)이 등장했습니다. 그러나 표준화된 플랫폼의 부족으로 인해 공정하고 비용 효율적인 비교가 어렵습니다. 또한, 복잡한 방법론 설계는 개별 구성 요소의 구체적인 기여도를 가리고 있습니다. 예를 들어, LFG는 트리 오브 씽킹, 몇 가지 예시 제공, 몬테카를로 트리 탐색 및 자연어 생성을 통합하지만, 각 기술의 경쟁 우위가 별도로 정량화되지 않았습니다. 이러한 문제점을 해결하기 위해, 우리는 다음과 같은 특징을 가진 최초의 경쟁 평가 프레임워크인 LATTEArena를 소개합니다: (1) 15가지 대표적인 방법을 재사용 가능한 구성 요소로 분해하는 6차원 분류 체계; (2) 통제된 비교를 위한 표준화된 모듈형 아레나; (3) 성능, 비용 및 견고성을 포함한 다각적 평가; (4) 각 기술의 경쟁 우위를 정량적으로 분석하는 구성 요소 수준의 제거 실험. 광범위한 평가를 통해 우리는 16가지 주요 결과를 밝혀냈습니다. 여기에는 다음이 포함됩니다: (1) 트리 오브 씽킹과 몬테카를로 트리 탐색은 최적의 비용 효율성을 달성합니다. (2) RPN 및 코드 출력 형식은 각각 분류 및 회귀 작업에서 우위를 차지합니다. 우리는 이 모듈형 프레임워크와 4000개 이상의 실행 로그를 공개하여 연구자들이 새로운 기술을 기존 기술과 비교하고 LATTE 분야를 발전시킬 수 있도록 지원합니다.
Feature engineering remains essential for tabular data analysis, and Large Language Models (LLMs) have emerged as a promising paradigm for automating this process, giving rise to LLM-powered AuTomated Tabular feature Engineering (LATTE). However, the absence of standardized platforms prevents fair, cost-aware comparisons. Furthermore, complex methodological designs obscure the specific contributions of individual components; for example, although LFG integrates Tree-of-Thought, few-shot demonstrations, Monte Carlo Tree Search, and natural language generation, the isolated impact of each technique's competitive edge remains unquantified. To address these challenges, we introduce LATTEArena, the first competitive evaluation framework featuring: (1) a six-dimensional taxonomy decomposing 15 representative methods into reusable components; (2) a standardized modular arena for controlled comparison; (3) multi-dimensional assessments covering performance, cost, and robustness; and (4) component-level ablation quantifying each technique's competitive edge. Through extensive evaluations, we reveal 16 key findings, including: (1) Tree-of-Thought with Monte Carlo Tree Search achieves optimal cost-effectiveness; (2) RPN and Code output formats dominate classification and regression tasks, respectively. We publicly release the modular framework and over 4000 execution logs, enabling researchers to seamlessly pit new techniques against existing ones and advance LATTE.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.