2607.25228v1 Jul 28, 2026 cs.CL

LLM 기반 기호화된 의사 결정 과정을 통해 해석 가능한 컬럼 어노테이션

Interpretable Column Annotation with LLM-Symbolized Decision Process Materialization

Zhenchang Xing
Zhenchang Xing
Citations: 434
h-index: 12
Jianwei Wang
Jianwei Wang
Citations: 116
h-index: 5
Liming Zhu
Liming Zhu
Citations: 28
h-index: 3
Mengqi Wang
Mengqi Wang
Citations: 14
h-index: 2
Qing Liu
Qing Liu
Citations: 53
h-index: 5
Xiwei Xu
Xiwei Xu
Citations: 161
h-index: 6
Michael Bain
Michael Bain
Citations: 6
h-index: 1
Wenjie Zhang Unsw Sydney
Wenjie Zhang Unsw Sydney
Citations: 0
h-index: 0
Data61
Data61
Citations: 23
h-index: 3
Csiro
Csiro
Citations: 0
h-index: 0

컬럼 어노테이션(CA)은 컬럼 유형 어노테이션(CTA) 및 컬럼 속성 어노테이션(CPA)을 포함하며, 테이블의 컬럼 의미와 그들 간의 의미적 관계를 파악하는 것을 목표로 합니다. 최근 CA 방법들은 주로 다양한 신경망 모델을 사용하여 컬럼 표현을 학습하고 이를 직접 레이블 범주에 매핑하는데, 이는 (1) 모델의 해석 가능성과 적응성을 저해하고, (2) 풍부한 레이블 의미를 간과하여 궁극적으로 정확도를 제한합니다. 이러한 한계를 극복하기 위해, 우리는 LLM 기반의 해석 가능한 CA 프레임워크인 SymCA를 제안합니다. SymCA는 컬럼 어노테이션을 전역에서 지역으로의 기호화된 의사 결정 과정으로 구현합니다. SymCA는 두 가지 구성 요소로 이루어집니다: (1) 전역 스켈레톤 유도, 이는 레이블 공간에 의미론적 스켈레톤을 구축하고, (2) 로컬 서브스트레이트 진화, 이는 스켈레톤 내에서 예측 서브스트레이트를 발전시킵니다. 특히, 해석 가능한 의사 결정 과정을 유지하면서 레이블 의미를 활용하기 위해, 전역 스켈레톤 유도 모듈은 LLM을 사용하여 후보 하이퍼님(hypernym)-영감적인 트리 구조의 의미론적 스켈레톤을 생성하고, Minimum Bayes Risk (MBR) 기반의 합의 전략을 사용하여 생성 변동에 강건한 스켈레톤을 선택합니다. 또한, 각 내부 노드는 자식 노드들을 구별하는 데 서로 다른 증거가 필요하므로, 로컬 서브스트레이트 진화 모듈은 각 내부 노드를 실행 가능하고 발전 가능한 예측 서브스트레이트로 구현합니다. 여러 번의 진화 과정을 거치면서, 각 서브스트레이트는 현재 연산자 집합을 사용하여 해석 가능한 랜덤 포레스트 분류기를 훈련하고, LLM을 활용하여 노드별 연산자 수정을 제안하며, 탐색-활용 전략을 사용하여 유망한 서브스트레이트를 우선시합니다. 광범위한 실험 결과는 SymCA가 정확하고, 견고하며, 해석 가능하며, Micro-F1에서 평균 6.42%, Macro-F1에서 평균 11.03%로 최첨단 기준 모델보다 우수한 성능을 보인다는 것을 보여줍니다.

Original Abstract

Column annotation (CA), including column type annotation (CTA) and column property annotation (CPA), aims to identify the meanings of table columns and the semantic relationships among them. Recent CA methods usually use various neural models to learn column representations and directly map them to label categories, thereby (1) sacrificing model interpretability and adaptivity, and (2) overlooking rich label semantics and ultimately limiting accuracy. To address these limitations, we propose SymCA, an LLM-empowered interpretable CA framework that materializes column annotation as a global-to-local symbolic decision process. SymCA consists of two components: (1) global skeleton induction, which constructs a semantic skeleton over the label space, and (2) local substrate evolution, which evolves predictive substrates within the skeleton. Specifically, to exploit label semantics while preserving an interpretable decision process, the global skeleton induction module leverages LLMs to generate candidate hypernym-inspired tree-structured semantic skeletons and employs a Minimum Bayes Risk (MBR)-based consensus strategy to select a robust skeleton against generation variance. Since different internal nodes require different evidence to distinguish among their child nodes, the local substrate evolution module materializes each internal node as an executable and evolvable predictive substrate. Over multiple evolution rounds, each substrate trains an interpretable random forest classifier with the current operator set, leverages the LLM to propose node-specific operator modifications, and uses an exploration-exploitation strategy to prioritize promising substrates. Extensive experiments demonstrate that SymCA is accurate, robust, and interpretable, outperforming the strongest baselines by an average of 6.42% in Micro-F1 and 11.03% in Macro-F1.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!