2607.24688v1 Jul 27, 2026 cs.DB

규모와 생성 모델을 넘어: 언어 모델 기반 개체 매칭 이해

Beyond Scale and Generation: Understanding Language Model-based Entity Matching

Zeyu Zhang
Zeyu Zhang
Citations: 45
h-index: 4
Xue Li
Xue Li
Citations: 15
h-index: 2
Iacer Calixto
Iacer Calixto
New York University
Citations: 1,573
h-index: 19
Sebastian Schelter
Sebastian Schelter
Citations: 63
h-index: 4
Paul Groth
Paul Groth
Citations: 32
h-index: 4

개체 매칭은 동일한 실제 개체를 참조하는 레코드를 식별합니다. 언어 모델은 바이-인코더, 크로스-인코더 및 생성형 매처 아키텍처를 통해 이 작업에 적용될 수 있습니다. 그러나 이전 연구에서는 종종 매처 아키텍처와 모델의 기본 구조, 모델 변형(다른 사전 훈련 목표를 반영), 그리고 모델 크기 간의 차이를 혼동하여 성능 향상의 원인을 명확하게 파악하기 어렵게 만듭니다. 우리는 Qwen3 패밀리의 세 가지 매처 아키텍처, 세 가지 모델 변형 및 세 가지 모델 크기를 포함하는 통제된 요인 실험을 통해 이 문제를 해결하고, 총 9개의 데이터 세트에서 1,215회의 파인튜닝을 수행했습니다. 또한, 데이터 세트 간의 일반화 성능과 계산 비용도 평가했습니다. 그 결과, 바이-인코더 모델에서는 모델 변형이 매우 중요하며, 임베딩 기반 변형은 더 나은 초기 설정 및 다운스트림 매칭 성능을 예측할 수 있는 유리한 표현 형상을 제공합니다. 크로스-인코더는 각 레코드를 독립적으로 표현하는 대신 레코드 쌍을 함께 인코딩하기 때문에 바이-인코더보다 일관되게 우수한 성능을 보이지만, 더 큰 모델은 이 격차를 부분적으로 좁힙니다. 생성형 매처가 항상 크로스-인코더보다 뛰어난 것은 아닙니다. 오히려, 데이터 분포의 변화, 즉 레코드 스키마의 미묘한 차이 및 데이터 세트 간 일반화와 같은 상황에서 장점을 보입니다. 또한, 더 큰 모델은 단축 학습에 더 의존하며, 반드시 더 나은 성능을 나타내지는 않습니다. 이러한 결과는 매처 아키텍처 간의 성능 차이를 유발하는 요인을 명확히 하고, 향후 연구 및 벤치마크 설계 방향을 제시합니다. 이는 아키텍처 선택과 모델 수준의 요인을 보다 명확하게 분리하고, 데이터 분포 변화와 데이터 세트 간 일반화 성능을 명시적으로 평가해야 합니다. 우리는 실험 결과, 코드, 학습 스크립트 및 평가 데이터를 https://github.com/Jantory/llm-trained-matcher 에서 공개합니다.

Original Abstract

Entity matching identifies records that refer to the same real-world entity. Language models can be adapted to this task through bi-encoder, cross-encoder, and generative matcher architectures. However, prior studies often conflate matcher architecture with differences in model backbone, model variant(reflecting different pretraining objectives), and model size, making it difficult to isolate the sources of performance gains. We address this issue through a controlled factorial study spanning three matcher architectures, three model variants and three model sizes from the Qwen3 family, and nine datasets, totaling 1,215 fine-tuning runs. We also evaluate cross-dataset transferability and computational cost. Our results show that model variant is critical for bi-encoders: embedding-oriented variants provide stronger initialization and more favorable representation geometry predictive of downstream matching performance. Cross-encoders retain a consistent advantage over bi-encoders because they jointly encode record pairs rather than representing each record independently, although larger models partially narrow this gap. Generative matchers do not universally outperform cross-encoders. Instead, their advantages concentrate under distribution shift, including subtle unseen differences in record schemas and cross-dataset transfer. We further find that larger models rely more heavily on shortcut learning and therefore do not necessarily perform better. These findings clarify the factors underlying performance differences across matcher architectures and motivate future research and benchmark designs that better disentangle architectural choices from model-level factors while explicitly evaluating distribution shift and cross-dataset transferability. We release our experimental results, code, training scripts, and evaluation data at https://github.com/Jantory/llm-trained-matcher.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!