2607.29120v1 Jul 31, 2026 cs.LG

교육 과정이 중요하다: 합성 데이터를 활용한 데이터 효율적인 관계형 PFN 사전 학습

Curriculum Matters: Data-Efficient Relational PFN Pretraining with Synthetic Data

Mohammad Sadeq Abolhasani
Mohammad Sadeq Abolhasani
Citations: 23
h-index: 3
Viswanath Ganapathy
Viswanath Ganapathy
Citations: 197
h-index: 9

관계 Prior-Data Fitted Networks (PFNs)는 RDB-PFN과 같이, 수백만 개의 합성 작업을 통해 다중 테이블 관계형 데이터베이스에 대한 베이지안 추론을 근사합니다. 우리는 이 패러다임에 대한 세 가지 상호 관련된 질문을 조사합니다. 첫째, 구조적으로 다른 합성 생성기인 PluRel이 RDB-PFN의 prior를 대체할 수 있는가? 둘째, 합성 데이터를 PFN에 제시하는 순서가 다운스트림 성능에 얼마나 영향을 미치는가? 셋째, 관계형 데이터가 도입되기 전에 PFN은 단일 테이블 합성 사전 학습만으로 얼마나 많은 관계 추론 능력을 얻을 수 있는가? 모든 실험에서 PluRel을 유일한 합성 데이터 소스로 사용한 결과, 다음과 같은 사실이 밝혀졌습니다. (i) 7개에서 17개의 열로 점진적으로 스키마 복잡성을 넓히는 단일 테이블 교육 과정은 약 13,300개의 합성 테이블(RDB-PFN의 보고된 초기 학습 레시피보다 약 45배 적은 단일 테이블 데이터 세트)만을 사용하여 23개 작업의 표 형식 벤치마크에서 평균 ROC-AUC 0.703을 달성하는 반면, 동일한 데이터를 한 번에 학습하면 ROC-AUC는 0.541로 감소합니다. (ii) 약 5,500개의 PluRel 데이터베이스만을 사용하여 처음부터 학습된 관계형 교육 과정은 19개 작업의 RelBench/4DBInfer 벤치마크에서 평균 ROC-AUC 0.638을 달성하며, 이는 RDB-PFN이 보고한 성능의 약 88%에 해당하고, 관계형 합성 데이터는 약 220배 적게 사용했습니다. (iii) 단일 테이블 교육 과정 모델은 어떠한 관계형 적용 없이 직접 관계형 벤치마크에서 평가되었을 때 0.631이라는 값을 얻었으며, 이는 전용 관계형 파이프라인과 거의 일치합니다. 이러한 결과들을 종합적으로 고려할 때, 특정 관계형 생성기나 원시 합성 데이터 규모 자체보다 교육 과정 설계 및 합성 데이터의 다양성이 관계형 PFN 사전 학습에 더 중요한 영향을 미칠 수 있습니다.

Original Abstract

Relational Prior-Data Fitted Networks (PFNs) such as RDB-PFN approximate Bayesian inference over multi-table relational databases by pretraining on millions of synthetic tasks. We investigate three intertwined questions about this paradigm. First, can a structurally different synthetic generator PluRel substitute for RDB-PFN's prior? Second, how much does the order in which synthetic data is presented to the PFN affect downstream performance? Third, how much relational reasoning can a PFN acquire from single-table synthetic pretraining alone, before any relational data is introduced? Using PluRel as the sole synthetic data source across all experiments, we find: (i) a progressive single-table curriculum that gradually widens schema complexity from 7 to 17 columns reaches 0.703 average ROC-AUC on the 23-task tabular benchmark using only approximately 13,300 synthetic tables (approximately 45x fewer single-table datasets than RDB-PFN's reported warm-up recipe), while the same data trained all-at-once collapses to 0.541 ROC-AUC; (ii) a relational curriculum trained from scratch on only approximately 5,500 PluRel databases reaches 0.638 average ROC-AUC on the 19-task RelBench/4DBInfer benchmark, recovering 88% of RDB-PFN's reported performance with approximately 220x less relational synthetic data; and (iii) the single-table curriculum model, evaluated directly on the relational benchmark without any relational adaptation, achieves 0.631, nearly matching the dedicated relational pipeline. Together, these findings suggest that curriculum design and synthetic data diversity may matter more for relational PFN pretraining than the specific relational generator or raw synthetic scale alone.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!