메타-러닝을 위한 지식 그래프 임베딩과 메타-특징 통합
Integrating Meta-Features with Knowledge Graph Embeddings for Meta-Learning
웹 상에 존재하는 방대한 머신러닝 기록은 메타-러닝에 중요한 기회를 제공합니다. 메타-러닝은 과거 실험 결과를 활용하여 성능을 향상시키는 기술입니다. 두 가지 중요한 메타-러닝 과제는 다음과 같습니다. 첫째, 목표 데이터셋에 대한 파이프라인 성능 예측(PPE), 둘째, 유사한 성능 패턴을 보이는 데이터셋을 식별하는 데이터셋 성능 기반 유사성 추정(DPSE). 기존 접근 방식은 주로 데이터셋의 메타-특징(예: 인스턴스 수, 클래스 엔트로피 등)을 사용하여 데이터셋을 수치적으로 표현하고 이러한 메타-러닝 과제를 해결합니다. 그러나 이러한 접근 방식은 종종 사용 가능한 과거 실험 결과 및 파이프라인 메타데이터를 간과합니다. 이는 데이터셋-파이프라인 상호 작용을 파악하여 성능 유사성 패턴을 포착하는 능력을 제한합니다. 본 연구에서는 KGmetaSP라는 지식 그래프 임베딩 기반 접근 방식을 제안합니다. KGmetaSP는 기존 실험 데이터를 활용하여 이러한 상호 작용을 파악하고 PPE 및 DPSE 모두의 성능을 향상시킵니다. 데이터셋과 파이프라인을 통합된 지식 그래프(KG) 내에 표현하고, 파이프라인에 독립적인 메타-모델을 위한 PPE 및 거리 기반 검색을 위한 DPSE를 지원하는 임베딩을 생성합니다. 제안하는 접근 방식을 검증하기 위해, 144,177개의 OpenML 실험 데이터를 포함하는 대규모 벤치마크를 구축하여 풍부한 교차 데이터셋 평가를 수행했습니다. KGmetaSP는 단일 파이프라인 독립적인 메타-모델을 사용하여 정확한 PPE를 가능하게 하고, DPSE 성능을 기존 방식보다 향상시킵니다. 제안하는 KGmetaSP, 지식 그래프(KG), 그리고 벤치마크는 공개되어 메타-러닝 분야의 새로운 기준점을 제시하며, 공개된 실험 데이터를 통합된 KG로 구축하는 것이 이 분야에 어떻게 기여하는지를 보여줍니다.
The vast collection of machine learning records available on the web presents a significant opportunity for meta-learning, where past experiments are leveraged to improve performance. Two crucial meta-learning tasks are pipeline performance estimation (PPE), which predicts pipeline performance on target datasets, and dataset performance-based similarity estimation (DPSE), which identifies datasets with similar performance patterns. Existing approaches primarily rely on dataset meta-features (e.g., number of instances, class entropy, etc.) to represent datasets numerically and approximate these meta-learning tasks. However, these approaches often overlook the wealth of past experimental results and pipeline metadata available. This limits their ability to capture dataset - pipeline interactions that reveal performance similarity patterns. In this work, we propose KGmetaSP, a knowledge-graph-embeddings approach that leverages existing experiment data to capture these interactions and improve both PPE and DPSE. We represent datasets and pipelines within a unified knowledge graph (KG) and derive embeddings that support pipeline-agnostic meta-models for PPE and distance-based retrieval for DPSE. To validate our approach, we construct a large-scale benchmark comprising 144,177 OpenML experiments, enabling a rich cross-dataset evaluation. KGmetaSP enables accurate PPE using a single pipeline-agnostic meta-model and improves DPSE over baselines. The proposed KGmetaSP, KG, and benchmark are released, establishing a new reference point for meta-learning and demonstrating how consolidating open experiment data into a unified KG advances the field.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.