2607.00926v1 Jul 01, 2026 cs.LG

생성적 메타학습에서의 인간-기계 협업: 모델 및 알고리즘

Human-Machine Collaboration on Generative Meta-Learning: Model and Algorithm

Samuel Kaski
Samuel Kaski
Citations: 55
h-index: 2
M. P. Unni
M. P. Unni
Citations: 5
h-index: 1

머신러닝 모델을 학습 데이터 분포와 다른 환경으로 일반화하는 것은 여전히 중요한 과제이며, 특히 목표 도메인의 데이터가 완전히 또는 부분적으로 없을 때 더욱 그렇습니다. 본 논문에서는 전문가의 직관을 활용하여 데이터 생성 과정을 안내함으로써 이러한 도메인 격차를 해소하는 새로운 프레임워크인 Generative Meta-Learning with Human Feedback (GMHF)를 제안합니다. 일반화 오차에 대한 이론적 분석을 바탕으로, 생성된 데이터 분포를 인간이 인지하는 목표 물리 법칙과 일치시키는 것이 위험을 크게 완화한다는 것을 보여주는 경계를 도출했습니다. GMHF는 Conditional Neural ODE (cNODE)를 생성 디지털 트윈으로 사용하고, 강화 학습(RL) 에이전트와 결합하여 이러한 통찰력을 구현합니다. 에이전트는 반복적으로 생성된 궤적의 잠재적인 물리 매개변수를 조정하며, 이를 통해 메타러너를 관측되지 않은 목표 분포로 효과적으로 유도합니다. 비선형 Duffing oscillator에 대한 실험 결과는 GMHF가 전문가의 신뢰도가 증가함에 따라 배포 손실을 크게 줄이며, 생성된 데이터와 목표 데이터 간의 차이가 신뢰할 수 있는 피드백 하에서 감소한다는 것을 보여주어, 본 연구에서 예측한 발산 최소화 메커니즘을 직접적으로 뒷받침합니다. 또한, 비동적 확률 모델에 대한 추가 실험은 프레임워크가 ODE로 제어되는 시스템을 넘어 확장될 수 있음을 확인하며, 인간-AI 협업이 분포 변화 하에서의 강력한 일반화를 위한 엄격하고 효과적인 촉매제가 될 수 있음을 입증합니다.

Original Abstract

Generalizing machine learning models to environments that differ from their training distribution remains a critical hurdle, particularly when data from the target domain is entirely or partially unavailable. We propose Generative Meta-Learning with Human Feedback (GMHF), a novel framework that bridges this domain gap by leveraging expert intuition to guide data synthesis. Grounded in a theoretical analysis of generalization error, we derive bounds demonstrating that aligning the distribution of generated data with human beliefs regarding the target physics significantly mitigates risk. GMHF operationalizes this insight by employing a Conditional Neural ODE (cNODE) as a generative digital twin, coupled with a Reinforcement Learning (RL) agent. The agent iteratively refines the latent physical parameters of the generated trajectories based on feedback, effectively steering the meta-learner toward the unobserved target distribution. Empirical validation on a nonlinear Duffing oscillator shows that GMHF substantially reduces deployment loss as expert reliability increases, and that the divergence between generated and target data falls under reliable feedback, directly corroborating the divergence-minimisation mechanism predicted by our theory. Further experiments on a non-dynamical probabilistic model confirm that the framework extends beyond ODE-governed systems, establishing human-AI collaboration as a rigorous catalyst for robust generalisation under distribution shift.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!