2608.13504v1 Aug 13, 2026 cs.LG

희소 직교 회귀 기법: 방정식 발견, 근사 및 통합을 위한 스펙트럴 프레임워크

Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration

S. Džeroski
S. Džeroski
Citations: 18,903
h-index: 65
S. Roman
S. Roman
Citations: 182
h-index: 8
L. Todorovski
L. Todorovski
Citations: 4,014
h-index: 32

본 논문에서는 희소 스펙트럴 프레임워크인 Sparse Orthogonal Regression Technique (SORT)를 개발합니다. SORT는 노이즈가 많고 불규칙하게 샘플링된 데이터로부터 직교 기저 확장을 학습하는 데 사용됩니다. SORT는 L1 정규화 회귀를 사용하여 관측값으로부터 직접 확장 계수를 추정하며, 명시적인 사분법 또는 해석적 내적 평가를 피합니다. 핵심 응용 분야는 데이터 기반의 상미분 방정식 발견입니다. 벡터장은 선택된 직교 기저로 표현되며 희소 계수 확장을 통해 학습됩니다. 이는 심볼릭 회귀, 문법 기반 발견 및 SINDy 스타일의 희소 식별에 대한 보완적인 접근 방식이며, 먼저 간결한 스펙트럴 표현을 복구하여 보다 간단한 해석적 형태를 찾는 검색을 안내할 수 있습니다. 동역학 시스템 실험에서, 선택된 직교 기저가 문제에 잘 맞을 경우 SORT는 기존의 라이브러리 기반 희소 회귀 방법을 능가하거나 동일한 성능을 보이며, 희소 샘플링, 노이즈가 많은 미분 추정 및 표현 불일치 상황에서도 더 안정적인 성능을 나타냅니다. 특정 예시를 통해 이러한 표현 방식이 유용한 이유를 설명합니다. 즉, 유한 라이브러리가 문제에 특화된 비선형성을 포함하지 못하는 경우, 결과 모델은 실패할 수 있습니다. SORT는 이러한 불일치로부터 완전히 자유롭지는 않지만, 일반적인 항 중에서 선택하는 취약한 방식을 피하고, 문제 영역에 적합하도록 설계된 기저를 사용하는 방식으로 문제를 전환합니다. 또한 실험 결과는 모델의 차수가 증가함에 따라 지배적인 저차 계수가 유지됨을 보여주며, 이는 일관성 있는 모델 성장을 뒷받침합니다. 방정식 발견 외에도, 동일한 학습된 확장은 복잡하고 고차원적인 적분 근사 및 추정에 사용될 수 있습니다. 전체적으로 SORT는 시스템 식별, 근사 및 통합을 위한 재사용 가능한 중간 표현을 제공하며, 기저 설계 문제를 과학 모델링의 명시적인 부분으로 만듭니다.

Original Abstract

We develop the Sparse Orthogonal Regression Technique (SORT), a sparse spectral framework for learning orthonormal-basis expansions from noisy and irregularly sampled data. SORT estimates expansion coefficients directly from observations using L1-regularized regression, avoiding explicit quadrature or analytic inner-product evaluation. The central application is data-driven discovery of ordinary differential equations: vector fields are represented in chosen orthogonal bases and learned as sparse coefficient expansions. This provides a complementary route to symbolic regression, grammar-based discovery, and SINDy-style sparse identification by first recovering a compact spectral representation, which can later guide searches for simpler analytic forms. Across the dynamical-system experiments, SORT matches or improves upon library-based sparse-regression baselines when the basis is well adapted to the problem, and shows more stable degradation under sparse sampling, noisy derivative estimates, and representation mismatch. Specific examples illustrate why this representation is useful: if a finite library misses the problem-specific nonlinearity, the resulting model can fail. SORT is not immune to mismatch, but it shifts the problem away from brittle selection among generic terms to basis design adapted to the problem domain. The experiments also show that dominant low-order coefficients persist as model order increases, supporting order-consistent model growth. Beyond equation discovery, the same learned expansion supports nonlinear approximation and estimation of complex, high-dimensional integrals by coefficient readout. Overall, SORT provides a reusable intermediate representation for system identification, approximation, and integration, while making basis design an explicit part of the scientific modeling problem.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!