2608.02887v1 Aug 03, 2026 cs.LG

일반화된 복지 최적화를 통한 인구 집단에 강건한 특징 선택 방법

Population-Robust Feature Selection via Generalized Welfare Optimization

A. Turcan
A. Turcan
Citations: 112
h-index: 5
Ruiqi Lyu
Ruiqi Lyu
Citations: 10
h-index: 2
Bryan Wilder
Bryan Wilder
Citations: 5
h-index: 1

어떤 특징을 수집할지는 실제 적용 시 중요한 결정이며, 동일한 제한적인 설문지, 검사 패널 또는 센서 세트가 여러 다양한 인구 집단을 대상으로 사용될 수 있습니다. 기존의 특징 선택 방법은 일반적으로 하나의 큰 인구 집단에 최적화되는 반면, 기존의 강건성 확보 방법은 각 인구 집단마다 별도의 공유 모델을 학습하는 경향이 있습니다. 본 논문에서는 PopFS라는 방법을 제안합니다. PopFS는 하나 이상의 공유 가능한 특징 집합을 학습하며, 이는 인구 집단 간의 차이에 강건하면서도 각 인구 집단은 자체적인 모델을 훈련할 수 있도록 합니다. PopFS는 조정 가능한 복지 목표를 사용하며, 이를 통해 연구자들은 전체 예측 성능과 소외된 인구 집단의 보호 사이의 균형을 맞출 수 있습니다. 이 목표를 대규모로 적용 가능하도록 만들기 위해, PopFS는 먼저 다중 작업 희소 학습을 사용하여 후보 풀을 줄인 다음, 유망한 추가 및 교체 항목을 순위화하여 직접적으로 특징 집합에 대해 검색하고, 짧은 목록에 포함된 항목만 전체적으로 재조정합니다. 5개의 표 형식 및 공공 보건 데이터 세트에서 추출한 6가지 예측 작업의 8가지 인구 집단 분할 결과, PopFS는 수천 개의 후보 특징에 대해 확장 가능하면서도 평균 성능과 최악의 성능을 보이는 인구 집단의 성능 모두에서 일관되게 우수한 결과를 달성합니다. 추가적으로, 43개 주를 대상으로 하는 COVID-19 예측 연구에서 복지 목표를 변경하면 전체적인 성능 변화가 거의 없으면서도 서비스가 부족한 주들의 성능을 향상시킬 수 있으며, 선택된 증상 신호에 대한 해석 가능한 변화를 얻을 수 있음을 보여줍니다. 저희의 코드는 https://github.com/Rachel-Lyu/PopFS 에서 확인할 수 있습니다.

Original Abstract

Choosing which features to collect is a deployment decision: the same limited questionnaire, test panel, or sensor set may need to serve several heterogeneous populations. Standard feature-selection methods typically optimize for one large population, while existing robust approaches tend to learn one shared model for every population. We introduce PopFS, a method for learning one shared, deployable feature set that is robust to population differences while letting each pop- ulation train its own model. PopFS uses a tunable welfare objective that lets practitioners balance overall predictive ben- efit against stronger protection of the populations that benefit least. To make this objective practical at scale, PopFS first uses multitask sparse learning to reduce the candidate pool, then searches directly over hard feature sets by ranking promising additions and swaps and fully refitting only a shortlist. Across eight population splits from six prediction tasks drawn from five tabular and public-health datasets, PopFS consistently achieves strong average and worst-population performance while scaling to thousands of candidate features. A 43-state COVID-19 nowcasting study further shows that changing the welfare objective can improve the least-served states with lit- tle change in average performance and yields an interpretable change in the selected symptom signals. Our code is available at https://github.com/Rachel-Lyu/PopFS.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!