2608.04669v1 Aug 05, 2026 cs.LG

이중 가격을 통한 차별화: 용량 제약 하에서의 엔드 투 엔드 정책 학습

Differentiating Through Dual Prices: End-to-End Policy Learning Under Capacity Constraints

Mahdi Salmani
Mahdi Salmani
Citations: 2
h-index: 1
M. Haghi
M. Haghi
Citations: 0
h-index: 0
Nima Kelidari
Nima Kelidari
Citations: 0
h-index: 0

많은 사회 서비스는 주택 지원이나 병원 개입과 같이 희소 자원을 한 번에 한 명씩 방문하는 사람들에게 할당합니다. 각 방문자에게 즉시 결정을 내려야 하며, 모든 자원의 장기적인 사용량은 용량을 초과해서는 안 됩니다. 본 연구에서는 관찰 데이터로부터 이러한 자원 배분 정책을 학습하는 방법을 다룹니다. 기존 방식은 의사 결정에 대한 고려 없이 수행됩니다. 즉, 각 옵션별로 회귀 분석을 통해 결과를 예측하고, 제한된 자원의 가격을 설정한 후, 예측 결과에서 가격을 뺀 값이 가장 큰 옵션을 방문자에게 할당합니다. 본 연구에서는 이러한 과정을 엔드 투 엔드로 학습하며, 배포된 정책의 가치를 이중 가격 자체를 통해 미분하여 추정합니다. 정확하지만 비선형인 모델과, 기댓값으로 용량 제약을 만족하는 볼록 완화 모델, 그리고 스무딩 온도와 옵션 수에 따라 제한적인 성능 저하가 발생하는 모델 두 가지를 분석합니다. 모든 방법은 자원의 사용량이 용량만큼 보충되는 큐잉 시뮬레이션을 통해 평가되었습니다. 여섯 개의 데이터 세트에서, 엔드 투 엔드 방식은 지연 비용과 관계없이 배포 조정된 가치 지수에서 상위 순위를 차지했습니다. 특히, 용량 제약이 중요한 경우, 기존 방식은 이러한 제약을 자주 위반하며 훨씬 더 긴 대기 시간을 초래합니다. 7만 명 규모의 병원 환자 데이터 세트에서도 엔드 투 엔드 학습은 상당한 수준의 정책 가치를 달성했으며, 이는 성능이 비슷한 신경망 기반 모델보다도 우수했습니다. 기존 방식의 회귀 분석은 측정 가능한 실제 값 예측에는 여전히 더 강력하지만, 자원이 실제로 부족하고 실현 가능성이 중요한 환경에서는 엔드 투 엔드 학습이 더 적합합니다.

Original Abstract

Many social services assign scarce resources, such as housing assistance or hospital interventions, to people who arrive one at a time: each arrival must receive a decision immediately, and the long-run usage of every resource must stay within its capacity. We study how to learn such an assignment policy from logged observational data. The standard pipeline is decision-blind: fit one outcome model per arm by regression, price each capacitated resource from the fitted models, and assign each arrival the arm whose predicted outcome minus price is largest. We instead train the outcome models end-to-end, differentiating an off-policy estimate of the deployed policy's value through the dual prices themselves. We study two formulations: an exact nonconvex one, and a convex relaxation whose optimum always satisfies the capacity constraints in expectation and which is suboptimal by at most a term linear in the smoothing temperature and logarithmic in the number of arms. Every method is evaluated in a queueing simulation with resources replenished at their capacity rates. Across six datasets, the two end-to-end variants take the top slots on a deployment-adjusted value index at every delay cost, including zero; when capacities are binding, decision-blind baselines frequently violate them and incur much longer queueing delays. On the largest dataset, a hospital cohort of seventy thousand patients, end-to-end training also achieves significantly higher policy value, a margin that survives a capacity-matched neural baseline. Flexible decision-blind regression remains the stronger pure predictor where ground truth is measurable; end-to-end training is best suited to settings where resources are genuinely scarce and feasibility matters.

0 Citations
0 Influential
0.5 Altmetric
2.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!