인터뷰가 있는 매칭 시장에서의 밴딧 학습
Bandit Learning in Matching Markets with Interviews
양면 매칭 시장은 양측의 선호도에 의존하지만, 선호도를 완벽히 평가하는 것은 종종 비현실적이다. 따라서 참가자들은 제한된 횟수의 인터뷰를 진행하여 초기 단계의 노이즈가 포함된 인상을 얻고 이를 바탕으로 최종 결정을 내린다. 본 연구는 인터뷰를 양측에 부분적인 선호도 정보를 제공하는 '저비용 힌트'로 모델링하여, 인터뷰가 포함된 매칭 시장에서의 밴딧 학습을 연구한다. 본 프레임워크는 기업 측의 불확실성을 허용한다는 점에서 기존 연구와 차별화된다. 즉, 구직자(에이전트)와 마찬가지로 기업도 자신의 선호도를 확신하지 못할 수 있으며, 덜 선호하는 에이전트를 고용하는 등 초기 채용 과정에서 실수를 범할 수 있다. 이를 해결하기 위해 우리는 '전략적 보류'(해당 라운드에서 고용하지 않기로 선택하는 것)를 허용하도록 기업의 행동 공간을 확장함으로써, 최적에 미치지 못하는 채용으로부터 회복할 수 있게 하고 별도의 조정이 필요 없는 분산 학습을 지원한다. 우리는 (i) 전지적 인터뷰 할당자가 존재하는 중앙 집중형 환경과, (ii) 두 가지 유형의 기업 측 피드백이 존재하는 분산형 환경을 위한 새로운 알고리즘들을 설계하였다. 모든 환경에서 본 알고리즘들은 시간에 독립적인 후회(regret)를 달성하는데, 이는 인터뷰 없이 안정적 매칭을 학습할 때 알려진 $O(\log T)$ 후회 한계에 비해 상당한 개선을 이룬 것이다. 또한, 약간의 구조적 조건이 주어진 시장에서는 분산형 알고리즘의 성능이 에이전트 및 기업 수의 다항식 배수 범위 내에서 중앙 집중형 알고리즘의 성능에 필적함을 보여준다.
Two-sided matching markets rely on preferences from both sides, yet it is often impractical to evaluate preferences. Participants, therefore, conduct a limited number of interviews, which provide early, noisy impressions and shape final decisions. We study bandit learning in matching markets with interviews, modeling interviews as \textit{low-cost hints} that reveal partial preference information to both sides. Our framework departs from existing work by allowing firm-side uncertainty: firms, like agents, may be unsure of their own preferences and can make early hiring mistakes by hiring less preferred agents. To handle this, we extend the firm's action space to allow \emph{strategic deferral} (choosing not to hire in a round), enabling recovery from suboptimal hires and supporting decentralized learning without coordination. We design novel algorithms for (i) a centralized setting with an omniscient interview allocator and (ii) decentralized settings with two types of firm-side feedback. Across all settings, our algorithms achieve time-independent regret, a substantial improvement over the $O(\log T)$ regret bounds known for learning stable matchings without interviews. Also, under mild structured markets, decentralized performance matches the centralized counterpart up to polynomial factors in the number of agents and firms.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.