블랙박스를 조종할 수 있을까? 협력 에이전트를 활용한 추천 시스템의 제어 가능성 중심 평가
Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents
추천 시스템은 블랙박스와 같이 작동하여 사용자나 규제 기관이 시스템의 출력 결과를 특정 의도에 맞게 조정하거나 시스템의 동작을 감사하는 데 어려움을 겪습니다. 이러한 제어 불가능성은, 즉 시스템이 명시적인 지침에 반응할 수 있는 능력 부족은 기존 평가 방식에서 해결되지 않은 중요한 문제입니다. 이 문제를 해결하기 위해, 우리는 제어 가능성을 체계적으로 평가하기 위한 협력 다중 에이전트 프레임워크인 CtrlBench-Rec을 제안합니다. 우리는 목표 콘텐츠 발견, 관심 프로필 형성, 그리고 인기 편향 완화라는 세 가지 기본적인 과제를 정의하여 명시적인 명령으로부터 암묵적인 표현 조작, 그리고 최종적으로 알고리즘 편향 극복에 이르는 다양한 측면의 제어 가능성을 측정합니다. 실제 데이터셋과 다양한 추천 모델에 대한 광범위한 실험 결과, 우리 프레임워크는 제어 가능성을 효과적으로 정량화하고 시스템의 중요한 병목 현상을 드러내는 것으로 나타났습니다. 특히, 장기적인 콘텐츠에 대한 지침을 제공하는 데 지속적인 어려움이 존재합니다. CtrlBench-Rec은 제어 가능한 추천 연구, 알고리즘 감사 및 사용자 권한 강화를 위한 최초의 표준 도구 모음을 제공합니다. 저희 코드는 https://github.com/caskcsg/CtrlBenchRec 에서 공개되어 있습니다.
Recommender systems operate as Black-Boxes, leaving users and regulators unable to steer their outputs toward specific intentions or audit their behavior. This lack of controllability, defined as the system's ability to respond to explicit guidance, remains an unaddressed dimension in existing evaluation paradigms. To fill this gap, we propose CtrlBench-Rec, a collaborative multi-agent framework for systematic assessment of controllability. We formalize three fundamental tasks: target content discovery, interest profile shaping, and popularity bias mitigation, which together measure steerability from explicit commands to implicit representation steering and finally to overcoming algorithmic biases.Extensive experiments on real-world datasets and multiple recommendation models demonstrate that our framework effectively quantifies controllability and exposes critical system bottlenecks, most notably persistent resistance to guiding long tail content. CtrlBench-Rec provides the first standardized toolkit for controllable recommendation research, algorithmic auditing, and user empowerment. Our code is released on https://github.com/caskcsg/CtrlBenchRec.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.