2607.27744v1 Jul 30, 2026 cs.LG

ROCS: 요청 기반 컴퓨팅 공유를 통한 효율적인 대규모 추천 시스템

ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation

Wenlin Chen
Wenlin Chen
Citations: 144
h-index: 2
Santanu Kolay
Santanu Kolay
Citations: 222
h-index: 8
Ellie Wen
Ellie Wen
Citations: 479
h-index: 9
Chunqiang Tang
Chunqiang Tang
Citations: 57
h-index: 3
Liang Luo
Liang Luo
Citations: 182
h-index: 5
Yuxin Chen
Yuxin Chen
Citations: 167
h-index: 6
Sijia Chen
Sijia Chen
Citations: 131
h-index: 3
Buyun Zhang
Buyun Zhang
Citations: 204
h-index: 6
Jian Jiao
Jian Jiao
Citations: 758
h-index: 8
Boda Li
Boda Li
Citations: 0
h-index: 0
Tongyi Tang
Tongyi Tang
Citations: 19
h-index: 2
Ao Cai
Ao Cai
Citations: 0
h-index: 0
Zijian Shen
Zijian Shen
Citations: 61
h-index: 4
Ryan Dick
Ryan Dick
Citations: 0
h-index: 0
Neng Shi
Neng Shi
Citations: 0
h-index: 0
Bin Yu
Bin Yu
Citations: 0
h-index: 0
Jianbo Xiao
Jianbo Xiao
Citations: 0
h-index: 0
Shuyao Bi
Shuyao Bi
Citations: 0
h-index: 0
Hongtao Yu
Hongtao Yu
Citations: 0
h-index: 0
Zhuoran Zhao
Zhuoran Zhao
Citations: 0
h-index: 0
Shuqi Yang
Shuqi Yang
Citations: 0
h-index: 0
Qianru Li
Qianru Li
Citations: 17
h-index: 1
Wei Ling
Wei Ling
Citations: 19
h-index: 2
Sihan Zeng
Sihan Zeng
Citations: 19
h-index: 2
Long-zhu Jin
Long-zhu Jin
Citations: 23
h-index: 2
Jiaxin Lu
Jiaxin Lu
Citations: 96
h-index: 4
Jiawei Li
Jiawei Li
Citations: 0
h-index: 0
Yichen Ruan
Yichen Ruan
Citations: 12
h-index: 2
Birmingham Guan
Birmingham Guan
Citations: 0
h-index: 0
Zijian Li
Zijian Li
Citations: 0
h-index: 0
Zeliang Chen
Zeliang Chen
Citations: 120
h-index: 5
Xiaohan Wei
Xiaohan Wei
Citations: 138
h-index: 7
G. Musumeci
G. Musumeci
Citations: 74
h-index: 4
Venkateshan Ranganathan
Venkateshan Ranganathan
Citations: 41
h-index: 1
Yantao Yao
Yantao Yao
Citations: 141
h-index: 2
Haoyu Wang
Haoyu Wang
Citations: 3
h-index: 1
Zhengkai Zhang
Zhengkai Zhang
Citations: 7
h-index: 1
Wenyi Xie
Wenyi Xie
Citations: 0
h-index: 0
Han Liu
Han Liu
Citations: 0
h-index: 0
Yuanwei Fang
Yuanwei Fang
Citations: 328
h-index: 7
Yang Chen
Yang Chen
Citations: 5
h-index: 1
Zikun Liu
Zikun Liu
Citations: 166
h-index: 5
Yinbin Ma
Yinbin Ma
Citations: 15
h-index: 2
Yong-Jae Lee
Yong-Jae Lee
Citations: 0
h-index: 0
Jian Sun
Jian Sun
Citations: 0
h-index: 0
Zhengyu Zhang
Zhengyu Zhang
Citations: 28
h-index: 2
Yuchen Hao
Yuchen Hao
Citations: 10
h-index: 2

최신 추천 모델은 특징 상호 작용 및 시퀀스 모듈을 확장하여 예측 정확도를 높이지만, 운영 비용 제약으로 인해 시스템의 확장성에 한계가 있습니다. 본 연구에서는 요청 기반 컴퓨팅 공유(ROCS)라는 새로운 모델링 및 추론 패러다임을 제안합니다. ROCS는 추천 추론의 고유한 특성을 활용하는데, 각 사용자 요청은 많은 후보 항목에 대해 평가되지만, 요청 측 특징은 모든 후보 항목에서 공유됩니다. ROCS는 요청-후보 항목 간 상호 작용을 최대한 늦게 수행하고, 후보 항목에 의존적인 표현을 분리하며, 모델의 상당 부분을 후보 항목당 한 번이 아닌 요청당 한 번만 평가하여 추론 효율성을 크게 향상시키는 동시에 예측 정확도를 유지하거나 개선합니다. 이러한 패러다임을 구현하기 위해, 특징 상호 작용 아키텍처에서 후보 항목 분리를 강제하는 일반화된 레이어 마스킹(GLM)과 시퀀스 아키텍처에 요청 기반 공유를 확장하는 딥 크로스 어텐션(DCA)을 개발했습니다. 또한 효율적인 GPU 배포를 지원하기 위해, ROCS 모델 실행 속도를 크게 향상시키는 커널 내 브로드캐스트 최적화(IKBO)를 공동 설계했습니다. 공개 벤치마크 실험 결과, ROCS는 다양한 추천 모델에서 품질-효율성 균형을 지속적으로 개선하는 것으로 나타났습니다. 실제 규모의 워크로드에서는 ROCS가 품질 저하 없이 검색 모델의 초당 처리량(QPS)을 최대 3배 향상시키고, 짧은 동영상 순위 모델에서 상대적인 LogLoss를 0.5% 개선하고 QPS를 50% 증가시켰습니다. ROCS는 광고 및 일반 콘텐츠 영역, 검색 및 순위 결정 단계 등 다양한 대규모 추천 시스템에 배포되었으며, 추론 복잡도의 두 배 이상의 차이를 보이며 온라인 성능을 크게 향상시키면서 인프라 비용을 절감했습니다.

Original Abstract

Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far systems can scale. In this work, we propose Request-Oriented Compute Sharing (ROCS), a modeling and inference paradigm that exploits a unique property of recommendation inference: each user request is evaluated against many candidates, while request-side features are shared across candidates. ROCS defers request-candidate interactions as late as possible, isolates candidate-dependent representations, and evaluates substantial portions of the model once per request rather than once per candidate, significantly improving inference efficiency while maintaining or improving prediction quality. To realize this paradigm, we develop Generalized Layer Masking (GLM) to enforce candidate isolation in feature-interaction architectures, and Deep Cross Attention (DCA) to extend request-oriented sharing to sequence architectures. To support efficient GPU deployment, we co-design In-Kernel Broadcast Optimization (IKBO) that significantly accelerates ROCS model execution. Experiments on public benchmarks show that ROCS consistently improves the quality-efficiency tradeoff across recommendation backbones. On production-scale workloads, ROCS achieves up to a 3x QPS improvement on retrieval models without quality degradation and a 0.5% relative LogLoss improvement with a 50% QPS gain on a short-form video ranking model. ROCS has been deployed across large-scale recommendation systems spanning ads and organic surfaces, retrieval and ranking stages, and more than two orders of magnitude in inference complexity, delivering significant online gains at reduced infrastructure cost.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!