ROCS: 요청 기반 컴퓨팅 공유를 통한 효율적인 대규모 추천 시스템
ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
최신 추천 모델은 특징 상호 작용 및 시퀀스 모듈을 확장하여 예측 정확도를 높이지만, 운영 비용 제약으로 인해 시스템의 확장성에 한계가 있습니다. 본 연구에서는 요청 기반 컴퓨팅 공유(ROCS)라는 새로운 모델링 및 추론 패러다임을 제안합니다. ROCS는 추천 추론의 고유한 특성을 활용하는데, 각 사용자 요청은 많은 후보 항목에 대해 평가되지만, 요청 측 특징은 모든 후보 항목에서 공유됩니다. ROCS는 요청-후보 항목 간 상호 작용을 최대한 늦게 수행하고, 후보 항목에 의존적인 표현을 분리하며, 모델의 상당 부분을 후보 항목당 한 번이 아닌 요청당 한 번만 평가하여 추론 효율성을 크게 향상시키는 동시에 예측 정확도를 유지하거나 개선합니다. 이러한 패러다임을 구현하기 위해, 특징 상호 작용 아키텍처에서 후보 항목 분리를 강제하는 일반화된 레이어 마스킹(GLM)과 시퀀스 아키텍처에 요청 기반 공유를 확장하는 딥 크로스 어텐션(DCA)을 개발했습니다. 또한 효율적인 GPU 배포를 지원하기 위해, ROCS 모델 실행 속도를 크게 향상시키는 커널 내 브로드캐스트 최적화(IKBO)를 공동 설계했습니다. 공개 벤치마크 실험 결과, ROCS는 다양한 추천 모델에서 품질-효율성 균형을 지속적으로 개선하는 것으로 나타났습니다. 실제 규모의 워크로드에서는 ROCS가 품질 저하 없이 검색 모델의 초당 처리량(QPS)을 최대 3배 향상시키고, 짧은 동영상 순위 모델에서 상대적인 LogLoss를 0.5% 개선하고 QPS를 50% 증가시켰습니다. ROCS는 광고 및 일반 콘텐츠 영역, 검색 및 순위 결정 단계 등 다양한 대규모 추천 시스템에 배포되었으며, 추론 복잡도의 두 배 이상의 차이를 보이며 온라인 성능을 크게 향상시키면서 인프라 비용을 절감했습니다.
Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far systems can scale. In this work, we propose Request-Oriented Compute Sharing (ROCS), a modeling and inference paradigm that exploits a unique property of recommendation inference: each user request is evaluated against many candidates, while request-side features are shared across candidates. ROCS defers request-candidate interactions as late as possible, isolates candidate-dependent representations, and evaluates substantial portions of the model once per request rather than once per candidate, significantly improving inference efficiency while maintaining or improving prediction quality. To realize this paradigm, we develop Generalized Layer Masking (GLM) to enforce candidate isolation in feature-interaction architectures, and Deep Cross Attention (DCA) to extend request-oriented sharing to sequence architectures. To support efficient GPU deployment, we co-design In-Kernel Broadcast Optimization (IKBO) that significantly accelerates ROCS model execution. Experiments on public benchmarks show that ROCS consistently improves the quality-efficiency tradeoff across recommendation backbones. On production-scale workloads, ROCS achieves up to a 3x QPS improvement on retrieval models without quality degradation and a 0.5% relative LogLoss improvement with a 50% QPS gain on a short-form video ranking model. ROCS has been deployed across large-scale recommendation systems spanning ads and organic surfaces, retrieval and ranking stages, and more than two orders of magnitude in inference complexity, delivering significant online gains at reduced infrastructure cost.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.