OneReason 기술 보고서
OneReason Technical Report
OneRec 계열의 생성형 추천 모델은 짧은 동영상, 라이브 스트리밍, 광고 및 전자 상거래와 같은 다양한 실제 서비스에 광범위하게 적용되었습니다. 그러나 이러한 생성형 모델은 확장성의 이점을 얻을 수 있지만, 아이템 토큰만을 사용하여 의미 있는 연쇄적 사고(Chain-of-Thought, CoT) 시퀀스를 구성할 수 없기 때문에 추론 능력을 활성화하기 어렵습니다. 대규모 언어 모델(LLM) 분야에서 '답변 전에 생각한다'는 추론 기반 패러다임의 성공에 영감을 받아, 우리는 생성형 추천 시스템에서의 추론 능력 탐색을 위한 초기 연구(OneRec-Think, OpenOneRec)를 수행했습니다. 그러나 예상치 못한 현상을 발견했는데, 사고 모드가 비사고 모드보다 우월하지 않았습니다. 다중 모달 언어 모델에서 CoT의 강건성에 대한 최근 연구 결과를 바탕으로, 추천 시스템에서의 효과적인 추론은 두 가지 요소에 달려 있다고 주장합니다. 첫째는 지각 능력으로, 이는 아이템 토큰을 해당 어휘 의미에 연결하는 능력입니다. 둘째는 인지 능력으로, 이는 사용자의 행동 시퀀스를 일관된 잠재적 관심 영역으로 재구성하는 능력입니다. 따라서 우리는 다음과 같은 요소들을 포함하는 OneReason을 제안합니다: (1) 사전 훈련 단계에서 강력한 아이템 토큰 지각 능력을 확보하고, (2) 지도 학습(SFT) 단계에서 세 수준의 인지 강화 CoT 형식을 추천 작업에 적용하며, (3) 강화 학습(RL) 단계를 통해 사고 능력을 향상시키기 위한 특화 후 통합 훈련 방식을 사용합니다.
Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. However, these generative models can only benefit from the scaling advantage, while their reasoning ability is hard to activate, since we cannot construct meaningful Chain-of-Thought (CoT) sequences consisting of itemic tokens only. Inspired by the success of the reasoning-style ``think before answer'' paradigm in the LLM field, we conduct preliminary studies (i.e., OneRec-Think, OpenOneRec) to explore reasoning capability in generative recommendation. Nevertheless, we notice an unexpected phenomenon: the thinking mode does not show advantages over the non-thinking mode. Drawing insights from recent findings on CoT robustness in multi-modal language models, we argue that effective reasoning in recommendation rests on two factors: perception, the ability to ground itemic tokens in their underlying language semantics, and cognition, the ability to reorganize a user's behavior sequence into coherent latent interest points. We therefore propose OneReason, which includes: (1) strong itemic token perception in pre-training, (2) a three-level cognition-enhanced CoT format for recommendation tasks in SFT, and (3) a specialize-then-unify training recipe in RL to enhance the thinking ability.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.