다양한 사고 패턴이 대규모 언어 모델의 추론 능력을 향상시킨다
Diverse Thinking Schemata Elicit Better Reasoning in Large Language Models
대규모 추론 모델(LRM)은 복잡한 수학 문제를 해결하기 위해 상세한 추론 과정을 생성하는 능력으로 인해 많은 관심을 받고 있습니다. 본 연구에서는 추론 과정의 중요한 두 가지 측면, 즉 추론 단계 간의 구체적인 전환 패턴을 나타내는 '추론 전환'과 모델이 생성하는 다양한 해법 경로를 반영하는 '답안 후보'에 주목합니다. 우리는 이 두 가지 측면을 통칭하여 '사고 패턴'이라고 정의합니다. 우리는 사고 패턴의 다양성과 모델 성능 사이에 상관관계가 있음을 확인했으며, 이를 바탕으로 사고 패턴의 다양성을 높여 추론 능력을 더욱 향상시킬 수 있다는 가설을 세웠습니다. 이를 위해, 본 연구에서는 모델에게 사고 패턴에 대한 인식을 부여하고, 강화 학습을 통해 다양성을 장려하며, 추론 시에도 다양한 추론을 유도하는 프레임워크인 '다양한 사고 패턴 정책 최적화(DiScO)'를 제안합니다. 여러 수학적 추론 벤치마크에서의 실험 결과는 DiScO가 기존의 그룹 상대 정책 최적화 방법보다 일관되게 우수한 성능을 보임을 보여줍니다. 정확도 외에도, 인간 전문가의 분석 결과는 DiScO가 모델이 초기 오류에서 회복하는 능력을 크게 향상시킨다는 것을 나타냅니다. 종합적으로 볼 때, 본 연구는 사고 패턴의 다양성이 중요한 역할을 한다는 점을 시사하며, 다양성 측면에서의 추가적인 연구 개발이 유망한 방향임을 제시합니다.
Large reasoning models (LRMs) have attracted increasing attention for their ability to solve complex mathematical problems by generating extended reasoning chains. In this work, we focus on two critical yet underexplored aspects of the reasoning process: reasoning transitions capturing the distinct transitions between reasoning steps and answer candidates reflecting the variety of solution paths produced by the model. We collectively define these two aspects as thinking schemata. We observe a correlation between the diversity of thinking schemata and model performance, which motivates us to enhance diversity as a means to further improve reasoning potential. To this end, we propose Diverse Schemata Policy Optimization (DiScO), a framework that first endows the model with schemata awareness, then encourages diversity through reinforcement learning, and further promotes diverse reasoning at inference time. Experiments on multiple mathematical reasoning benchmarks demonstrate that DiScO consistently outperforms standard group relative policy optimization. Beyond accuracy, human-annotated analyses show that DiScO substantially improves the model's ability to recover from erroneous initial attempts. Overall, our work suggests the important role that diversity of the thinking schemata plays and points to scaling along the diversity dimension as a promising research direction.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.