Intern-S2-Mobius: 분리된 지식과 추론 기능을 갖춘 기초 모델
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
본 논문에서는 전역적으로 공유되는 메모리(FFN)를 통해 지식 벡터를 저장하고, 여러 개의 추론 모듈(Self-Attn)이 반복적인 구성적 추론을 수행하는 Mobius-v0 아키텍처를 소개합니다. 추론 모듈은 숨겨진 상태를 캐시 및 전달 매개체로 사용하여 필요한 지식 벡터를 메모리에서 반복적으로 검색하며, 동시에 지식을 추론 연산자로 다시 전송합니다. 이러한 지식-추론 분리 아키텍처를 통해 Mobius는 더 나은 지식 압축과 추론 효율성을 달성합니다. Mobius-v0 아키텍처를 기반으로: 1) 처음부터 학습된 7B 모델은 62.6%의 기준 데이터만 사용하여 7B Transformer 기본 모델과 유사한 성능을 보입니다. 2) Qwen3.5-35B에서 지속적으로 사전 학습된 Intern-S2-Mobius는 유사한 성능을 달성하면서, 전체 추론 속도를 약 4배 향상시킵니다.
We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iteratively achieve compositional reasoning. Using hidden states as cache and carrier, reasoners repeatedly query memory for required knowledge-vectors, while the knowledge is transmitted back to reasoning operators. Through this knowledge-reasoning-separation architecture, Mobius achieves better knowledge compression and reasoning efficiency. Built upon Mobius-v0 architecture: 1) Our 7B model trained-from-scratch achieves similar downstream score as a 7B Transformer baseline with 62.6% of baseline's training data. 2) Our Intern-S2-Mobius, continually-pretrained from Qwen3.5-35B, achieves similar downstream score while delivering nearly 4x end-to-end inference speedup.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.