2608.14290v1 Aug 14, 2026 cs.AI

Intern-S2-Mobius: 분리된 지식과 추론 기능을 갖춘 기초 모델

Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Weida Wang
Weida Wang
Shanghai AI Laboratory
Citations: 189
h-index: 9
Bowen Yang
Bowen Yang
University of Science and Technology of China
Citations: 1,032
h-index: 5
Dahua Lin
Dahua Lin
Citations: 2,451
h-index: 22
Rui Wang
Rui Wang
Citations: 36
h-index: 2
Lixin Gu
Lixin Gu
Citations: 4,000
h-index: 6
Huanze Tang
Huanze Tang
Citations: 493
h-index: 4
Haijun Lv
Haijun Lv
Citations: 565
h-index: 7
Kai Chen
Kai Chen
Citations: 18
h-index: 3
Jiaye Ge
Jiaye Ge
Citations: 3,509
h-index: 7
Ermo Hua
Ermo Hua
Citations: 669
h-index: 11
Youbang Sun
Youbang Sun
Citations: 779
h-index: 9
Ning Ding
Ning Ding
Citations: 97
h-index: 4
Bowen Zhou
Bowen Zhou
Citations: 956
h-index: 9
Yicheng Gu
Yicheng Gu
Citations: 536
h-index: 8
Dawei Liu
Dawei Liu
Citations: 84
h-index: 2
Haozheng Hou
Haozheng Hou
Citations: 0
h-index: 0
Biqing Qi
Biqing Qi
Citations: 2,561
h-index: 23
Jie Hou
Jie Hou
Citations: 0
h-index: 0
Xiangyu Hong
Xiangyu Hong
Citations: 81
h-index: 2
Minxi Jin
Minxi Jin
Citations: 10
h-index: 2
Cheng Liang
Cheng Liang
Citations: 29
h-index: 1
Han Lv
Han Lv
Citations: 4,009
h-index: 6
Ningsheng Ma
Ningsheng Ma
Citations: 17
h-index: 2
Jian Qian
Jian Qian
Citations: 0
h-index: 0
Shiya Su
Shiya Su
Citations: 0
h-index: 0
Zhongbo Tian
Zhongbo Tian
Citations: 98
h-index: 3
Ting Wang
Ting Wang
Citations: 45
h-index: 4
Baiting Wu
Baiting Wu
Citations: 1
h-index: 1
Haochen Ye
Haochen Ye
Citations: 417
h-index: 4
Shan Yu
Shan Yu
Citations: 0
h-index: 0
Xiaoyi Yu
Xiaoyi Yu
Citations: 0
h-index: 0
Qi Zeng
Qi Zeng
Citations: 0
h-index: 0
Qi Zhang
Qi Zhang
Citations: 0
h-index: 0
Ming Zhang
Ming Zhang
Citations: 128
h-index: 5

본 논문에서는 전역적으로 공유되는 메모리(FFN)를 통해 지식 벡터를 저장하고, 여러 개의 추론 모듈(Self-Attn)이 반복적인 구성적 추론을 수행하는 Mobius-v0 아키텍처를 소개합니다. 추론 모듈은 숨겨진 상태를 캐시 및 전달 매개체로 사용하여 필요한 지식 벡터를 메모리에서 반복적으로 검색하며, 동시에 지식을 추론 연산자로 다시 전송합니다. 이러한 지식-추론 분리 아키텍처를 통해 Mobius는 더 나은 지식 압축과 추론 효율성을 달성합니다. Mobius-v0 아키텍처를 기반으로: 1) 처음부터 학습된 7B 모델은 62.6%의 기준 데이터만 사용하여 7B Transformer 기본 모델과 유사한 성능을 보입니다. 2) Qwen3.5-35B에서 지속적으로 사전 학습된 Intern-S2-Mobius는 유사한 성능을 달성하면서, 전체 추론 속도를 약 4배 향상시킵니다.

Original Abstract

We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iteratively achieve compositional reasoning. Using hidden states as cache and carrier, reasoners repeatedly query memory for required knowledge-vectors, while the knowledge is transmitted back to reasoning operators. Through this knowledge-reasoning-separation architecture, Mobius achieves better knowledge compression and reasoning efficiency. Built upon Mobius-v0 architecture: 1) Our 7B model trained-from-scratch achieves similar downstream score as a 7B Transformer baseline with 62.6% of baseline's training data. 2) Our Intern-S2-Mobius, continually-pretrained from Qwen3.5-35B, achieves similar downstream score while delivering nearly 4x end-to-end inference speedup.

0 Citations
0 Influential
11.5 Altmetric
57.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!