DFM Mimir v1: 허용 가능한 학습 데이터만을 사용하여 10억 파라미터 규모의 개방형 HRM 모델이 제공하는 최첨단 성능
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
현재 대규모 언어 모델 개발은 종종 윤리적으로 문제가 될 수 있는 방대한 데이터셋에 의존하여, 오픈 소스 및 윤리적 데이터 사용을 지향하는 연구자들에게 높은 장벽으로 작용합니다. 본 논문에서는 계층적 추론 모델(HRM) 아키텍처를 기반으로 10억 파라미터 규모의 언어 모델인 Mimir v1을 소개합니다. Mimir v1은 처음부터 학습되었으며, 허용 가능한 학습 데이터만을 사용하여 영어 분야에서 뛰어난 성능을 제공하며, 특히 덴마크어 분야에서 새로운 최고 수준의 성능을 달성했습니다. 161개의 다양한 데이터셋으로 학습된 Mimir v1은 원래 HRM-Text 1B 모델보다 우수한 성능을 보이며, Qwen 3.5 4B 및 Gemma 4 E2B와 같은 더 큰 최첨단 모델과 경쟁합니다. 본 모델은 Hugging Face Hub에서 확인할 수 있습니다: https://huggingface.co/danish-foundation-models/DFM-Mimir
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, that is trained from scratch and delivers highly competitive performance for English and sets a new state of the art for Danish using only permissible post-training data. Trained on a mixture of 161 datasets, Mimir v1 outperforms the original HRM-Text 1B and competes with larger frontier models like Qwen 3.5 4B and Gemma 4 E2B, tested across 20 benchmarks for English, Math & Code and Danish. The model is available on the Hugging Face Hub: https://huggingface.co/danish-foundation-models/DFM-Mimir
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.