Signal DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissively-Licensed Data
Summary
Current large language model development typically relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data practices. The paper introduces Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning Model architecture, trained from scratch using only permissively licensed post-training data. Despite the restrictive data-sourcing constraint and small parameter count, the model delivers highly competitive performance in English. The authors describe this as setting frontier-level results at this parameter scale. The work is positioned as a proof point that ethically sourced, openly licensed training data need not sacrifice model capability. This is relevant to the broader AI governance debate over training-data provenance.
Classification
Evidence 1
- DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissively-Licensed Data arXiv (cs.AI) 2026-08-13 accessed 2026-08-16T10:59:38+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-dd2ed27b33fe
