- Source
- arXiv
- Published
- Runtime
- 0:00
- Snippets
- 4
A conversation between
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
§02
Snippets
-
Mimir v1, a 1B-parameter model trained only on permissible data, matches or exceeds larger frontier models like Qwen 3.5 4B across 20 English, math, code, and Danish benchmarks.
Demonstrates that ethical data sourcing and open development need not sacrifice performance, lowering barriers for reproducible research.
-
Mimir v1 achieves state-of-the-art performance on Danish benchmarks, the first 1B model to do so with only permissible post-training data.
Shows that careful data curation and multi-dataset mixing can unlock strong performance in underserved languages without proprietary corpora.
-
Mimir v1 combines 161 permissible datasets through a curated mixture strategy, delivering competitive performance without access to private or restricted training corpora.
Validates that thoughtful data engineering and curation can compensate for smaller model size and limited proprietary data.
-
The HRM architecture, applied to Mimir v1, scales effectively to 1B parameters and rivals larger frontier models on diverse benchmarks.
Suggests alternative architectures deserve exploration and may offer competitive advantages when paired with good training practices.
§03
Synthesis
The Core Achievement
Mimir v1 challenges the assumption that frontier-level language model performance requires access to massive, legally murky datasets. The authors built a 1-billion-parameter model trained entirely on permissible data—datasets with clear licensing and ethical sourcing—and showed it competes with larger, less-scrupulous peers. On English benchmarks, Mimir v1 matches or beats models like Qwen 3.5 4B and Gemma 2 8B. For Danish, it sets a new state-of-the-art, a significant result for under-resourced languages often neglected in AI development.
How It Works
Mimir v1 uses the Hierarchical Reasoning Model (HRM) architecture, trained from scratch on a curated mixture of 161 datasets. Rather than relying on web-scraped data of unknown provenance, the team assembled datasets with explicit permissions—academic papers, open corpora, and licensed resources. This constraint is tight: permissible data is far smaller and more fragmented than the internet-scale pools most frontier labs use.
The key insight is that diversity and curation can partially compensate for data size. By mixing 161 sources, the model sees varied reasoning patterns, code examples, and language styles. This breadth appears to help the model generalize despite the total volume being constrained by ethical sourcing requirements.
The model was tested across 20 benchmarks spanning English, mathematics, coding, and Danish. Performance was evaluated against established baselines and contemporary peers—the 1B HRM-Text baseline, as well as larger commercial models like Qwen 3.5 4B.
Why This Matters
For research accessibility: Many researchers and smaller organizations cannot or will not use datasets obtained through aggressive web scraping or terms-of-service violations. Mimir v1 demonstrates that competitive performance is achievable under these constraints, lowering barriers to entry.
For language diversity: Most frontier models focus on English. Mimir v1's state-of-the-art Danish performance hints at a path forward for other low-resource languages. If permissible data strategies work for Danish, they may work for other under-served languages with committed communities.
For reproducibility and governance: Models built on openly licensed data are easier to audit, reproduce, and understand. There's no hidden debt from undisclosed training data. This matters for organizations subject to data governance regulations or concerned about intellectual property risks.
The trade-off: The authors' achievement doesn't mean permissible data is free. It requires careful curation, sourcing from academic repositories and licensed corpora, and accepting that you won't have billions of unlabeled web pages. But for many use cases—especially in regulated industries or non-English contexts—that trade-off is worth it.
Mimir v1 is released on Hugging Face, signaling the authors' commitment to transparency and community reuse. The result is incremental, not revolutionary—the model isn't larger or faster than alternatives, and permissible data remains a constraint. But it resets the conversation: frontier performance and ethical sourcing are not mutually exclusive.
Mine your own.
Lode is a workbench, not a feed. Paste a YouTube URL. The model proposes a transcript, a set of quote-grounded snippets, a synthesis essay, and the fan-out. You decide what stays.