Knowledge recovered
53.00 vs 50.14 on the matched 700-question development set, alongside a 1.42-point gain on the pooled knowledge composite.
Noema model family
More capable. More reliable. Still completely local. An open 2B model with recovered knowledge, stronger code, precise instructions, and shorter reasoning.
Why 1.5 exists
The previous Noema release developed useful strengths in mathematics, code, and structured responses, but paid a measurable price in broad knowledge. Noema 1.5 was trained specifically to recover that deficit without giving up the capabilities that made Noema useful. The result is a compact model with better knowledge retention, stronger code generation, more precise instruction following, improved multi-turn consistency, and shorter reasoning traces. No retrieval, external tools, or inference-time answer repair were used in the reported results; the improvements live in the weights.

Download Noema 1.5 into the Noema model library, choose a device-aware runtime preset, and keep every prompt and response on hardware you control.
The new balance
Matched development gates show progress from Noema 2B; the final lockbox compares the frozen release with stock Qwen3.5-2B.
53.00 vs 50.14 on the matched 700-question development set, alongside a 1.42-point gain on the pooled knowledge composite.
72.09 vs 65.06 for stock Qwen3.5-2B on strict prompt-level instruction following; the paired result is statistically significant.
Pass@1 increased from 50.00 to 53.66 over Noema 2B and remained favorable against the stock foundation in final testing.
Relative increase in conversations satisfying every tested turn: 17.88% for Noema 1.5 vs 14.77% for stock Qwen3.5-2B.
Mean completion length fell from 7,969 to 6,895 tokens while the thinking-accuracy point estimate remained favorable.
The verified mathematics development gate moved from 78.00 to 92.00 over the previous Noema release.
Two frozen comparison stages
Percentages unless noted · identical settings within each comparison
Final IFEval prompt-strict improvement over stock Qwen3.5-2B; statistically significant at p = 0.000475.
Relative lift in complete three-turn Multi-IF conversations over the stock model.
Shorter mean thinking traces with a favorable accuracy point estimate in the final evaluation.
Matched development gates
Paired development and selection results against the exact Noema 2B starting checkpoint. These gates informed model selection; they are not a third arm in the untouched final evaluation.
One-time final evaluation
The frozen Noema 1.5 candidate and stock Qwen3.5-2B used identical prompts, generation settings, and graders. MMLU-Pro was the preregistered knowledge-retention endpoint; point estimates are reported even when they were not statistically resolved.
Publisher-reported model context
Noema 1.5 occupies a strong middle ground: balanced knowledge and science reasoning, strict instruction following, and a compact package designed for local deployment.
Protocol noteDirectional context, not a leaderboard. External scores are reported by each model publisher and were not produced in Noema's controlled harness. Prompt templates, reasoning modes, sampling, token budgets, benchmark revisions, quantization, and graders may differ. Only stock Qwen3.5-2B was evaluated under Noema's paired protocol above.
The open sub-2B field
Published results from current openly available models below two billion total parameters. The protocol label beside each model is part of the comparison.
ReadoutNoema reports stronger MMLU-Pro and GPQA-Diamond results than LFM2.5-1.2B Thinking and narrowly exceeds EXAONE's non-reasoning results. EXAONE leads both knowledge measures in reasoning mode, while LFM leads the instruction-focused result.
Selected larger edge models
Selected publisher results from larger local and edge-oriented models, shown on the same score scale for easier comparison.
ReadoutNoema's reported MMLU-Pro result slightly exceeds Phi-4 Mini's. Gemma, Ministral, Nemotron, and Qwen retain advantages on their published measures.
Evaluation discipline
The final comparison was run once after the candidate was frozen. No retrieval, external tools, answer repair, or outside model calls were used to produce benchmark answers.
Apache 2.0 / Open weights
Open weights under Apache 2.0. Download the official model and run it privately in Noema or another compatible local runtime.
Download Noema 1.5 2B