RC RANDOM CHAOS

IFM ships K2 Horizon: six open models from 0.9B to 375B, full training lifecycle included

· via Hacker News

Original source

K2 Horizon: A connected fleet of six open models

Hacker News →

IFM has released K2 Horizon, a family of six models sharing one architecture, vocabulary, and training methodology: a 375B-A23B sparse MoE flagship, a dense 32B, a 36B-A4B sparse model, and 7B, 3.7B, and 0.9B models aimed at phones, wearables, and other edge hardware. All ship with quantization support, and the smallest three are claimed as state-of-the-art in their size classes. The 36B-A4B model uses a new sparse attention scheme IFM calls Mixture-of-Value-Attention (MoVA) to approach dense-32B quality while activating only about 4B parameters per token.

The more notable move is transparency rather than raw scores. IFM is releasing not just final weights but the full development tree for each model — intermediate checkpoints, post-training branches, training code, configs, logs, evaluation results, and either the data itself or detailed construction recipes where redistribution isn’t possible. Models and code are Apache 2.0; datasets carry their own licenses (e.g. ODC-BY). The company frames this as the first fully open model fleet to expose the entire agentic post-training pipeline, letting researchers reproduce how reasoning, tool use, and planning emerge rather than treating a checkpoint as an opaque starting point.

The performance claims are vendor benchmarks and should be read as such — figures like a 0.9B model scoring above 48 on AIME 2026, plus SWE-bench and BrowseComp results for the 7B and 3.7B, await independent replication. IFM concedes the smallest models still struggle on long-horizon, recovery-heavy tasks like TerminalBench. The real bet here is that pairing competitive capability with a genuinely reproducible training recipe — extending the fully-open stance from IFM’s 2023 LLM360 work — is more useful to the research community than another set of closed weights.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.