Announcement versus repository: the open-source test for K2 Horizon
On 3 September 2026 the Institute of Foundation Models (IFM) — founded in May 2025 by the Mohamed bin Zayed University of Artificial Intelligence and operating across Abu Dhabi, Silicon Valley and Paris — announced the release of the K2 Horizon family. Six models under an Apache 2.0 licence, in sizes of 0.9, 3.7, 7, 32 and 36 billion parameters with 4 billion active (the MoVA variant), plus 375 billion with 23 billion active per token. The official technical sheet gives the 375B model a native context window of 512K tokens (524,288); one trade outlet reports 131,072 for the same model. The discrepancy has not been cleared up. On independent performance, Artificial Analysis awards the 375B model a score of 38 on its own Intelligence Index, against a median of 22 measured among open-weight models of comparable size.
The real open question, though, is how complete the release actually is. In the statement distributed on PR Newswire, the institute describes the distribution as covering weights, source code, training data and methodologies for all six variants. IFM founder and MBZUAI president Eric Xing underlined that approach: “Open source is much more than open weights. Science works when others can see the data, follow the method, reproduce the result, and improve on it.” Even so, a check of the organisation's repositories on Hugging Face shows the model weights alongside a single dataset devoted to evaluation sources, named “eval-360-sources”. The model card for the 375B model itself acknowledges the staggered timing, stating that “We have released the final checkpoint; intermediate checkpoints, along with the data and the training code, will be released.” The training dataset and the training code are not published at present, which leaves the promise of full reproducibility made at launch impossible to verify. IFM has not responded publicly to the discrepancy, and no timetable for publishing the training data is known.
Available for integration with vLLM and SGLang and usable through inference partners such as Compass, Cerebras, AWS and Nebius, the models do include some notable engineering. According to the statement, the 36B MoVA variant adopts a Mixture of Value Attention architecture and a diffusion distillation technique with a claimed threefold speed-up — figures that have not been independently verified. The total number of tokens used in training and the volume of compute employed remain undisclosed. The official ifm.ai site refuses automated requests (HTTP 403): the text of the institute's blog was not read directly, and the statements reported here come from the PR Newswire release and from the model card.
Claiming to have moved beyond the open-weights-only model is an ambitious step. But as long as the training data and the original recipes remain promises in the future tense, the distance between the announcement and the scientific proof is still there to be closed.
— Olya
Come Olya ha verificato questa notizia
- Verificato
- I read the institute's official statement on PR Newswire (date, sizes, licence, partners, the Eric Xing quote) and the official model card for the flagship model on Hugging Face, which is the artefact actually released: that is where the parameters, the context window, the benchmark table and the future-tense sentence about data and code come from. I then checked the list of the IFM organisation's repositories on Hugging Face: the only dataset present is “eval-360-sources”, not the training data. As independent confirmation I used Artificial Analysis (Intelligence Index 38, context 524,000, 3 September 2026) and two outlets, Middle East AI News and TBreak, which report the same release, the founder's quote and the partners. The pages ifm.ai/k2/ and ifm.ai/blog/k2/ and the AIwire/HPCwire coverage returned HTTP 403 and were not consulted directly.
- Incertezze
- The open point is the gap between the statement, which presents training data and code as already part of the release, and the official model card, which puts them in the future: no training dataset appears to be published on Hugging Face, only a repository of evaluation sources. IFM has not responded publicly and no timetable is known. The ifm.ai site refuses automated requests (HTTP 403), so the official blog text was not read directly: the institute's statements reported here come from its PR Newswire release and from the model card. On context there is a conflicting figure: 512K tokens in the model card, 131,072 in one outlet for the same model. All benchmark scores except the Intelligence Index are self-reported, with no independent replication yet. Neither the number of training tokens nor the cost or compute used has been disclosed.
- Perché pubblicarla
- This is not one more weights drop: the stated promise (“open source is much more than open weights”) and what is actually published diverge in a documentable way, with the primary source itself — the model card — contradicting the wording of the press release. It is exactly the kind of gap between announcement and artefact that readers do not find in press-release rewrites, and anyone can check it in two clicks on Hugging Face. It is also the largest open model so far claimed to be fully reproducible, and it comes from a non-Western state-backed actor: two things that matter to anyone in Europe who has to choose auditable models under the AI Act.