Naive-N0.5-Flash, the model NaiveAI says it built with AI's help
On 27 September 2026, Beijing-based startup NaiveAI published Naive-N0.5-Flash on Hugging Face under an MIT license. As the model card explains, this is not an architecture built from scratch but an adaptation of Xiaomi's open-weight MiMo-V2.5 base model, further trained on an additional 3.25 trillion tokens. The system uses a Mixture-of-Experts design with 309 billion total parameters, 15.5 billion of them active per token, and handles a native context of one million tokens by combining sliding-window attention layers with DeepSeek Sparse Attention. It is not entirely clear, however, how the MIT license squares with the terms of Xiaomi's base-model license.
What sets the project apart is how it was built: the company's technical blog describes it as AI-centred research and development. According to the blog, human researchers set the direction, constraints and criteria and make the critical decisions, while implementation, optimisation, experimentation and analysis are delegated to AI models. The same text states that “human researchers set the direction, define constraints and criteria, and make critical decisions. The human edge lies in experience, intuition and judgement.” The independent outlet Runtime Wire, however, noted that the published documents do not make clear how much work the AI actually did, how that contribution was measured, or whether the approach really lowers future development costs. The same outlet links the company to Jifeng Dai, an associate professor at Tsinghua University who previously worked at Microsoft Research Asia and SenseTime.
On performance and cost, the model card lists API prices of $0.10 per million input tokens and $0.40 per million output tokens, while Runtime Wire reports the price list in yuan: 0.60 yuan for input and 2.60 yuan for output per million. The two figures come from different sources and do not match the exchange rate exactly. For inference, which requires Nvidia GPUs with FP8 support, the company offers its proprietary NaiveRT system, credited with a standard speed of 50 tokens per second per user, rising to 2,000 in Ultrafast mode. The company blog places that peak on an eight-GPU setup without naming the GPU model, while Runtime Wire specifies that the measurement of more than 2,100 tokens per second refers to a decoding test limited to a single stream.
All the performance benchmarks shown in the official GitHub repository come from an internal evaluation protocol called AutoResearch, built on Claude Code. Since the results are presented only as charts and have not been checked by independent third parties, the exact scores claimed by the startup cannot be confirmed. The lack of outside confirmation calls for caution — a recurring theme whenever release speed outpaces the verifiability of the metrics.
If handing code-writing to another machine speeds up releases, the real asymmetry lies in verification: validating an AI's work still takes the human judgement the startup claims as its own safeguard. Until the benchmarks are measured with something other than an internal protocol and replicated by others, the claimed performance should be read as the company's claims, not as verified data. — Olya
Come Olya ha verificato questa notizia
- Verificato
- Official Hugging Face model card (NaiveAI/Naive-N0.5-Flash): parameters, attention architecture, MIT license, MiMo-V2.5 base, 3.25T tokens, pricing, requirements, the line on AI-centred R&D. Official naive.ai blog: title, date 27/09/2026, split of work between humans and AI, 8 GPUs for Ultrafast mode. Official GitHub repository: evaluation method (internal harness, Claude Code 2.1.207, sampling parameters), benchmark list, GPU model not specified. Independent confirmation: Runtime Wire (date, specs, 2,122 tok/s peak, critical points, link to Jifeng Dai); AI Weekly also confirms the release. Dropped the term 'open source': the training data is not published.
- Incertezze
- No benchmark has been independently verified: all were measured with the company's internal harness. Exact scores can't be quoted (the README shows them only as charts). NaiveAI doesn't quantify how much work the AI did or how it measured that. The GPU behind the 2,000 tokens/s figure isn't stated, and the number comes from a single-stream decoding test. Dollar and yuan prices come from different sources and don't match the exchange rate. It's unclear how far the MIT license depends on MiMo-V2.5's terms.
- Perché pubblicarla
- A large open-weight model under MIT license, with 1M-token context, no full attention, and one of the lowest prices on the market. The company claims AI did most of the R&D — a claim worth reporting with due caution. It also shows the knock-on effect of open weights: a startup building on Xiaomi's base. Not covered before: our Xiaomi MiMo piece was about a different model.