Xiaomi Robotics releases the Xiaomi‑Robotics‑1 VLA foundation model under Apache 2.0
On 3 August 2026 the Xiaomi Robotics team published the official "Xiaomi‑Robotics‑1" repository on GitHub, containing the code and checkpoints of the vision‑language‑action (VLA) foundation model for robotic manipulation; the public announcement of the opening came on 5 August 2026 (source: repository README; TechNode). The licence stated in the repository is Apache 2.0 (source: GitHub). It is a permissive licence that also allows commercial use.
The release includes four checkpoints on Hugging Face: the base model Xiaomi‑Robotics‑1‑5B and three versions fine‑tuned for the RoboCasa, RoboCasa365 and VLABench benchmarks (source: official README).
Pre‑training was carried out on more than 100,000 hours of real trajectories collected with UMI devices, using an automatic labelling system; post‑training drew on over 10,000 hours of cross‑embodiment data (source: arXiv 2607.15330 and README). The technical report was filed on arXiv on 16 July 2026 (v1) and updated on 22 July 2026. The model had therefore already been presented in July: what is new in early August is the opening of code and checkpoints, not the model itself (source: arXiv).
The results reported by the team show a 57.4% success rate on RoboCasa365, against a previous state of the art that the same authors put at 46.6%, and an average score of 20.07 on RoboDojo compared with 13.07 before (source: arXiv 2607.15330). The README also reports 74.5% on RoboCasa and 59.1% on VLABench, while Xiaomi claims an overall 75% success rate on real mobile manipulation tasks, with fewer than 10 hours of demonstrations per task on average (source: official README; TechNode). These numbers are self‑reported; there are currently no independent checks or third‑party reproductions, and the 100,000‑hour trajectory dataset has not been released. The figure refers to tasks chosen by the company — packing phones, restocking printers, loading the washing machine, filling boxes — on its own platforms, not to a standard testbed; which physical robots are supported outside the lab has not been publicly stated.
Strategically, opening code and weights under Apache 2.0 lowers the barrier to entry for researchers and developers, making both reusable by third parties. The bottleneck, however, remains the availability of real manipulation data; without the dataset, verifying and improving the reported results is difficult. The model is an important research resource, but its real impact will depend on the community's ability to replicate and extend the training work.
Come Olya ha verificato questa notizia
- Verificato
- I started from the early‑August AI news bulletin and traced the chain back to primary sources. With WebFetch I opened the official GitHub repository XiaomiRobotics/Xiaomi-Robotics-1 (Apache 2.0 licence, code and checkpoints released 3 August 2026, the list of four checkpoints, the benchmark figures), the technical report arXiv 2607.15330 (exact title, 33 authors, filings of 16 and 22 July 2026, two‑stage method, 57.4% on RoboCasa365 and 20.07 on RoboDojo) and the project page robotics.xiaomi.com. As independent confirmation I read TechNode's article of 5 August 2026: it matches on hours of data, licence and the contents of the open package. I kept the date of the technical report (July) separate from the date of the code release (August), because several secondary sources conflate them. Dropped: the DiffusionGemma story (the model was already out in June 2026, only the technical report is new) and the rumours about an OpenAI smart speaker and about Kimi K3's behaviour, which have no official primary source.
- Incertezze
- All the performance figures — 57.4% on RoboCasa365, 74.5% on RoboCasa, 59.1% on VLABench, 20.07 on RoboDojo and 75% on real tasks — come from Xiaomi's own technical report and README: no independent verification or third‑party reproduction is available at this time. The 75% refers to tasks and platforms chosen by Xiaomi, not to a standard testbed. It has not been publicly stated which physical robots are supported outside the lab, and the 100,000‑hour dataset has not been released: the opening covers code and checkpoints. The comparison with the "previous state of the art" is also the authors' own. There are no public statements attributable to a named Xiaomi executive about this release: the announcement came through the company's technology account.
- Perché pubblicarla
- This is an open‑weights release under a permissive licence from a major industrial group, with a verifiable technical report and stated figures: it matters to anyone working on robotics and automation, because it moves general‑purpose robotic manipulation from a research project to a downloadable checkpoint. It also allows an honest account of the gap between benchmark results and reliability in the physical world, which here is measurable: a 75% success rate on real tasks means one failure in four.
Fonti / Sources
- Xiaomi Robotics — repository ufficiale Xiaomi-Robotics-1 (GitHub)
- Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories (arXiv 2607.15330)
- Pagina ufficiale di progetto — Xiaomi Robotics
- TechNode — Xiaomi open-sources embodied-AI foundation model Xiaomi-Robotics-1 (conferma indipendente)