Garp Independent AI & technology journalism
Wednesday, July 22, 2026 Sign In · Join Subscribe
Latest Anthropic’s landmark $1.5B copyright settlement is approved

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  Robotics  /  Xiaomi-Robotics-1 shows that more data beats bigger models when training robots to move

Robotics

Xiaomi-Robotics-1 shows that more data beats bigger models when training robots to move

Xiaomi-Robotics-1 shows that more data beats bigger models when training robots to…

Xiaomi has released an AI model for robots that follows the same scaling pattern as large language models. Its performance improves as it trains on more data.

To build that dataset, Xiaomi largely avoided using physical robots. Robot AI faces a data problem that language models don’t: LLMs can train on huge parts of the public internet, while useful data on robot movement is scarce. Robots usually have to learn from scratch how to grip, lift, and put away objects. The standard method has people remotely guide a physical robot through every movement. That process is slow and expensive, and it often produces repetitive data from the same tasks in the same settings. That’s the gap Xiaomi-Robotics-1 is trying to close. The model is designed to follow spoken or written commands in unfamiliar environments without prior exposure and adapt to new tasks with little extra training. To get around the data bottleneck, Xiaomi mostly ditched real robots during data collection. Instead, the team used portable handheld grippers with attached cameras that a person simply picks up and operates by hand. This setup lets you record manipulation tasks in kitchens, offices, stores, factory floors, and outdoor spaces without a robot even being present. The result was over 100,000 hours of motion recordings. A dataset that large creates another problem because each recording needs a description the model can learn from. Labeling it all by hand wasn’t practical, so Xiaomi used another AI model to describe each motion segment in text. The team says it labeled the full dataset in about two weeks. Xiaomi then transferred that training to physical robots, including wheeled models and dual-arm systems.