Friction is key to making better robot world models
NVIDIA Corp. — contactile offers robotic hands and grippers equipped with its tactile sensors. | But, because conditioning on touch remains fundamentally incomplete, they cannot reliably generalize across novel surfaces and objects.
A new model class, VμA, proposes to fix that by making friction a first-class input. World models are the next frontier The most ambitious direction in robot learning today is the world model: a generalist model of physical reality that a robot can use to predict the consequences of its actions, plan across long horizons, and generalize to situations it has never encountered in training. If a robot’s internal model of the world is accurate enough, it does not need to memorize every task. Instead, it can reason its way through novel ones. This is a compelling vision, and the field is moving fast. But deploying world models in real robotic systems requires a step that receives less attention than the models themselves: conditioning. A world model must be conditioned on the robot’s current physical state before it can make useful predictions. And the quality of that conditioning determines whether the model’s predictions reflect reality, or merely an approximation of it. The conditioning problem: Touch is missing Current world model conditioning in robotics relies primarily on two inputs. These are visual observations from cameras, and end-effector position from joint encoders. Some of the most capable systems in the field condition on nothing more than this. For free-space motion tasks, it is often sufficient. For contact-rich manipulation, it is not enough. The moment a robot touches an object, the information that matters most — what is happening at the interface between fingertip and surface — is invisible to a camera and unresolvable from joint position alone. In many world model implementations, contact is not even encoded by dedicated sensing. Instead, it is inferred from motor currents in the robot’s joints.