Google Deepmind argues video generators already contain the world models computer vision has been missing
A new model from Google Deepmind called GenCeption uses a pre-trained video generation model as the basis for classic computer vision tasks. It achieves state-of-the-art performance in depth estimation, segmentation, and 3D pose estimation while needing very little training data.
Language models became versatile processing systems almost as a byproduct of learning to predict the next word. That task appears to require models to absorb grammar, world knowledge, and contextual relationships during training. This is the prevailing explanation for the emergent capabilities of large language models. Computer vision still lacks an equivalent training method. Specialized models dominate the field, including “Segment Anything” for segmentation and “Depth Anything” for depth estimation. Each uses its own architecture.Ad In a new paper, Deepmind researchers argue that large text-to-video models could bridge this gap. Generating realistic videos requires understanding the spatial geometry of a scene, how objects move, and basic physics. These models also use text descriptions during training, which links language to visual content. They can train on vast datasets because labeling the data is relatively cheap.AdDEC_D_Incontent-1 GenCeption repurposes a video generator for vision tasks GenCeption builds on Alibaba’s open source Wan2.1 video model. The main change is a simpler architecture. Diffusion models typically generate video from noise through many small steps. GenCeption instead produces a prediction in one forward pass, making it fast enough for practical computer vision tasks. The researchers use a simple method to make one model handle many tasks. GenCeption represents every output as a standard three channel RGB image, whether the result is a depth map, surface normal map, or segmentation mask. It also converts camera movement into an image representation.Ad A text prompt tells the model which task to perform, much like an instruction given to a chatbot.