NVIDIA open source SANA-WM: 2.6B parameter world model, generates 1-minute 720p video from a single image
NVIDIA's open source SANA-WM can generate up to 1 minute of 720p continuous video from a single image with only 2.6B parameters, supporting precise 6 degrees of freedom camera control. The distilled version can be generated in 34 seconds on a single RTX 5090, which is 36 times faster than comparable models. This open source world model will bring new possibilities in areas such as video generation, game development, and autonomous driving simulation.
Core conclusion
SANA-WM is NVIDIA's latest open source efficient minute-level world model. It can generate 720p high-definition video of up to 60 seconds with only 2.6B parameters, and supports precise 6-degree-of-freedom camera trajectory control.
Key Points
- Number of model parameters: only 2.6B (an order of magnitude smaller than similar models)
- Generation capability: single image input → 1 minute 720p continuous video output
- Training efficiency: only 213K public video clips, completed in 15 days on 64 H100
- Inference speed: the distilled version completes 60-second video generation in 34 seconds on RTX 5090
- License: Open Source (Apache 2.0), weights to be released soon
- Paper link: arXiv 2605.15178
Background: Current status and challenges of world models
World Model is a key frontier in the field of AI - it can simulate the temporal evolution process in the physical world and is crucial for scenarios such as autonomous driving, robot control, game development, and video generation.
However, there are three major bottlenecks in the current mainstream world model:
- Parameter scale is too large - Many industrial-grade models exceed 10B or even 50B parameters, and the deployment threshold is high
- Slow inference — It takes several minutes or even dozens of minutes to generate a 1-minute video
- Inaccurate camera control — Unable to achieve accurate trajectory tracking
SANA-WM simultaneously solves the above three problems through four core design innovations.
Four core architecture innovations
SANA-WM's architecture revolves around four key designs:
| Design | Function | Value |
|---|---|---|
| Hybrid linear attention | Frame-level Gated DeltaNet + softmax attention | Memory-efficient long-term context modeling |
| Dual-branch camera control | 6 degrees of freedom trajectory precise following | Video direction/viewing angle accurately controllable |
| Two-stage generation pipeline | Long video refiner processes first-stage output | Improves sequence quality and consistency |
| Robust annotation pipeline | Extract metrically accurate 6-DoF camera poses from public videos | High-quality spatiotemporally consistent action labels |
Efficiency comparison: SANA-WM vs industrial grade model
| Indicators | SANA-WM (2.6B) | LingBot-World | HY-WorldPlay |
|---|---|---|---|
| Number of parameters | 2.6B | ~10B+ | ~10B+ |
| Training data | 213K public clips | Large-scale private data | Large-scale private data |
| Training Compute | 64 H100 × 15 days | More GPUs × longer time | More GPUs × longer time |
| Output quality | Comparable | Baseline level | Baseline level |
| Throughput | 36× Higher | 1× | 1× |
| Camera Control | ✅ 6-DoF Precise Control | Partially Supported | Partially Supported |
| Open Source | ✅ Apache 2.0 | ❌ Closed Source | ❌ Closed Source |
Why is this important?
SANA-WM is much more than just "another video generation model."
First, the world model enters the "consumable" stage. The 2.6B parameters mean it can run on consumer-grade GPUs—the distilled version produced a 1-minute video in 34 seconds on an RTX 5090. This is completely different from the previous solution that required several H100 clusters.
Second, open source breaks the monopoly. The current minute-level world models are almost entirely monopolized by closed-source industrial solutions. SANA-WM is open source under the Apache 2.0 license, which means that individual developers and small and medium-sized teams can also use world model capabilities.
Third, verification of training efficiency. Achieving comparable quality using only 213K public video clips and 15 days of training proves that efficient world models do not require massive amounts of private data - a huge boon for academic research and the open source community.
Adaptation suggestions
For content creators and automation pipeline developers:
- Pay attention to the weight release — NVIDIA said that the weight is "coming soon" and it can be deployed locally for testing once it is officially released.
- Evaluate integration costs — 2.6B parameters + NVFP4 quantification means that a single card RTX 5090 can run, and the integration threshold is extremely low
- Explore automated video generation pipeline — Connect SANA-WM to n8n or customize workflow to realize automated production of images → videos
- Focus on Camera Control API — 6-DoF precise control is the core advantage of SANA-WM and can be used for automated trajectory generation
Reference videos and materials
Tool entry
Tool name that appears naturally in the text: NVIDIA
Internal link guidance
- Want to learn local deployment? See: How to run local AI models on M4 Mac with LM Studio: A complete 30-minute tutorial
- Real case: He Built an AI Automation Stack with Claude + n8n — $4K to $12K/mo in 6 Months
Topic hub
YouTube AI Content Policy Hub
Answer-style evergreen hub for AI labels, auto detection, and disclosure—not just breaking news.
Explore YouTube AI Content Policy Hub →Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
ChatGPT ads + content distribution
Sell compliance checklists and automated distribution
OpenClaw Agent short-video growth
Lean into hybrid workflows as labels get stricter