HiDream-O1-World Tops WBench Navi at 80.9, Betting on Persistent 3D Scenes
HiDream.ai’s new model is designed to make generated environments hold together through movement and edits. Its strongest disclosed evidence is a navigation benchmark; its larger film, robotics and production ambitions remain proposed uses.
Listen to this story
The audio brief
Story brief
3 key pointsHiDream.ai’s HiDream-O1-World launch puts scene persistence—not just initial image quality—at the center of interactive 3D generation. The model uses memory-based 3D priors and test-time adapter updates to preserve objects and relationships during navigation, editing, and interaction. Its strongest disclosed evidence is an 80.9 WBench Navi score, including 88.0 for consistency and 73.3 for physical performance. That...
- 01
WBench Navi covers 289 multi-turn cases and 1,058 interaction turns, making the result relevant but narrower than full-world simulation.
- 02
HiDream.ai claims first place on Navi and says it outperformed Tencent Hunyuan 1.5.
- 03
The model supports text, room-photo, and interactive inputs, including digital-twin creation and first- or third-person exploration.
HiDream.ai has released HiDream-O1-World, a model intended to turn text prompts, images, or interactive controls into explorable 3D worlds. The company is betting that retaining a coherent scene through camera movement and user actions will matter as much as generating the initial view.
Launched August 24, the model combines roaming, editing and interaction on HiDream.ai’s Unified Transformer, or UiT, architecture. HiDream.ai says users can move through a generated world in first- or third-person views, direct character actions, and trigger environmental changes such as rainfall.
The input options point to different creation paths. A short description can generate a world from scratch, while a single room photo can be used to construct a digital twin with a completed panorama, according to HiDream.ai. The company also says the model can work across real-world scenes, natural terrain, anime-style imagery and game-style rendering.
A scene record instead of a fresh redraw
The company’s technical contrast is between a system that can lose objects or drift as the camera changes and one that retains the scene’s structure. HiDream.ai says its Memory context encodes 3D priors about geometry and object relationships, while test-time training makes lightweight adapter updates during inference to keep representations aligned with the scene’s constraints.
That approach is meant to address three related failure modes: camera changes that blur or reset a scene, physical behavior that looks implausible, and objects that disappear after leaving the frame. Rather than generate each view without retained structure, HiDream.ai says the system carries geometry and relationships forward as the viewpoint changes.
HiDream.ai reported that HiDream-O1-World ranked first on WBench’s Navi sub-leaderboard with an average score of 80.9.
The company reported a Physical score of 73.3 on the Navi evaluation.
HiDream.ai reported a Consistency score of 88.0.
A navigation lead, not a complete deployment case
WBench was developed jointly by Meituan’s LongCat team and Fudan University. It evaluates 289 multi-turn interaction cases and 1,058 interaction turns across five dimensions and 22 metrics; its Navi board focuses on spatial navigation and viewpoint control. That focus makes the result relevant to HiDream-O1-World’s scene-persistence pitch.
HiDream.ai says the model topped Navi on its first entry and outperformed Tencent Hunyuan 1.5 there. But a leaderboard centered on navigation and camera control tests a narrower proposition than a system’s claimed ability to simulate varied physical events or serve as a production environment.
The separate physics claim
- HiDream.ai says its training data includes simulated rigid-body collisions, fluid motion, soft-body deformation and gravity-driven projectiles.
- The company says test-time training also adapts to new objects or materials during inference, and reports gains of 13.6% in visual plausibility and 12.7% in causal fidelity over an unspecified industry average.
- The company says its character handling extends across people, animals and fictional characters, adjusting motion to each form.
From interactive scenes to proposed simulators
Holding a layout steady as a viewpoint moves and producing plausible collisions or deformation are related, but distinct, demands. HiDream.ai proposes the model for interactive film and games, embodied-AI simulation, and 3D scene production; those uses extend beyond the Navi benchmark’s stated navigation focus. The associated DreamWorld research was accepted to ECCV 2026.
The company’s proposed simulation use is especially demanding: it describes virtual cities, factories and interiors as testbeds for robotics, autonomous driving and smart manufacturing. The launch materials frame those settings as alternatives to costly or risky real-world testing, but the disclosed benchmark result is specifically about interactive navigation rather than those downstream applications.
Sources
- vir.com.vnHiDream.ai releases HiDream-O1-World interactive artificial intelligence model