Modelspublished

HiDream-O1-World Tops WBench Navi at 80.9, Betting on Persistent 3D Scenes

HiDream.ai’s new model is designed to make generated environments hold together through movement and edits. Its strongest disclosed evidence is a navigation benchmark; its larger film, robotics and production ambitions remain proposed uses.

By 3 min read
HiDream-O1-World Tops WBench Navi at 80.9, Betting on Persistent 3D Scenes

Listen to this story

The audio brief

About 1:45
0:001:45
Read transcript
HiDream.ai says its new HiDream-O1-World model has taken the top spot on WBench Navi, scoring 80.9 on a benchmark built around interactive navigation. The result is notable because the model is designed to preserve a scene as you move through it, rather than redraw each viewpoint from scratch. HiDream.ai reports 88.0 for consistency and 73.3 for physical performance. WBench Navi covers 289 multi-turn cases and 1,058 interaction turns, so it is meaningful evidence for navigation and viewpoint control, but a narrower test than full-world simulation. HiDream.ai says the model also beat Tencent Hunyuan 1.5 on that leaderboard. The system accepts text, room photos, and interactive controls. It can generate a world, build a digital twin from a single room image, and support first- or third-person exploration. Technically, HiDream.ai describes a Memory context that stores 3D geometry and object relationships, alongside lightweight test-time adapter updates intended to keep those representations aligned as scenes change. The company also reports 13.6 percent higher visual plausibility and 12.7 percent higher causal fidelity than an unspecified industry average. Its proposed uses include games, film, robotics, autonomous driving, and manufacturing, while related DreamWorld research was accepted to ECCV 2026. The key constraint is clear: the strongest disclosed evidence is still navigation. The open question is whether that persistence holds up for demonstrated physics and real downstream deployments.

Story brief

3 key points

HiDream.ai’s HiDream-O1-World launch puts scene persistence—not just initial image quality—at the center of interactive 3D generation. The model uses memory-based 3D priors and test-time adapter updates to preserve objects and relationships during navigation, editing, and interaction. Its strongest disclosed evidence is an 80.9 WBench Navi score, including 88.0 for consistency and 73.3 for physical performance. That...

  1. 01

    WBench Navi covers 289 multi-turn cases and 1,058 interaction turns, making the result relevant but narrower than full-world simulation.

  2. 02

    HiDream.ai claims first place on Navi and says it outperformed Tencent Hunyuan 1.5.

  3. 03

    The model supports text, room-photo, and interactive inputs, including digital-twin creation and first- or third-person exploration.

HiDream.ai has released HiDream-O1-World, a model intended to turn text prompts, images, or interactive controls into explorable 3D worlds. The company is betting that retaining a coherent scene through camera movement and user actions will matter as much as generating the initial view.

Launched August 24, the model combines roaming, editing and interaction on HiDream.ai’s Unified Transformer, or UiT, architecture. HiDream.ai says users can move through a generated world in first- or third-person views, direct character actions, and trigger environmental changes such as rainfall.

The input options point to different creation paths. A short description can generate a world from scratch, while a single room photo can be used to construct a digital twin with a completed panorama, according to HiDream.ai. The company also says the model can work across real-world scenes, natural terrain, anime-style imagery and game-style rendering.

A scene record instead of a fresh redraw

The company’s technical contrast is between a system that can lose objects or drift as the camera changes and one that retains the scene’s structure. HiDream.ai says its Memory context encodes 3D priors about geometry and object relationships, while test-time training makes lightweight adapter updates during inference to keep representations aligned with the scene’s constraints.

That approach is meant to address three related failure modes: camera changes that blur or reset a scene, physical behavior that looks implausible, and objects that disappear after leaving the frame. Rather than generate each view without retained structure, HiDream.ai says the system carries geometry and relationships forward as the viewpoint changes.

The disclosed navigation result
80.9Navi average

HiDream.ai reported that HiDream-O1-World ranked first on WBench’s Navi sub-leaderboard with an average score of 80.9.

73.3Physical

The company reported a Physical score of 73.3 on the Navi evaluation.

88.0Consistency

HiDream.ai reported a Consistency score of 88.0.

A navigation lead, not a complete deployment case

WBench was developed jointly by Meituan’s LongCat team and Fudan University. It evaluates 289 multi-turn interaction cases and 1,058 interaction turns across five dimensions and 22 metrics; its Navi board focuses on spatial navigation and viewpoint control. That focus makes the result relevant to HiDream-O1-World’s scene-persistence pitch.

HiDream.ai says the model topped Navi on its first entry and outperformed Tencent Hunyuan 1.5 there. But a leaderboard centered on navigation and camera control tests a narrower proposition than a system’s claimed ability to simulate varied physical events or serve as a production environment.

The separate physics claim

  • HiDream.ai says its training data includes simulated rigid-body collisions, fluid motion, soft-body deformation and gravity-driven projectiles.
  • The company says test-time training also adapts to new objects or materials during inference, and reports gains of 13.6% in visual plausibility and 12.7% in causal fidelity over an unspecified industry average.
  • The company says its character handling extends across people, animals and fictional characters, adjusting motion to each form.

From interactive scenes to proposed simulators

Holding a layout steady as a viewpoint moves and producing plausible collisions or deformation are related, but distinct, demands. HiDream.ai proposes the model for interactive film and games, embodied-AI simulation, and 3D scene production; those uses extend beyond the Navi benchmark’s stated navigation focus. The associated DreamWorld research was accepted to ECCV 2026.

The company’s proposed simulation use is especially demanding: it describes virtual cities, factories and interiors as testbeds for robotics, autonomous driving and smart manufacturing. The launch materials frame those settings as alternatives to costly or risky real-world testing, but the disclosed benchmark result is specifically about interactive navigation rather than those downstream applications.

Sources

  1. vir.com.vnHiDream.ai releases HiDream-O1-World interactive artificial intelligence model