Multimodal scene input
Atlas takes text, images, video, and 3D as input.
Audio and video / product dossier
Generates camera-controlled 1440p video from text, images, video, or 3D inputs.
Product brief
Atlas is World Labs' omni world model. It takes text, images, video, and 3D, then generates camera-controlled 1440p video up to a minute, reconstructs scenes from a few photos, and simulates space-time for robotics. Early access.
Why we selected it
The multimodal world-model framing and camera control make this more technically distinct than a standard text-to-video generator, though access is early.
Product preview saved with our daily selection
Capability scan
Only capabilities supported by the product information we collected are listed here.
Atlas takes text, images, video, and 3D as input.
It generates camera-controlled 1440p video up to one minute long.
It can reconstruct scenes from a few photos.
The description states that Atlas simulates space-time for robotics.
Best-fit use cases
FAQ
Atlas takes text, images, video, and 3D.
It generates camera-controlled video at 1440p, up to one minute long.
Yes. The description says it reconstructs scenes from a few photos.
Yes. It is described as simulating space-time for robotics.
Atlas is listed as early access.