Vidu releases Q4 video preview with voice references and 4K output
The web and API release gives creators controls for characters, voices and camera work. Its performance improvements remain company claims, and two public sources quote different starting prices.
Released October 7, 2026, Vidu Q4 Preview gives creators more ways to steer a scene before generation: Reference-to-Video accepts up to 15 images and three audio references, while Image-to-Video animates a single image. ShengShu documents API access and output options from 540p to 4K, though the default image-to-video result is five seconds at 720p and clips top out at 16 seconds. Performance claims are unverified, and published starting prices conflict: $0.014 versus $0.045 per second, with no explanation for the difference.
01
ShengShu is offering a 30% discount on both generation modes through November 30, 2026.
02
The reference-to-video API accepts MP3 voice samples lasting 3–12 seconds and supports five aspect ratios.
03
ShengShu also advertises 10-bit color, automatic camera switching, and consistency across multiple camera positions; these are documented capabilities, not guarantees.
ShengShu Technology released Vidu Q4 Preview on October 7, 2026, giving video creators a new model through its web platform and developer API. The company says it combines voice references, up to 15 reference images and 4K output, aiming to give production teams more control over performances rather than just sharper-looking clips.
ShengShu is pitching the release to short-drama producers, advertising agencies, social creators and film teams. Its launch announcement promises more natural emotional delivery, improved lighting and steadier camera moves. Those are company claims, not independent performance findings; the more concrete change is the set of inputs creators can use to direct a scene.
Unite.AI’s account of Vidu’s product page and documentation describes two generation modes. One starts from a single image. The other accepts a collection of images and optional audio, giving creators separate reference material for subjects and voices. The distinction changes what a creator supplies before generation begins.
Image-to-Video animates one uploaded image using a text prompt. Listed clip lengths run from 3 to 16 seconds.
Reference-to-Video accepts one to 15 images and zero to three audio references. The product page lists durations from 1 to 16 seconds.
The image references are intended to anchor characters, clothing, props, products and environments. ShengShu says reference-voice support keeps a character’s voice consistent while allowing emotional delivery. Its claimed performance improvements also cover the coordination of facial expressions, body movement and speech—not just the appearance of individual frames.
For fast action, ShengShu says the camera follows fights and chases more coherently, with cuts better coordinated with movement. It also claims explosions, fireworks and particles blend more naturally into their surroundings. These promises concern how a scene moves and how effects interact with it, rather than resolution alone.
For developers, the API—a way for software to request generation—identifies the model as viduq4-preview. Unite.AI reports that its image-to-video endpoint accepts a single starting image in PNG, JPEG, JPG or WebP format, up to 50MB. Prompts can contain up to 20,000 characters, while the default output is a five-second clip at 720p.
Audio generation is enabled by default for that endpoint and can include dialogue and sound effects. The reference-to-video endpoint accepts MP3 voice references lasting 3 to 12 seconds each. It also offers five aspect ratios, including square, vertical and widescreen formats, with 16:9 as the default.
Reference-to-video prompts can include tags for the supplied subjects. The documentation also lists automatic camera switching, consistency across multiple camera positions and simultaneous audio-video output. These are documented capabilities, not a measured guarantee that every generated cut will preserve continuity.
Output options span 540p, 720p, 1080p, 2K and 4K. ShengShu also advertises 10-bit color for professional workflows, pairing its performance pitch with a higher-end output specification.
The maximum listed video length is 16 seconds, even though output reaches 4K. Returned creation links remain valid for 24 hours. ShengShu is releasing this preview ahead of the full Q4 model and says feedback on performance, creative control and production needs will shape the final release.
The pricing deserves a closer look than the launch headline suggests. Unite.AI quotes a lower starting figure of $0.014 per second, while ShengShu’s own release gives the resolution-based rates above. The outlet also says final pricing, supported resolutions, features and usage terms may vary by plan and region.
The reason for the difference between those starting prices is unresolved. Separately, ShengShu offers a 30% discount on image-to-video and reference-to-video generation through November 30, 2026. That is a time-limited promotion for the two modes, not a resolution of which starting price applies to a particular account or generation request.
Editorial illustration for Vidu releases Q4 video preview with voice references and 4K output.
Sources
globenewswire.comShengshu Technology Launches Vidu Q4 Preview, A Next-Generation AI Video Model Built for Lifelike Performances
unite.aiVidu Releases Q4 Preview of Next-Generation Flagship AI Video Model
Reader comments
Newest comments first. Replies stay oldest first.