A useful route when a shot needs a defined first/last frame, native audio, or extension of a Veo-generated clip.
- Input
- Text, starting image; up to 3 reference images (standard/Fast)
- Length
- 4, 6 or 8 seconds; references and higher resolutions require 8s
- Output
- 720p, 1080p, 4K; extension is 720p
- Audio
- Native audio, always on
- Controls
- First/last frame; extension on standard/Fast, not Lite
What makes it useful
- Frames help specify where a shot starts and ends.
- Audio is generated alongside the picture.
- Fast and Lite offer different cost/resolution tradeoffs.
Where to pay attention
- Reference, duration and resolution combinations have constraints.
- Extension accepts supported Veo outputs; it is not a general footage editor.
- Preview-model access and regional person-generation rules can change.
A first workflow to try
- Approve a product or character image before animating.
- Request one camera move and explicit sound direction.
- Check the first, middle and last frame; extend only a usable take.
For our miniature-car shot, start with image-to-video and a single dolly-in. Check body geometry before trying dialogue.
Cost & accessGemini API: standard $0.40/s at 720p/1080p; Fast $0.10/s at 720p; Lite $0.05/s at 720p. 4K and other resolutions differ.
Google now recommends Omni Flash as its default video-generation route. Its distinctive workflow is refining a clip through follow-up instructions.
- Workflow
- Video generation and multi-turn editing through the Interactions API
- Output
- 360p or 720p; 1080p and 4K are upscaled
- References
- API video references: up to 3 clips, 3s each; audio ignored
- Editing limits
- Uploaded footage ≤10s; upload editing unavailable in EEA, Switzerland and UK
- Audio
- Generated audio; uploaded audio references and voice editing unsupported in this API
What makes it useful
- Follow-up edits preserve the conversational clip context.
- Can edit a generated result without starting a new prompt from scratch.
- Useful to explore changes in look or a specific visual element.
Where to pay attention
- The general model overview is broader than the current API supports.
- Upload-region and footage-length limits matter for real projects.
- Ask explicitly for a continuous shot; the default can include multiple shots.
A first workflow to try
- Generate a short, simple scene with no cuts.
- Change one thing in a follow-up and ask to keep everything else the same.
- Check that the untouched product, background and timing actually remain stable.
Choose this when iteration on an existing generated shot matters more than a precise first/last-frame pipeline.
Cost & accessCheck current Gemini API pricing for this model. Veo rates do not apply to Omni Flash.
A short-shot model within a broader filmmaking platform. Start with an image when composition is already approved, then describe the motion.
- Input
- Text-to-video or image-to-video
- Length
- 2–10 seconds
- Output
- 720p, 24 or 25fps
- Framing
- Text: 16:9; image: 16:9, 9:16, 1:1, 4:3, 3:4 or 21:9
- Access
- Standard plan or higher; 12 credits per generated second
What makes it useful
- Variable short durations suit a shot-based workflow.
- Image-to-video supports several useful delivery shapes.
- Documented motion-focused prompting is a clear starting point.
Where to pay attention
- Native model output is 720p.
- Other Runway models and tools have different capabilities and costs.
- Premium ProRes/PNG export adds credits and has plan restrictions.
A first workflow to try
- Create an approved first frame with the intended composition.
- Describe camera and subject motion rather than re-describing every pixel.
- Choose the shortest take that covers the action, then finish in your timeline.
A practical starting point for controlled visual inserts and product motion. This is an editorial fit, not a quality ranking.
Cost & access12 Runway credits/s. ProRes/PNG sequence generation adds 5 credits/s on supported premium plans. Credits are not dollars or Pika credits.
Video 3.0 and 3.0 Omni bring shot-based storytelling and native audio into the same family. Omni adds richer reference and storyboard controls.
- Length
- Up to 15 seconds
- Audio
- Native audio; announcement names English, Chinese, Japanese, Korean and Spanish
- References
- Image/video elements; Omni can extract character appearance and voice from video
- Storyboarding
- Omni: per-shot duration, shot size, perspective and camera movement
- Pricing / resolution
- Not verified from a current public primary specification; check the selected host
What makes it useful
- Per-shot direction is useful when the order of scenes matters.
- Character appearance and voice can be referenced together in Omni.
- Multi-language audio opens more options than a silent clip workflow.
Where to pay attention
- Video 3.0 and Omni are different model variants.
- Launch-day access is not evidence of your current plan entitlement.
- Text and identity consistency are provider claims to test on your own content.
A first workflow to try
- Prepare a character reference with the correct voice and rights.
- Write a small sequence with one clear action in each shot.
- Check speaker attribution, continuity and exact text before accepting the take.
Consider Omni for a short character-led sequence with explicit cuts and voices. Keep the initial storyboard simple.
Cost & accessCheck the current Kling plan and exact model. We do not assign an unverified dollar-per-second rate.
A reference-heavy workflow for shaping motion and finishing the image, including HDR and EXR output options.
- Length / output
- Up to 20 seconds at 1080p, per the release announcement
- Keyframes
- Up to 16 keyframes
- Performance
- Skeletal/gesture tracking and facial tracking for up to 8 faces
- Finishing
- Native HDR and 16-bit EXR; reframe and background replacement
- Interface
- Ray3.2 API and application workflows; select the model explicitly
What makes it useful
- Multiple keyframes let you describe a trajectory across the shot.
- Performance controls are relevant to modifying existing footage.
- HDR/EXR support matters when the finishing pipeline needs it.
Where to pay attention
- A more complex pipeline than making a disposable social clip.
- Resolution, length and output format materially change the credit bill.
- Ray3.14 restrictions should not be assumed to describe Ray3.2.
A first workflow to try
- Pick the finishing format before generating variations.
- Use keyframes to define the meaningful moments of the shot.
- Check motion between frames and review the result in the intended color pipeline.
Start here when the problem is shaping or finishing a shot with references, rather than finding a broad visual idea.
Cost & accessRay3.2 720p SDR T2V/I2V: 100 credits/5s or 300/10s. HDR costs 2× and EXR 3×. Video editing and reframe use different tables.
Native Firefly video generation can sit close to an existing Adobe editing workflow, with frame and composition references.
- Native duration
- 5 seconds at 24fps in the documented editor workflow
- Keyframes
- First and last images
- Reference
- Composition guided by existing video
- 4K
- Available through a separate upscaling workflow; not a native-output claim here
- Model choice
- Firefly and partner models have separate features, credits and terms
What makes it useful
- Useful when the team already edits and delivers in Adobe tools.
- Reference composition and frames offer concrete direction.
- Adobe documents native Firefly training and commercial positioning.
Where to pay attention
- Five-second native clips require a shot-and-edit approach.
- A partner model inside Firefly is not the native Firefly model.
- Commercial positioning is not a guarantee that every output clears third-party rights.
A first workflow to try
- Bring an approved frame or composition reference.
- Generate a short insert using the selected native model.
- Edit, add exact titles and audio, then upscale only if delivery requires it.
Consider native Firefly for short inserts in an Adobe finishing workflow and evaluate the exact model terms for a client project.
Cost & accessCheck current Firefly plan and generative credit requirements. Partner generation is billed differently.
The September 2026 app is a creative workspace with multiple underlying video models, character/product tools and sound tools. It is not a single video model.
- Platform
- Video, character and product studios with a multi-model catalog
- Longer sequences
- App advertises up to 30s multi-shot video; model-dependent
- Models
- Catalog includes Seedance 2.5, MiniMax H3 and Pika 2.5 among others
- Commercial license
- Current pricing: Creator/Fancy yes; Free/Starter no
- Credits
- New app credits differ from legacy app, API and iOS credits
What makes it useful
- One workspace for trying different models and creative tasks.
- Character/product studios can shorten a social-content workflow.
- Sound tools sit alongside picture tools.
Where to pay attention
- A Pika result can be generated by another provider’s model.
- An unwatermarked download does not automatically include commercial rights.
- Different models consume different credits; old pricing is not current-app pricing.
A first workflow to try
- Choose the underlying model, not just the Pika app.
- Prepare the character or product reference, then make a short test.
- Check license tier and credit usage before a batch or client delivery.
A useful entry point for everyday social creation when you want several tools in one interface.
Cost & accessMonthly pricing: Starter $10/900 credits; Creator $35/3,150. Seedance 2.5 720p/5s example: 122 credits. Check current terms.
An audio-video model aimed at longer sequences and rich multimodal direction, with timestamp-level editing in the provider’s release.
- Length
- Up to 30 seconds per generation; multiple extension rounds
- References
- Provider model: up to 30 images, 10 videos and 10 audio clips
- Editing
- Timestamp-level audio/video edits; perspective and reference-based editing
- Access
- Release names Jimeng and Doubao; also listed in Pika’s model catalog
- Host caveat
- An app may expose fewer controls, lengths or references than the model announcement
What makes it useful
- Longer individual takes can cover more of a small story.
- Rich reference packs support multiple subjects and scene directions.
- Timestamp-based edits are relevant when one part of a take needs changing.
Where to pay attention
- A model-level reference ceiling is not every host app’s upload limit.
- Longer sequences need careful continuity and audio review.
- Availability, region and cost depend on the access platform.
A first workflow to try
- Build a small consistent reference pack rather than adding every image.
- Write timed beats with explicit cut and sound intentions.
- Review the entire sequence, then edit the problematic moment if supported by the host.
Consider it for a sequence whose references and timeline matter more than one spectacular isolated shot.
Cost & accessHost-dependent. Pika lists 122 credits for a 5s 720p generation; that does not price a 30s run or the native platform.
A multimodal model that can combine visual and sound references with generation and video editing.
- Length / output
- Up to 15 seconds at 2K
- Audio
- Native stereo sound
- Input
- Text, images, video and audio context
- Editing
- Video-to-video motion transfer and natural-language edits
- Access
- Hailuo and supported hosts; exact exposed controls depend on the interface
What makes it useful
- Picture and stereo sound can be generated together.
- Motion-reference workflows offer a starting point for specific camera ideas.
- Video editing can target changes to an existing scene.
Where to pay attention
- Provider claims about brand/text fidelity still need exact-output checks.
- A Hailuo 2.3 spec or price does not describe H3.
- Published comparisons against competitors are provider-reported, not our benchmark.
A first workflow to try
- Define what each reference contributes: subject, movement or sound.
- Start with a short product shot and one motion reference.
- Check readable branding, geometry and sound placement in the actual export.
Consider H3 when a product or motion-reference shot also needs integrated sound.
Cost & accessCheck H3 pricing in the selected interface. We do not infer a rate from older Hailuo models or provider relative-price claims.