Get Started

Model capabilities and supported inputs

Vidu Q4 Preview for Reference-Led Video

Your character already has a face, your scene has a mood, and the next step is a moving shot. Start from images, give the moment one clear action, and review the result in ArcLoop.

Start creating
Actual Vidu Q4 Preview result · 3 seconds · 540p

Vidu Q4 Preview at a glance

Input
One image in image-to-video; 1–15 images in reference-to-video.
Duration
3–16 seconds per generation.
Resolution
540p, 720p, 1080p, 2K, and 4K.
Audio
Generated audio can be enabled or disabled. Audio references are supported in reference-to-video, up to 3.
Aspect ratio
Image-to-video follows the starting image. Reference-to-video offers 16:9, 9:16, 1:1, 3:4, and 4:3.

Choose the input that matches your shot

Vidu Q4 Preview is a video model with image-to-video and reference-to-video modes. Image-to-video animates one starting frame. Reference-to-video uses 1–15 images and can include up to 3 audio references. It requires image input: this integration does not accept a text-only request or a reference video. Choose a mode around the material you already have, then describe what changes in the shot.

Choose the model around your source material

Keep a finished composition

Use image-to-video when the first frame already has the shot you need. The single image establishes the opening view and output ratio; Q4 Preview does not take a separate ending frame.

Combine visual references

Reference-to-video accepts 1–15 images. Use it when identity and scene references matter more than copying one opening composition. References guide the result; they do not guarantee exact continuity.

Plan the delivery format

Choose 3–16 seconds and 540p through 4K in the current controls. Audio generation is optional; reference-to-video also accepts up to 3 audio references. These are supported options, not quality scores.

Vidu Q4 Preview questions

Is Vidu 4 the same name as Vidu Q4 Preview?

This page covers Vidu Q4 Preview, the model available in this ArcLoop workflow. Use the full name when choosing the model so that you do not confuse it with other Vidu releases.

Can I generate with text alone?

No. This integration requires an image. Use one image for image-to-video, or 1–15 images for reference-to-video. A prompt directs the movement; it does not replace the image input.

Can I supply a video or an ending frame?

Reference video input is not supported here. Image-to-video takes one starting frame and does not accept a separate ending frame. Use a supported image workflow for the shot.

How much does a generation cost?

Check the credit estimate in the generator before submitting. Duration, resolution, and available settings can affect the charge. This page does not promise a fixed price or free generation.