Vidu Q4 Preview
How to run image-to-video and reference-to-video on the 7 Oct 2026 Preview: stills, prompts, duration, resolution, native audio, and limits.
Last updated: 2026-10-09
The public Preview that launched on 7 October 2026 is an image-first video model. You start from a still (or a pack of stills), describe motion and camera, and get a clip of 3–16 seconds with optional native audio. viduq4.app is an independent composer for that model — not the official product site. It exists so you can try a shot without writing request JSON.
This page is the practical guide: which job type to pick, how to prepare files, what the knobs actually change, and where the Preview stops. Dollar math lives on pricing. Request fields and official routes live on the API notes. If you are choosing a starting point against a text-first model, see vs Sora 2.
What you can generate
Two job types share the same model id (viduq4-preview):
- Image-to-video — one start frame plus a prompt. The still becomes the first frame; the model extends motion from there.
- Reference-to-video — 1–15 stills for character, product, or style, and optionally up to three short voice clips. Use this when the same subject has to survive cuts inside one take.
On this site the composer labels them Vidu Q4 and Vidu Q4 Refs. They are different jobs, not a single “prompt only” slider.
Typical knobs on both paths:
- Duration: 3–16 seconds. This composer offers 3, 5, 8, 12, and 16.
- Resolution: 540p, 720p, 1080p, 2K, 4K (default 720p).
- Aspect ratios: 16:9, 9:16, 1:1, 3:4, 4:3. Image-to-video output follows the still you start from.
- Audio: on or off per job. On means dialogue and ambience are generated with the picture; off means a silent plate.
Official materials also call out camera switching inside a single take and 10-bit color. Treat those as Preview capabilities, not a promise that every prompt will cut like an edited sequence.
Image-to-video: a working loop
Use this path when you already like a frame and want that frame to move.
- Lock a still. png, jpeg, jpg, or webp. One image only. Keep the subject large enough to read at 720p; a tiny face in a wide plate will smear.
- Write motion, not a second description of the picture. The model can see the still. Tell it what happens next: who moves, how the camera moves, whether anyone speaks.
- Pick duration from the beat, not from habit. A blink-and-turn wants 3–5 seconds. A walk-to-window-and-look-back wants 8–12. Save 16 seconds for a take that actually has two beats.
- Start at 720p. Use 540p only when you are checking whether the motion idea works. Move to 1080p or 2K once the action is stable. 4K is for plates you will grade or crop.
- Decide audio before you iterate. If you need a silent plate for your own score, turn audio off on the first successful motion pass so you are not paying for sound you will mute.
- Submit from the homepage composer. The panel shows the credit estimate before you run. Download the result from generation history when the job succeeds.
A prompt that usually behaves: subject action + camera + optional line of dialogue. Example shape: “She turns toward the window, pauses, then steps forward. Camera dollies in slowly. Soft room tone. She says, ‘I left it on the table.’”
A prompt that usually fails: restating wardrobe, restating the room, and asking for a three-act story in five seconds.
Reference-to-video: keeping a subject consistent
Use this path when the brief is a person, product, or look that has to stay recognizable across cuts.
Still pack (1–15 images). Think in roles, not in a dump of screenshots:
- One or two clear hero portraits (front, three-quarter).
- One full-body or product three-quarter if scale matters.
- One environment or lighting reference if the scene is not in the hero still.
- Optional style stills (color, costume, set dressing) only if those details are not already in the hero frames.
More images are not automatically better. Conflicting wardrobes, mixed art styles, or two different faces for “the same character” will fight each other. Fifteen slots exist for complex packs — a talking-head ad often needs three.
Voice (0–3 MP3 clips, about 3–12 seconds each). Official docs allow these so a character can speak in a consistent voice. Clean, dry reads work better than music beds. If you do not need a cloned voice, skip the audio references and write the line in the prompt instead.
Prompting with multiple stills. Name the role of each image in the prompt (“the woman from the first still walks the aisle holding the bottle from the second still”) rather than hoping the model guesses which face owns which product.
On this site, pick Vidu Q4 Refs, drop the pack in the reference well, set duration / resolution / audio, and submit. The same credit estimate appears before you run.
Native audio
Audio is a per-job flag, default on in this composer. When it is on, the model aims to generate dialogue and ambience with the picture — including lip motion when a line is in the prompt. When it is off, you get a silent plate.
Practical rules:
- Put spoken lines in quotes in the prompt if you want them.
- Do not ask for a licensed song. You will not get a real track; you will get mush.
- If you will replace sound in an NLE, generate silent. Mixing generated room tone under a new score is extra work.
- Voice references on the reference path are for who speaks, not for background music.
Duration, resolution, and aspect
Duration. Official image-to-video allows 3–16 seconds (default 5). Reference-to-video is documented as 1–16 seconds on some vendor pages; this composer still exposes 3, 5, 8, 12, and 16 for both jobs so the UI stays consistent. Shorter takes are cheaper at list and easier to control. If a 16-second prompt is failing, split it into two 8-second shots.
Resolution. 540p is a draft. 720p is the default and enough for most social crops. 1080p is the first “client review” size. 2K and 4K cost more per second at list — see pricing — and take longer to finish. Do not jump to 4K to “fix” a bad motion prompt.
Aspect. Image-to-video keeps the still’s aspect. If you need 9:16, crop or generate the still in 9:16 first. Reference-to-video lets you set 16:9, 9:16, 1:1, 3:4, or 4:3 on the job. Match the destination: Stories and TikTok-shaped placements want 9:16; YouTube and most ads want 16:9; 1:1 is for feeds that still crop squares.
What the Preview does not do
- No text-to-video. A sentence alone is not a valid job. Source or generate a still first, then animate it.
- No minute-long timeline. One job is one take up to 16 seconds. Stitch longer stories in an editor.
- No subject-library call. Official notes say the Preview does not currently support the older “subject invocation” method. You pass images (and optional voice files) on the job itself.
- Not a color-managed finishing tool. 10-bit is useful headroom, not a substitute for a grade.
- Not the official web app. This site calls the model through authorized API channels and bills workspace credits. ShengShu’s own product surface is separate.
If you only have a sentence and no locked still, a text-first model may be a better concepting tool — that split is the whole point of the comparison page.
Stitching a longer piece
Treat each generate as a shot, not as an episode.
- Write a shot list with one action per clip.
- Generate a hero still (or photograph one) per shot so continuity has a source of truth.
- Run image-to-video for locked frames; use reference-to-video when the same face or bottle must recur.
- Keep duration, lens height, and lighting language consistent across prompts (“same apartment, tungsten practicals, 35mm, eye-level”).
- Edit in any NLE. Handles, J-cuts, and titles happen after the model.
A 45-second social ad is typically four to eight Preview jobs plus an edit, not one heroic 16-second prompt.
Quality checklist before you spend a 4K take
- Face or product is large in the still, not a small insert.
- Prompt describes change: motion, camera, light, speech.
- Duration matches the number of beats (one beat per ~3–5 seconds is a useful ceiling).
- Aspect matches the still and the placement.
- Audio flag matches the finish (keep or replace).
- Reference pack does not mix two identities for one character.
- You looked at the credit estimate on the composer and accepted it.
Common failure modes: over-long prompts that re-describe the still; 16-second prompts with five camera moves; reference packs with different people tagged as one hero; 4K used as a first draft.
Try it here
Open the homepage composer. Choose Vidu Q4 for a single start frame or Vidu Q4 Refs for a pack. Drop in stills, set duration and resolution, read the estimate, submit. You need an account to run a job; the draft prompt and references stay in the panel if you are sent to sign-in.
This site is not affiliated with ShengShu. For official field lists see the image-to-video API reference. For how this wallet maps list rates to credits, read pricing.