Back to home

Vidu Q4 vs Sora 2

Image-to-video and multi-reference consistency versus OpenAI text-to-video: how you start a shot, limits, access, and which workflow viduq4.app actually runs.

Last updated: 2026-10-09

This is a workflow choice, not a single “better model” score. The Preview that launched on 7 October 2026 starts from images and references. Sora 2 is OpenAI’s video model, known for text-to-video with synced sound (and, in some product surfaces, an optional image anchor). viduq4.app runs the Preview — it does not host Sora.

Use this page to decide how you want to start a shot. Specs for the image-first model live on the Preview guide. List rates and credits live on pricing. Official request shape lives on the API notes. Treat OpenAI’s own pages as source of truth for Sora duration, resolution, and availability — those numbers move independently.

A locked product still: a corked glass bottle on a wooden table against a plain backdrop

How you start a clip

This Preview does not take a text-only prompt. You need at least one still:

  • Image-to-video: one start frame + motion / camera text. The still is the first frame.
  • Reference-to-video: 1–15 stills and up to three voice clips so a character or product stays recognizable across cuts inside one take.

Sora 2 is the text-first cousin: describe a scene and get moving picture and sound without uploading a frame. Some OpenAI surfaces also accept a reference image as a visual anchor for the first frame, but the product is still built around a sentence. That is useful for concepting from nothing. It is the wrong tool if an art director already signed off on a key still and you must not drift off it.

If you are coming from sentence-only prompting, generate or photograph a still first, then animate it here. The still is the contract; the prompt is the motion.

Consistency and references

Pick the Preview’s reference-to-video path when identity is the job: the same face, bottle, jacket, or room across cuts. You pass the pack on every job (1–15 stills). Official notes say the Preview does not currently support a stored “subject invocation” library — there is no character id to reuse; you re-attach the pack.

Sora-family products have shipped their own character-reference tools on some API snapshots (upload a character, reuse an id). Those tools live in OpenAI’s stack. They are not a toggle on this composer. Do not assume a Sora character id will mean anything to viduq4-preview.

Decision shortcut:

  • Locked hero still, must match → image-to-video on this site.
  • Locked identity across several angles → reference-to-video on this site.
  • No still yet, exploring a scene → a text-first model in its own product.
  • Need the same person next week without re-uploading a pack → check whether your text-first vendor still offers a character library; this Preview does not.

Length, resolution, and audio

On the Preview you pick 3–16 seconds, up to 4K with 10-bit color, and native audio on or off. Aspect follows the still for image-to-video (16:9, 9:16, 1:1, 3:4, 4:3). Multi-shot camera moves inside one take are part of the Preview pitch — still a single generate, not an NLE.

Sora 2’s public product has been framed around physics-looking motion and dialogue from a prompt, usually inside OpenAI’s own apps. Published developer docs have listed sizes such as 1280 by 720 and 720 by 1280 (and higher tiers on a Pro id) and clip lengths in a small set of seconds. OpenAI has also documented extensions that append more time onto an existing clip, which this Preview does not offer as a first-class “continue this MP4” control. Developer availability for Sora has changed over 2026 — including withdrawal of some API snapshots — so verify on OpenAI’s site before you plan a pipeline around it.

Neither model is an editor. A 60-second story is still a sequence of short generates plus a timeline.

Access and what this site is

  • Preview (viduq4-preview) — public web + API as of 7 October 2026. This independent site wraps it in a credit-priced composer. See the API page for official api.vidu.com routes. You will not paste a vendor key into the homepage panel.
  • Sora 2 — OpenAI product surface (app / ChatGPT / whatever OpenAI is shipping the week you read this). There is no Sora toggle in this generator.

viduq4.app is not affiliated with ShengShu or OpenAI. “Vidu” marks belong to ShengShu Technology; Sora belongs to OpenAI. This page compares workflows, not corporate products you can buy as a bundle.

Decision criteria by job

Product film / e-commerce. You already have pack shots. Image-to-video (or reference-to-video if the bottle must turn and stay on-model). Crop the still to the placement aspect first.

Performance / talking head. Reference pack of the person + optional voice clip + a line in the prompt. Preview path. A text-first model will invent a face unless you give it a reference, and even then it may not hold identity as tightly as a dedicated reference job.

Concept / mood / “what if the city flooded.” Text-first. You do not have a locked frame yet, and you should not pretend you do.

Social series with a recurring character. Preview reference-to-video for the recurring identity; edit privately. Budget for many 5–8 second shots, not one 16-second miracle.

Pitch deck from a paragraph. Text-first to get any moving picture, then (optionally) lock a still from a frame grab and re-animate here if the client picks a look.

High-resolution finishing. Preview list goes to 4K. Confirm Sora’s current max size on OpenAI’s pages if you need a side-by-side export spec. Do not assume 4K exists on every Sora SKU.

Prompting differences

Image-first prompts should not re-describe the still. The pixels already carry wardrobe, set, and face. Write verbs: turns, steps, rack-focuses, dollies, says. Put dialogue in quotes if you want native audio to try a line.

Text-first prompts have to carry the whole frame: who, where, time of day, lens, wardrobe, motion. That is more writing, and more chances to drift when you iterate.

Reference-to-video prompts should bind roles to stills (“the woman from still 1 hands the bottle from still 2 to the man from still 3”) so the pack is not a pile of unexplained pictures.

If a Preview job ignores your camera move, shorten the clip and remove extra clauses. If a text-first job ignores a locked brand color, stop fighting the sentence and start from a still.

Cost and operations (high level)

This site bills workspace credits from the vendor estimate, shown before submit. 10 credits = $1. List rates for this model run from $0.045/sec at 540p to $0.39/sec at 4K, with a 30% image-to-video / reference-to-video promo on some providers through 30 November 2026. Full tables and plan math: pricing.

Sora’s list price, if any, lives on OpenAI’s billing pages and has never been this site’s checkout. Do not convert one vendor’s invoice into the other with a homemade multiplier. Compare workflows and access, then compare the number each UI shows you before you run.

Ops differences that matter more than a score:

  • This composer requires a still; you cannot “just type.”
  • This composer submits one job from the panel at a time.
  • Official Preview result URLs on api.vidu.com are documented as 24-hour links — copy them if you self-host.
  • Sora access may be consumer-app-only, API-only, or unavailable in your region this week. Check OpenAI.

When to pick which

Choose the Preview on this site when you already have a frame, a character sheet, or a voice you must keep, and you want stills-to-motion up to 2K/4K with optional native sound. That is the path the homepage is built for.

Choose Sora 2 in OpenAI’s products when the brief is a sentence and you do not yet have a locked still — and when that product is actually available to you.

Choose both, in sequence when you concept in a text-first tool, grab a frame you like, and finish motion from that still here so the approved picture cannot drift.

Try a Preview clip on the homepage — pick Vidu Q4 or Vidu Q4 Refs, drop in a reference, read the estimate, generate.