Skip to main content

Video to Seedance Prompt

Video to prompt reverse-engineers an existing clip into a shot-by-shot Seedance prompt: subject, scene, camera, lighting and per-shot timing, plus the first and last frame extracted as images you can feed straight into image-to-video.

How it works

  1. 1Upload a clip. The browser pulls the first, middle and last frame locally; only those frames and the video itself are sent.
  2. 2A vision model reads the video with per-frame timestamps and returns a structured breakdown — subjects, scene, and one entry per shot.
  3. 3That breakdown is rendered into two prompts, one for Seedance 2.5 and one for Seedance 2.0.
  4. 4Open it in the creator with the first and last frame already attached, and generate.

Why two prompt versions

Seedance 2.0 does not respond to timestamps — it only responds to shot numbers. Handing it a 2.5 prompt silently discards the segmentation, so each version is rendered separately rather than trimmed from the other.

Seedance 2.5Seedance 2.0
Shot segmentationInteger-second timestamps (0s-3s:)Shot numbers (Shot 1:)
Aspect ratioAny ratio in [0.4, 2.5]Six fixed ratios
Max length30 seconds15 seconds

Limits and cost

Formats
MP4, MOV, WebM
Maximum size
50 MB
Maximum length
60 seconds
Cost
3 credits, plus 1 for clips over 30 seconds
Reference frames
First, middle and last, uploaded with the result
Eligibility
Accounts that have topped up

What it does not recover

  • Dialogue and sound effects. This pass reads frames only, so any line of dialogue would be invented rather than transcribed.
  • An exact copy. A prompt is a lossy description of a video; the result matches subject, staging, camera and pacing, not every pixel.
  • Negative constraints. Seedance only supports negative control for subtitles and audio, so the output uses positive description throughout.

Questions

What does video to prompt actually return?
Two prompts and a set of frames. The Seedance 2.5 version segments the clip with integer-second timestamps; the Seedance 2.0 version uses shot numbers instead, because 2.0 does not respond to timestamps. Alongside them you get the first, middle and last frame as images.
Can it recover the dialogue in a video?
No. The analysis reads sampled frames, not the audio track, so dialogue and sound-effect fields are deliberately left empty. Reading speech off lip movement in still frames is not recoverable, and filling those fields would mean inventing them.
How close will the generated video be to the original?
Close in subject, staging, camera movement, lighting and pacing — not pixel-identical. Carrying the extracted first and last frame into image-to-video pins the opening and closing images, which is what makes the result actually resemble the source rather than merely describe it.
What does it cost?
3 credits per extraction, plus 1 credit for clips longer than 30 seconds. Credits are pay-as-you-go with no subscription, and a failed extraction is not charged.
Which video files are supported?
MP4, MOV and WebM, up to 50 MB and 60 seconds.
Who can use it?
Accounts that have topped up. Each extraction calls a paid vision model, so it is not offered on unfunded accounts.
Video to Prompt · Seedance Prompt From a Clip | AiuniVid