The Complete Guide to Seedance 2.5: Prompts, References, Editing, and 30-Second Storytelling
Learn how to direct Seedance 2.5 with structured prompts, multimodal references, timed scenes, precise edits, extensions, keyframes, and production-ready templates.
Seedance 2.5 is designed less like a one-shot text-to-video toy and more like a compact production workflow. You can describe a scene, assign jobs to image, video, and audio references, plan a 30-second sequence, refine an existing clip, or use rough visual blocking to guide a polished result.
This guide focuses on the practical part: how to turn an idea and a pile of assets into instructions the model can follow.
The Core Prompt Formula
A dependable Seedance 2.5 prompt answers seven questions:
Subject + environment + action + camera + look + audio + constraints
- Subject: Who or what must remain recognizable?
- Environment: Where and when does the scene happen?
- Action: What changes from the first moment to the last?
- Camera: What is the framing, lens feel, movement, and edit rhythm?
- Look: What lighting, color, texture, and medium define the image?
- Audio: What dialogue, music, ambience, and effects belong on the timeline?
- Constraints: What must stay fixed, and what must not appear?
Write the main creative intent first. Add references and timing second. Put non-negotiable continuity rules last. A prompt should read like a concise director's brief, not a bag of visual adjectives.
Map Every Reference to a Role
Uploading more material does not automatically produce more control. The model needs to know what each asset contributes.
Use explicit assignments such as:
@Image 1defines the lead character's face, hair, proportions, and wardrobe.@Image 2defines the location, palette, and lighting only.@Video 1defines choreography and timing, but not identity or clothing.@Video 2defines camera movement only.@Audio 1defines the musical beat and overall duration.
If two references disagree, state which one wins. For example: “Use the person from @Image 1; borrow only the body motion from @Video 1; ignore the performer and background in that video.” This prevents a motion reference from quietly replacing your cast or art direction.
Choose references scene by scene
Do not spend the entire reference budget on every shot. Select the smallest useful set for each scene:
- Character dialogue: identity sheet, wardrobe frame, location frame, and voice or dialogue reference.
- Dance or action: character sheet, motion video, camera-motion video if needed, and rhythm track.
- Product commercial: product packshots from useful angles, material close-ups, set reference, and a camera or lighting reference.
- Storyboard sequence: numbered boards or keyframes, character lock, and one consistent style frame.
- Video edit: the source clip plus only the references required for the requested replacement or restyle.
For multi-character work, name each person and reserve a distinct reference group for them. Repeat the mapping near the timeline whenever a character returns after several shots.
Dreamina Reference Limits and Prompt Syntax
The Dreamina guide describes a maximum of 50 references in total, with these category limits:
- Up to 30 images, each input image no larger than 4K.
- Up to 10 videos, with all reference video duration totaling no more than 30 seconds.
- Up to 10 audio files, with all reference audio duration totaling no more than 30 seconds.
The 4K figure above is an input-image limit, not a promise of 4K video output.
Dreamina also documents lightweight syntax for separating sound and text instructions:
- Music:
(soft analog synth pulse, gradually building) - Sound effect:
<ceramic cup placed on a wooden counter> - Dialogue:
{I knew you would come back.} - Subtitle or on-screen caption:
【Three hours earlier】
For non-Chinese dialogue, put the language before the line. You can also specify a regional variety, accent, and delivery, for example: English, Irish accent, quietly amused: {You took your time.}
Treat these markers as clear labels, not magic commands. The surrounding scene, speaker, timing, and emotional direction still matter.
Write 30 Seconds as Timed Beats
Thirty seconds is long enough for a scene to drift if the prompt contains only one action. Divide it into a small number of meaningful phases. Four to six beats are usually easier to direct than thirty one-second instructions.
A useful dramatic shape is:
- 0–5s — Establish: location, subject, and visual question.
- 5–12s — Develop: the main action begins and gains direction.
- 12–20s — Escalate: a reveal, complication, or change in camera energy.
- 20–27s — Payoff: the decisive action or emotional turn.
- 27–30s — Resolve: a readable final image with room to land.
Each beat should specify what changes. “The camera moves” is vague; “a slow dolly-in changes a waist-up two-shot into a tight close-up while she realizes the letter is hers” gives the movement a purpose.
Keep transitions causal. The subject's final position, gaze, momentum, lighting, and screen direction in one phase should give the next phase somewhere logical to begin.
Lock Parameters Before Directing Motion
References establish possibilities; locks establish priorities. Put a compact lock block before the timeline when consistency matters.
Lock only what must not drift:
- Character identity, age range, body proportions, hair, and wardrobe.
- Product shape, label placement, materials, and brand colors.
- Location layout, time of day, weather, and primary light direction.
- Aspect ratio, visual medium, color treatment, and camera axis.
- Required props, character count, and spoken language.
Then separate controlled variation from fixed parameters: “Wardrobe and face remain unchanged; expressions progress from guarded to relieved.” This tells the model that emotional change is intentional rather than identity drift.
Do not issue incompatible locks. A fixed focal length conflicts with a requested zoom-lens effect; a locked noon sun conflicts with a sunset transition. Decide which parameter carries the story.
Direct Observable Emotion
Words such as “sad” or “nervous” name an internal state but do not tell the camera what to see. Convert emotion into behavior:
- Her smile arrives late and disappears when she notices the empty chair.
- He rehearses the greeting under his breath, wipes his palm on his coat, then knocks.
- She avoids eye contact, inhales once, and finally holds his gaze on the last line.
- His shoulders loosen only after the off-screen door closes.
Pair one facial cue with posture, breath, gaze, or hand behavior. For dialogue, add vocal delivery and a reaction beat after the line. Small, timed actions usually read more convincingly than a demand for “extreme emotion.”
Use Professional Camera Language Precisely
Camera terms are useful when they describe a visible production choice:
- Shot size: extreme wide, wide, medium, medium close-up, close-up, extreme close-up.
- Angle: eye level, low angle, high angle, overhead, over-the-shoulder, Dutch angle.
- Movement: pan, tilt, dolly in or out, truck left or right, crane, pedestal, orbit, handheld follow.
- Lens behavior: wide-angle spatial depth, natural perspective, compressed telephoto background, shallow depth of field, rack focus.
- Staging: two-shot, profile composition, foreground occlusion, deep staging, clean single, shot-reverse-shot.
- Editing: hard cut, match cut, whip-pan transition, motivated cut, continuous take, J-cut, L-cut.
Do not stack movements simply because they sound cinematic. “A slow lateral track as the cyclist accelerates” is clearer than “pan, orbit, crane, zoom, and handheld.” Define the subject-camera relationship and one dominant move per beat.
First and Last Frames, Keyframes, and Storyboards
Different planning assets solve different problems.
First and last frames
Use a first frame to lock the opening composition. Add a last frame when the destination matters just as much: a product must end centered, a transformation must reach a specific design, or a character must arrive at an exact pose.
Describe the path between them. Without an action bridge, the model may morph from one state to the other rather than stage a physical transition.
Multiple keyframes
Use keyframes when a shot has several required visual milestones. Number them in order, assign approximate times, and state whether the camera cuts or moves continuously between them. Keep neighboring frames compatible in character, geography, and lighting.
Storyboards
A storyboard controls shot order and composition. Tell the model whether each panel represents a cut, a camera destination, or an action peak. Do not assume visible panel numbers or annotations should appear in the final video; explicitly exclude board borders, labels, arrows, and production notes.
Rough and refined blockouts
A rough blockout—simple shapes, an animatic, a 3D white model, or green-screen staging—is useful for timing, spatial layout, and camera paths. A refined blockout adds clearer poses, lens choices, lighting direction, and edit points. In both cases, state what should be preserved and what should be replaced by the final style.
The official ByteDance Seed page specifically identifies white-model control and green-screen editing as Seedance 2.5 production capabilities.
One-Click Video Without Losing Control
For a fast one-click workflow, prepare one compact production packet:
- A clear prompt using the core formula.
- A small, role-labeled reference set.
- A lock block for identity, wardrobe, setting, and format.
- A timed sequence for anything longer than a single simple action.
- An end-state and a short exclusion list.
Then generate the full shot in one pass. “One click” works best when the planning has already removed ambiguity. If the first result misses one isolated element, edit that element rather than rewriting the entire concept.
Edit Existing Video with Surgical Instructions
Seedance 2.5 can be directed to modify existing video. Write edit prompts in four parts:
Keep + change + affected time or region + continuity requirement
Example:
Keep the source camera movement, actor performance, timing, shadows, and ambient audio. From 06–11 seconds, replace only the paper cup in her right hand with the blue ceramic mug from
@Image 1. Preserve hand contact, reflections, scale, occlusion, and motion blur. Do not alter her face, wardrobe, background, or the rest of the timeline.
For green-screen material, state how the new environment should match the subject: contact shadows, color spill, depth of field, perspective, and interactive light. For white-model footage, identify whether the blockout controls staging, motion, camera, or all three.
Extend Forward or Backward with Continuity
ByteDance's official page says a generated video can be extended twice. An extension should begin from a continuity handoff, not a new synopsis.
For a forward extension, restate the final observable state:
- Subject position, facing direction, pose, and motion vector.
- Camera height, framing, lens feel, and movement.
- Lighting direction, weather, and background activity.
- Wardrobe, props, damage, and other accumulated changes.
- Audio ambience, musical phase, and unfinished dialogue.
For a backward extension, define the required opening state of the existing clip and reverse-engineer a plausible lead-in. The new segment must arrive at the original first frame with matching momentum, exposure, subject placement, and sound.
Avoid repeating the same action at the seam. Ask for continuation from the exact contact, step, glance, or camera movement already underway. If possible, use the boundary frame and the source clip together.
Design Seamless Transitions
A seamless transition needs a shared visual or physical bridge. Reliable options include:
- Match on action: a hand movement continues across the cut.
- Shape or color match: a circular lamp becomes the moon; a red fabric fills the frame and reveals a new location.
- Occlusion: a foreground object wipes across the lens.
- Whip pan: directional blur carries the same speed and direction into the next scene.
- Lighting bridge: a flash, shadow, or exposure change motivates the transformation.
- Sound bridge: the next scene's audio begins before the visual cut, or the previous sound continues after it.
Specify the seam's direction, speed, framing, and audio. “Transition seamlessly” alone leaves too many choices open.
Copy-Ready Prompt Templates
Replace the bracketed fields and reference numbers with your own material.
30-second cinematic scene
Create a 30-second [genre] scene in [aspect ratio].
REFERENCE MAP
@Image 1: lock [character] identity, hair, proportions, and wardrobe.
@Image 2: use only for [location], palette, and light direction.
@Video 1: use only for [movement/choreography], not identity or setting.
@Audio 1: use as the timing and musical reference.
LOCKS
Keep [identity/product/location parameters] unchanged throughout. Maintain [camera axis, lighting, style, character count].
TIMELINE
0–5s: [establishing composition and first action].
5–12s: [development, camera relationship, observable emotion].
12–20s: [escalation or reveal].
20–27s: [payoff action and strongest camera beat].
27–30s: [resolved final image and audio landing].
AUDIO
(music direction)
<key synchronized sound effect>
[Language, accent, delivery]: {dialogue}
Exclude [only the most important unwanted changes, artifacts, text, logos, or extra subjects].
Product video from multiple references
Create a clean 15-second premium product film for the object in @Image 1.
Use @Image 2 for the side profile and @Image 3 for material detail only. Preserve the exact silhouette, proportions, finish, controls, and logo placement from every angle. Use @Image 4 only for the warm studio palette.
0–4s: Macro detail of [feature], slow lateral track, narrow highlight moving across the surface.
4–10s: Pull back into a three-quarter hero view as [physical product action] occurs with realistic weight and contact.
10–15s: Controlled half-orbit ending on the exact front view from @Image 1, centered with negative space for copy to be added in post.
<precise product sound>, (restrained minimal score). No generated captions, extra controls, warped branding, floating parts, or shape changes.
Emotion-led dialogue scene
Two characters wait on an empty late-night train platform. Lock the identities and wardrobes from @Image 1 and @Image 2. Cool overhead light, soft rain beyond the platform roof, naturalistic drama, restrained handheld camera.
Begin in a quiet waist-up two-shot. Character A grips a folded ticket, tries to speak, then looks away. Character B notices the gesture but does not interrupt. Slowly push into Character A as their shoulders settle.
English, London accent, barely above a whisper: {I kept the ticket.}
<distant train brakes, rain on metal roofing>
(very soft low strings entering only after the line)
Hold the reaction for two seconds: Character B exhales, looks at the ticket, and gives a small disbelieving smile. Keep eyelines, rain direction, background geography, and facial identity consistent. No subtitles.
Precise source-video edit
Edit @Video 1 without changing its duration, shot order, camera motion, performance, or audio timing.
Change only [object/person/region] during [time range]. Replace it with [description or @Image reference]. Match the source perspective, scale, lighting, reflections, motion blur, occlusion, and physical contact in every affected frame.
Keep [protected elements] exactly as in the source. Do not modify any frame outside the stated range. No new text, logos, props, or people.
Forward extension
Extend @Video 1 forward from its exact final frame. Begin with [subject] at [position and pose], moving [direction and speed], while the camera continues its [existing movement]. Match the existing lens feel, exposure, light direction, weather, wardrobe, props, environment, ambience, and music phase.
Continue the unfinished action: [next causal action]. Do not replay the ending or pause at the seam. Over [duration], develop toward [new payoff], then finish on [specific final composition]. Preserve identity and spatial geography throughout.
Pre-Flight Checklist
Before generating, confirm:
- The prompt has one clear creative goal.
- Every reference has one named role, and conflicts have a priority rule.
- Character, product, setting, and format locks are explicit but not contradictory.
- A long clip has timed beats with visible changes.
- Actions connect physically across beats, cuts, and extensions.
- Camera terms describe one dominant choice per moment.
- Emotion is expressed through observable behavior and vocal delivery.
- Dialogue, music, sound effects, and captions use clear labels and timing.
- The final state is defined.
- Exclusions are short and directly relevant.
- Reference counts and combined durations stay within the Dreamina guide's stated limits.
- You have not treated the per-image 4K input ceiling as an output-resolution claim.
- You have permission to use the people, brands, footage, voices, and music in your references.
Limitations and Evidence Boundary
Seedance 2.5 is generative, so prompt compliance, identity, readable text, anatomy, object permanence, dialogue, and audio synchronization can still vary between results. Review every output before publishing, especially for brand assets, recognizable people, product claims, safety-sensitive actions, and licensed material.
The official ByteDance Seed page confirms single generations up to 30 seconds, up to two extensions, more precise reference understanding and editing, and production workflows involving white models and green screens. It does not publicly specify 4K output, API availability, pricing, regional availability, generation speed, or commercial licensing terms on that page. Do not infer those details from the 4K input-image limit in the Dreamina guide. Product interfaces and limits may also change after publication, so verify current settings in the service you use.
Put the Workflow Into Practice
Explore more ready-made structures, or start creating with Seedance 2.5 in SJinn.
