Guide
A practical Seedance 2.0 prompt structure for text, image, video and audio references, with reusable templates for generation, editing and extension.
Published · Updated
Write one primary action, then specify the subject, environment, camera, lighting, audio and continuity constraints. Assign a clear role to every reference asset, and change one variable at a time during testing.
This guide was checked on August 26, 2026 against ByteDance's official Seedance 2.0 capability page and the BytePlus plan page. It documents workflow structure, not a universal prompt formula. BytePlus currently presents conflicting 4K and 1080p resolution claims, so output limits must be verified in the live product before production.
Use a fixed order so failed generations are easier to diagnose. Keep the requested action physically coherent and avoid stacking unrelated scene changes into one short clip.
Name the main subject, one observable action and the setting. Example structure: [subject] [single action] in [environment], with [important continuity constraint].
Add shot size, camera path, speed, focus behavior, lighting direction and shadow behavior only after the action is clear.
Describe dialogue, ambience and timing separately. State which identity, wardrobe, object, background or motion must remain consistent.
Template: [subject and appearance] [single action] in [environment]. [shot size], camera [movement] at [pace], focus on [visual priority]. [lighting direction and quality], [shadow behavior]. Audio: [dialogue or sound], timed to [event]. Preserve [identity, wardrobe, object and background constraints].
A brushed-steel coffee grinder rotates once on a clean studio table. Medium close-up, slow 30-degree orbit, focus locked on the front dial. Soft key light from camera left with a crisp contact shadow. Audio: one quiet mechanical click at the end. Preserve the logo, dial markings and table position.
Tell the model what each uploaded asset controls. Do not assume that a reference image, video or audio clip will automatically be interpreted as identity, composition, motion or sound guidance.
Use image 1 for the character identity and wardrobe; use image 2 for the room layout and color palette. Animate only [action]. Keep facial features, clothing details and background geometry unchanged. Camera [movement]; lighting [description]; audio [description].
Use video 1 only for movement timing and camera rhythm; do not copy its subject or location. Use audio 1 for beat and event timing. Render [new subject and scene], with [continuity constraints].
Continue from the final frame without a cut. Extend [action] for [requested duration], preserve direction of travel, camera velocity, lighting, ambience and subject identity, then end on [specific final state].
Start with a short draft, inspect the failure mode, and revise one field. Keep reference assets and any exposed random seed fixed while comparing camera, motion or audio variants.
Review motion coherence, identity continuity, reference adherence, camera execution and audio synchronization separately instead of using one overall impression.
Set duration and resolution in the product controls or API rather than embedding them only in prose. Recheck the live account because the checked BytePlus page contains conflicting resolution statements.