AI Video Prompt Structure: Subject, Motion, Camera, Duration
Ai video prompt structure: subject, motion, camera, duration for Kling, Hailuo, Luma, and Runway Gen-4. SMCD framework, examples, and workflow.
Turn any photo into an AI prompt — free
No sign-up required. Works with Midjourney, FLUX, DALL-E.
Try Image to Prompt →An ai video prompt needs four decisions named in order: subject, motion, camera, duration. Subject is who or what owns the frame. Motion is what changes over time. Camera is how the lens behaves. Duration is how many seconds the clip runs. Skip one pillar and Kling AI, Hailuo, Luma Dream Machine, or Runway Gen-4 fills the gap with random drift. This guide teaches the SMCD stack so you write paste-ready lines by hand, judge generator output, and turn a reference still into a motion brief. You leave with a reusable template, three worked examples, model dialect notes for mid-2026 hosts, and a soft path to PromptMake /image when your clip starts from a photo.
What ai video prompt structure means
Structure means the four pillars appear in every brief whether you label them or bury them in prose. Image prompts freeze one frame. Video prompts describe change across seconds. Most failed clips trace back to image language pasted into a video box: beautiful lighting, cinematic, 8K. Those words describe a poster. They never say what moves.
You need SMCD if you cut product teasers, social B-roll, concept pre-viz, or image-to-video tests on Kling or Luma. You also need it if a text tool drafts your first pass. The tool fills labels. You still approve subject nouns, one camera path, and a duration that matches the motion you wrote.
This page is a craft framework for ai video prompt structure. It is not a generator product tour and it is not the seven-field anatomy primer elsewhere on the blog. Those pages cover tool loops and extended fields like audio and continuity anchors. SMCD is the spine every host reads first. Add light, sound, and continuity after the four pillars hold.
Who should skip deep SMCD work: teams that only export still frames from Midjourney or FLUX and never touch motion UIs. Who should lean in: marketers testing five-second hero clips, creators running image-to-video on a locked portrait, and designers storyboarding motion before they spend credits on a twelve-second wish list.
The four pillars at a glance
Subject, motion, camera, duration. Write them in that order when you draft by hand. Read them in any order when you edit someone else's paste. Each pillar answers one question the renderer asks silently: what is on screen, what changes, how the frame travels, and how long the change runs.
Kling AI and Hailuo reward clear subject nouns plus explicit camera paths separated from subject motion. Luma Dream Machine reads natural sentences when duration stays short. Runway Gen-4 expects cinematic vocabulary but still collapses when you stack three camera moves in one block. The pillar names stay fixed. Only the paste dialect shifts.
Treat duration as a hard cap, not a vibe word. A four-second prompt with one motion beat outperforms a ten-second prompt with five beats on most consumer tiers as of August 2026. Match motion scope to seconds before you polish adjectives.
Subject: anchor the frame
Subject is the noun stack that must stay recognizable. Name material, color, scale, and count. "Matte white ceramic mug on oak sill, no print, single object" beats "coffee mug." For people, fix hair color, jacket tone, and age band once. Image-to-video routes inherit the upload as subject. Your text prompt then adds motion and camera on top of pixels the model already locked.
Keep one primary subject per clip unless the job is a deliberate two-person dialogue shot. Crowd scenes need explicit count language: "three runners mid-ground, faces soft, single hero runner sharp in foreground red bib." Vague crowd words produce extra limbs and costume drift between frames.
Subject also carries scene context when you write text-to-video from scratch: "sunlit kitchen counter, empty aside from the mug" grounds the prop. Without place, you get gray void or random clutter. Pair subject noun with one location clause before you touch motion.
Motion: timed change
Motion is verbs with timing. Write what happens in order across the clip length you plan to generate. "0–2s mug static; 2–4s steam curls once from the rim" gives Kling a readable spine. "Steam rises beautifully" gives the model room to invent a smoke column that fills the frame.
Separate subject motion from camera motion in the same sentence when both exist: "Mug stays fixed; camera pushes in slowly." Models merge the two when you write "dynamic cinematic reveal" and the prop spins while the lens orbits.
One primary motion beat per two to four seconds of duration is a safe default on Hailuo and Luma short clips. Save multi-beat montage language for longer Runway Gen-4 runs where the UI exposes ten-second lanes and you storyboard shot IDs.
Camera: one path per clip
Camera combines framing and movement. Framing: wide, medium, close-up, extreme close-up. Movement: locked tripod, slow dolly in, lateral truck left, gentle tilt up, static handheld micro-shake. Pick one move per generation pass. "Slow dolly in only" travels across hosts. "Epic orbit crane swoop" does not.
Write lens feel when the host supports it: "35mm spherical, shallow depth on subject, background soft." That line belongs near camera, not subject, because it describes how the frame sees the anchor object. Runway Gen-4 and Kling both respond to focal length cues in mid-2026 public docs.
When you run image-to-video, camera motion often matters more than subject motion because the still already fixed pose and light. "Locked subject; slow push-in from medium to close on the face" is a complete motion-plus-camera line for a portrait upload on Luma.
Duration: seconds as a design choice
Duration belongs in the UI slider on most hosts and in your prompt plan on every host. Decide seconds before you write beats. A four-second Kling clip can hold one push-in and one steam curl. A ten-second Runway clip can hold an establish cut plus a hero move if you split beats across timecodes in the text.
Avoid asking prose to carry resolution or frame rate unless the vendor docs for your route say otherwise. "4K masterpiece ten seconds" inside the prompt fights the container settings and adds no motion detail. Set length in the control panel. Spend prompt tokens on subject, motion, and camera.
If your first render fails, shorten duration before you rewrite subject. Half the time the model ran out of temporal budget for the beats you listed.
How to assemble SMCD without fluff
Field order beats elegant paragraphs on the first draft. SMCD assembly is a worksheet habit. You fill four lines, read them aloud, then merge into paste-ready prose if the host prefers sentences over labels.
Cut adjective stacks that do not change pixels: stunning, breathtaking, hyper-detailed. Keep nouns, verbs, and direction words. "Soft morning window light from camera left" sets light without eating motion budget. Save style spine lines for a second pass after SMCD holds.
Practice on an object within arm's reach so you can verify color and material against reality. A real mug, plant, or keyboard grounds subject nouns. You learn faster than copying forum prompts meant for someone else's product.
Draft subject and motion first
Line 1 subject: concrete noun, material, location, count. Line 2 motion: timed beats matched to your target seconds. Example for a five-second Luma text-to-video test: "Subject: terracotta pot with single snake plant, concrete windowsill, afternoon sun. Motion: 0–3s leaves still; 3–5s one leaf tip quivers slightly in a breeze."
Stop after two lines and picture the clip. If you cannot see the motion in the seconds you chose, delete a beat or add two seconds in the UI. Do not add camera until subject and motion read clean.
For image-to-video, subject line shrinks to change language: "Preserve uploaded portrait; face and wardrobe unchanged. Motion: subtle blink and slight head turn toward camera over 3s." The upload carries subject pixels. Text carries delta.
Add camera and duration last
Line 3 camera: framing plus one move. Line 4 duration: confirm the UI setting matches your beats. Example continuation: "Camera: medium shot, slow dolly in only, 50mm feel. Duration: 5s in host slider; prompt beats written for 0–5s."
Merge into dense prose for paste boxes that reject labels: "Medium shot of a terracotta pot with a single snake plant on a concrete windowsill in afternoon sun. Camera slow dolly in only. Plant still for three seconds, then one leaf tip quivers in a breeze. 50mm shallow depth, warm daylight from the right."
Read merged prose for duplicate motion verbs. If dolly and orbit both appear, delete one before you spend credits.
Step-by-step: reference still to video prompt
Many Kling, Hailuo, and Luma workflows start from a hero still you already approved in Midjourney, FLUX, or GPT Image. SMCD turns that still into a motion brief without re-describing every pixel from scratch.
Budget fifteen minutes the first time. Later runs take five when your SMCD shell exists. Keep a notes row per clip: subject anchors, motion beats, camera move, seconds, host name, date.
Image-to-video quality rises when the still already matches your aspect target. Crop before upload. A vertical portrait uploaded into a horizontal video lane forces the model to invent side space.
Step 1: Lock the hero frame
Generate or select one still that passes your static checklist: subject readable at thumbnail size, clean edges, no unwanted text, lighting direction obvious. That still is the subject anchor for image-to-video.
If the still is weak, fix it in image tools before motion. PromptMake /image can draft Midjourney, FLUX, DALL·E, Stable Diffusion, or Leonardo text from a reference photo when you need a cleaner hero. Pick Recreate Exactly to hold composition, Change Style when only medium drifted, Adjust Lighting when motion can wait but exposure cannot. Soft start: https://promptmake.net/image. Guest quota is about three image runs per day; free registration raises image quota separately from /text.
Export at the largest size your still tool allows. Downscaled uploads lose fine subject detail that motion models use for consistency.
Step 2: Layer motion and camera on the still
Upload the still to your video host. Write SMCD text that respects pixels already present. Lead with preserve language: "Maintain subject, lighting, and background from reference." Follow with motion and camera lines only.
Test the smallest motion first on Kling or Luma: slow push-in, locked subject, three to five seconds. A micro-win teaches what the host preserves from upload before you ask for hair wind and cloth ripple in the same pass.
Log the exact text that worked next to the still filename. Next week you swap subject by changing the image, not by rewriting motion grammar.
Step 3: Set duration and iterate one variable
Set seconds in the UI to match your beat list. Generate once. Score: Did subject drift? Did camera obey one path? Did motion finish inside the window? Change one pillar per retry. Drifted face: reduce motion adjectives. Mushy camera: delete extra move words. Early cut-off: add one second or remove a beat.
Runway Gen-4 and Hailuo sometimes expose motion strength sliders. Lower strength when upload identity matters. Raise strength when text-to-video needs bolder action from a blank frame.
After two clean clips on the same shell, save the SMCD template with bracket variables: [subject noun], [motion beat], [camera move], [seconds].
Worked SMCD examples
Three teaching copies below share one discipline: one subject, one motion spine, one camera path, explicit seconds. Adapt nouns to your product. Keep the pillar order when you edit.
Example A, product hero on Kling AI text-to-video: Subject: brushed steel water bottle, blue cap, no logo, centered on gray stone surface, studio void background. Motion: 0–2s bottle static; 2–5s condensation bead slides down the left side once. Camera: medium shot, slow dolly in only. Duration: 5s host setting. Paste prose: "Medium studio shot of a brushed steel water bottle with blue cap, no logo, on gray stone. Bottle static two seconds, then one condensation bead slides down the left side. Slow dolly in only, soft key light from upper left, 85mm product feel."
Example B, portrait image-to-video on Luma Dream Machine: Subject: preserve uploaded portrait, woman in olive jacket, overcast park background. Motion: 0–4s subtle smile grows, single blink at 2s. Camera: locked framing, no zoom. Duration: 4s. Paste prose: "Maintain reference portrait exactly. Subtle smile grows over four seconds, one natural blink at two seconds. Camera locked, no zoom, shallow depth preserved from source."
Example C, street scene on Runway Gen-4 text-to-video: Subject: wet asphalt crosswalk at dusk, single yellow taxi mid-frame, neon shop signs soft background. Motion: 0–3s taxi approaches slowly; 3–6s brake lights brighten once. Camera: low angle static tripod, 24mm wide. Duration: 6s. Paste prose: "Low wide shot of wet asphalt crosswalk at dusk, yellow taxi mid-frame approaching slowly for three seconds, brake lights brighten at three seconds, neon signs soft in background, static tripod, 24mm cinematic feel."
Weak contrast for study: "Cinematic taxi scene, beautiful rain, epic vibe, 8K, ten seconds of action." No subject anchor, no timed beat, no camera path, no honest duration match. SMCD fixes that line in four edits, not forty adjectives.
Common ai video prompt mistakes
Mistake 1: Image prompt pasted into video. Still captions freeze light and pose. Add motion and camera or accept invented chaos.
Mistake 2: Duration longer than beats. Twelve seconds of prose with one steam curl wastes Runway credits on empty tail frames.
Mistake 3: Two camera moves in one pass. Orbit plus dolly plus tilt collapses on Kling and Hailuo. One path per generation.
Mistake 4: Subject motion and camera motion fused into "dynamic reveal." Split them explicitly every time.
Mistake 5: Upload aspect fights output aspect. Crop the hero still to the video lane before image-to-video.
Mistake 6: Skipping duration planning. Beats written for eight seconds while the UI sits at four produces clipped or frozen motion.
Mistake 7: Rewriting all four pillars when one failed. Isolate camera if framing broke. Isolate motion if subject was fine but action wrong.
Mistake 8: Chasing model-specific lore before SMCD holds. Master the four lines, then read host release notes for dialect tweaks.
Model notes for Kling, Hailuo, Luma, Runway (2026)
Public names shift. Treat the notes below as an August 2026 snapshot. Confirm model pickers inside Kling AI, MiniMax Hailuo, Luma Dream Machine, and Runway before you paste old forum prompts.
Shared rule across all four: SMCD first, then lighting spine, then optional audio if the host generates sound. Kling and Runway expose longer clips on paid tiers. Luma and Hailuo often shine on three-to-five-second image-to-video tests.
Image-to-video routes on these hosts assume your still is the subject pillar. Text adds motion, camera, and duration intent. Text-to-video routes need full subject prose in line one.
Kling AI and Hailuo
Kling AI (Kuaishou) remains a strong pick for image-to-video on product and character stills as of mid-2026. Write preserve language on uploads. Keep camera to one move. Kling reads separated clauses well: subject block, motion block, camera block.
Hailuo from MiniMax competes on short cinematic clips with clear motion verbs. Timed beats help: "0–2s static, 2–4s hair lifts in wind once." Avoid montage lists inside one Hailuo pass. Chain clips in edit if you need multiple beats.
Both hosts punish conflicting motion adjectives. "Explosion of movement" plus "calm minimal scene" in one prompt yields neither. Pick energy level once in the motion pillar.
Free tiers on Kling and Hailuo rotate daily credits. Spend them on SMCD-correct five-second tests, not on twelve-second first drafts.
Luma Dream Machine and Runway Gen-4
Luma Dream Machine favors readable sentences when duration stays at five seconds or less. Image-to-video on Luma works best with subtle motion on uploads: blink, breath, slow push-in, fabric shift.
Runway Gen-4 targets longer narrative clips and camera vocabulary trained on film language. SMCD still applies. Runway tolerates slightly richer camera words but not three moves at once. Use Gen-4 when you need ten-second text-to-video with explicit lens cues.
Runway image-to-video often pairs with Gen-4 still tools in the same suite. Lock subject in a still export, then promote the same frame to motion with SMCD text layered on top.
Luma text-to-video accepts scene prose if subject leads the sentence. Runway prefers cinematic ordering: shot type, subject, action, light, camera.
When a host adds native audio in 2026 builds, treat audio as a fifth optional field after SMCD. Name ambient sound or write "no music" so the model does not invent a score over your quiet product clip.
When PromptMake /image fits video prep
Video hosts consume pixels. PromptMake /image produces still-language and reference-aligned drafts before you open Kling or Luma. Use /image when you need a hero still that matches SMCD subject anchors, when your upload keeps drifting because the source photo was noisy, or when you want Midjourney or FLUX dialect from a reference without hand-translating flags.
Workflow: build subject still in your preferred image model, or upload a product photo to https://promptmake.net/image and copy formatted prompt text for iteration. Fix composition and light until the frame is thumbnail-clear. Export PNG. Upload to image-to-video. Paste SMCD motion and camera lines only.
Goal modes map cleanly to video prep. Recreate Exactly holds layout for image-to-video uploads. Create Variation builds alternate hero stills with the same SMCD shell. Adjust Lighting fixes exposure before motion when windows blow highlights.
Guest accounts receive about three /image runs per day. Free registration raises the image quota. Quota is separate from /text. Spend image runs on subject lock, not on hunting synonyms for motion you have not written yet.
PromptMake does not render video clips. It shortens the still step so SMCD motion lines attach to a frame worth animating.
FAQ
What is ai video prompt structure?
Ai video prompt structure is the SMCD stack: subject, motion, camera, duration. Subject names who or what owns the frame. Motion names timed change. Camera names framing and one move. Duration matches beats to seconds in the host UI. The four pillars appear in every strong Kling, Hailuo, Luma, or Runway brief whether you label them or write dense prose.
How is SMCD different from a seven-field video prompt?
SMCD is the minimum spine: subject, motion, camera, duration. Extended guides add scene context, lighting spine, continuity anchors, and audio as separate fields. Learn SMCD first when clips fail on basic motion. Add extended fields when you cut multi-shot ads with sound sync.
Which pillar should I write first?
Write subject and motion before camera and duration. Those two decide whether a five-second clip reads at all. Add camera once you can picture the action. Set duration in the UI to match your beat list. If a render fails, fix subject or motion before you rewrite camera vocabulary.
Does ai video prompt structure work for image-to-video?
Yes. Upload carries subject pixels on Kling, Hailuo, Luma, and Runway image-to-video routes. Text adds motion, camera, and duration intent. Lead with preserve language, then timed motion beats, then one camera path. Crop the still to output aspect before upload.
How long should my ai video prompt be?
Length matters less than pillar clarity. Four labeled lines or one dense paragraph both work when SMCD is complete. Avoid prompt walls that repeat adjectives. Match motion beats to seconds. One beat per two to four seconds is a safe starting ratio on consumer tiers as of August 2026.
When should I use PromptMake /image for video work?
Use /image when you need a cleaner hero still before image-to-video, when you want Midjourney or FLUX formatted text from a reference photo, or when upload drift traces back to a weak source frame. Soft path: https://promptmake.net/image. After the still holds, paste SMCD motion and camera lines into your video host.
Which video model should I start with in 2026?
Start with the host you can access today and short duration. Luma and Kling fit quick image-to-video tests on three-to-five-second clips. Runway Gen-4 fits longer text-to-video when you need ten-second camera language. Hailuo fits short cinematic beats with clear timed verbs. SMCD travels across all four; only paste dialect and credit pricing change.
Ready to generate your own prompts?
Free. No sign-up required. Works with all major AI models.