PromptMake
2026-08-24·15 min read

Image to Video Prompt Guide: Still → Motion Language

Write an image to video prompt that animates a locked still: preserve clauses, motion deltas, camera paths, and hero-frame prep for Kling, Luma, and Runway.

image to video promptimage to videovideo promptsklinglumarunwayimage promptsguide

Turn any photo into an AI prompt — free

No sign-up required. Works with Midjourney, FLUX, DALL-E.

Try Image to Prompt →

An image to video prompt tells a motion model how a locked still should change across a few seconds. You upload a frame you already trust. Text carries only the delta: what moves, how the camera travels, and what must stay identical to the upload. Kling AI, Hailuo, Luma Dream Machine, Runway Gen-4, and Pika all read that pattern on image-to-video routes as of mid-2026. This guide teaches still → motion language so you stop pasting poster captions into video boxes. You leave with preserve clauses, timed motion verbs, camera-only recipes, a still-prep checklist, host notes, and a soft path to PromptMake /image when the hero frame needs cleanup before you spend video credits.

What an image to video prompt is for

Image-to-video starts from pixels. Text-to-video starts from nouns. That split decides every line you write. On an image-to-video route the upload already fixed subject, wardrobe, lighting, and composition. Your image to video prompt adds timed change on top of those pixels. Re-describing the jacket color or the sky gradient fights the still and invites identity drift.

You need this craft when you animate Midjourney or FLUX heroes, product packshots, portrait stills, or concept art that already passed a static review. Marketers use it for five-second social loops. Designers use it for pre-viz before a longer edit. Creators use it when a face or logo must match the source frame.

This page stays on still → motion language. The SMCD structure article on this blog covers subject, motion, camera, and duration as a general spine for any video brief, including blank-frame text-to-video. Generator product pages teach tool loops. Here you learn the dialect that sits between a finished still and a short clip: preserve, delta, camera, seconds.

Skip this page if you only export stills and never open a video host. Lean in if your last Kling or Luma run reinvented the face, invented side scenery, or ignored a slow push-in you thought you wrote.

Still → motion language: the core craft

Motion language for a locked still has three jobs. First, tell the model the upload is the authority. Second, name one or two changes that fit the clip length. Third, name one camera path or declare a locked tripod. Everything else is noise. Adjective stacks that worked for Midjourney posters waste tokens here because the pixels already carry style.

Think in deltas, not scenes. A delta is a verb with timing: "steam curls once from the rim between 2s and 4s." A scene rewrite is "cinematic coffee mug on oak sill, soft morning light, 8K." Scene rewrite belongs in the still tool. Delta belongs in the image to video prompt.

Hosts differ in paste length and slider names, but they share the same failure mode. When preserve language is weak and motion language is vague, the model treats your still as a mood board and rebuilds the frame. Strong preserve plus weak motion still beats weak preserve plus poetic motion.

Preserve clauses that hold identity

Lead with a short preserve block. Write what must not change. "Maintain subject, wardrobe, face, lighting direction, and background from the reference image" is a working opener on Kling, Luma, Hailuo, and Runway image-to-video. Add one identity anchor when faces matter: "Keep eye color, hair length, and jacket tone identical to upload."

Name forbidden changes when the host has drifted on you before. "No costume change. No new props. No text overlays. No background rebuild." Negative preserve lines cost few tokens and stop common inventions. Keep the list short. Five bans beat twenty synonyms.

For products, preserve material and logo state: "Brushed steel bottle, blue cap, blank side, no logo added." If the still already shows a logo, say "Keep logo geometry and placement from reference." If the still is blank on purpose, say "Do not invent branding." Ambiguous logo language is how you get melted letters across frames.

Motion deltas: verbs with seconds

Write motion as ordered beats matched to your UI duration. "0-2s subject static; 2-5s hair lifts once in a light breeze" gives Hailuo and Kling a readable spine. "Hair moves beautifully in the wind" gives the model room to thrash the face.

Prefer micro-motion on first passes: blink, breath, smile growth, fabric shift, steam curl, leaf quiver, condensation bead. Micro-motion proves the host respects the upload. After a clean three-to-five-second test, raise energy one notch: walk cycle, pour, door open, car approach.

Separate subject motion from camera motion in the same prompt. "Portrait locked; camera slow dolly in only" is clear. "Dynamic cinematic reveal" merges lens and subject into one mushy verb. Models that merge those channels invent spin that your still never implied.

One primary motion beat per two to four seconds is a safe ratio on consumer tiers as of August 2026. A four-second clip holds one beat. A six-second clip can hold establish stillness plus one action. Save montage lists for an editor after you generate separate takes.

Camera paths on a locked still

Camera-only prompts are the most reliable image to video prompt pattern for portraits and product heroes. The subject stays put. The lens moves. "Medium shot preserved; slow push-in from medium to close over 4s; tripod-smooth; no orbit" travels across Luma and Kling with fewer identity failures than hair wind plus orbit in one pass.

Pick one move: locked tripod, slow dolly in, gentle dolly out, lateral truck left or right, slight tilt up, handheld micro-shake. Write the framing start and end when the host supports it: "Start medium, end close on face." Skip crane-plus-orbit stacks on a first credit spend.

Lens feel belongs next to camera, not in a style dump: "50mm product feel, shallow depth preserved from source." Depth language tells the model to keep the blur already in the still instead of sharpening the background into a new scene.

Step-by-step: still to motion brief

Budget fifteen minutes the first time you turn a hero still into an image to video prompt. Later runs take five once your preserve shell and camera recipes live in a notes file. Keep one row per clip: still filename, preserve line, motion beats, camera path, seconds, host, date, pass or fail.

Crop before upload. Vertical portraits into horizontal video lanes force invented side walls. Match the still aspect to the host output before you write a single verb. Export the largest clean PNG or JPEG your still tool allows. Soft, noisy, or over-compressed uploads lose edges the motion model needs for consistency.

Treat the first generation as a smoke test. Score three questions: Did the subject hold? Did one camera path land? Did motion finish inside the seconds? Change one variable per retry. Rewriting preserve, motion, and camera together hides which line failed.

Step 1: Approve the still as a motion base

Run a static checklist before you open the video host. Subject readable at thumbnail size. Clean edges on the hero silhouette. Lighting direction obvious. No accidental text. Background simple enough that a slow push-in will not reveal garbage at the frame edge.

If the still fails the checklist, fix it in image tools first. PromptMake /image can draft Midjourney, FLUX, DALL·E, Stable Diffusion, or Leonardo text from a reference photo when you need a cleaner hero. Use Recreate Exactly to hold composition, Adjust Lighting when exposure blocks motion tests, Create Variation when you want alternate stills that share one motion shell. Soft start: https://promptmake.net/image. Guest image quota is about three runs per day; free registration raises that cap separately from /text.

Name the file with intent: product-bottle-front-v3.png beats final2.png. Your notes row should point at a frame you can find next week.

Step 2: Write preserve, then one delta

Open a notes app. Line 1: preserve clause. Line 2: one motion beat with seconds. Line 3: camera path or "camera locked." Line 4: duration that matches the host slider. Do not merge into prose until the four lines read clean aloud.

Example shell for a portrait upload on Luma: "Preserve uploaded portrait, olive jacket, overcast park. Motion: subtle smile grows over 4s, one blink at 2s. Camera: locked framing, no zoom. Duration: 4s." That shell is a complete image to video prompt before you polish adjectives.

Example shell for a product upload on Kling: "Maintain brushed steel bottle and blue cap from reference; blank side; no logo. Motion: 0-2s static; 2-5s one condensation bead slides down the left side. Camera: medium shot, slow dolly in only. Duration: 5s."

Step 3: Paste, generate, iterate one pillar

Upload the still. Paste the shell. Set the duration slider to match your beats. Generate once. If the face drifts, shorten motion and strengthen preserve. If the camera ignores your path, delete extra move words and keep one verb. If the clip freezes early, remove a beat or add one second in the UI.

Lower motion strength when the host exposes a slider and identity matters. Raise strength only after a clean micro-motion pass. Log the exact text that worked next to the filename. Next week you swap the still and keep the motion grammar.

After two wins on the same shell, save bracket variables: [preserve anchors], [motion beat], [camera move], [seconds]. That template is your personal image to video prompt kit without a paid prompt vault.

Worked image to video prompt examples

Three teaching copies below share one discipline: upload carries subject, text carries delta. Adapt nouns. Keep preserve first and camera singular.

Example A, portrait on Luma Dream Machine: Upload a medium outdoor portrait. Prompt: "Maintain reference portrait exactly, including hair length, jacket color, and background trees. Subtle smile grows over four seconds with one natural blink at two seconds. Camera locked, no zoom, shallow depth preserved from source." Duration: 4s. Goal: identity hold with micro-expression only.

Example B, product hero on Kling AI: Upload a centered bottle still on gray stone. Prompt: "Preserve brushed steel water bottle with blue cap from reference; no logo; studio void background unchanged. Bottle static for two seconds, then one condensation bead slides down the left side once. Slow dolly in only, 85mm product feel, soft key from upper left preserved." Duration: 5s. Goal: product identity plus one readable beat.

Example C, concept art on Runway Gen-4: Upload a dusk street still with a yellow taxi mid-frame. Prompt: "Keep wet asphalt, neon soft background, and taxi colors from reference. Taxi approaches slowly for three seconds, brake lights brighten once at three seconds. Low wide framing held; static tripod; 24mm feel from source." Duration: 6s. Goal: scene continuity with timed vehicle motion.

Weak contrast for study: "Cinematic animation of this image, beautiful motion, epic camera, 8K, ten seconds." No preserve, no timed verb, no single camera path, duration longer than any planned beat. Fix that line by adding preserve, one delta, one camera move, and honest seconds before you spend another credit.

Common image to video prompt mistakes

Mistake 1: Pasting a full still caption into the video box. The model already sees the pixels. Caption rewrite invites a rebuild. Keep text on delta and preserve.

Mistake 2: Asking for motion the still cannot support. A profile portrait cannot "turn to face camera" without inventing the missing side of the head. Pick deltas that fit the visible pose.

Mistake 3: Two camera moves in one pass. Orbit plus dolly plus tilt collapses on Kling and Hailuo. One path per generation.

Mistake 4: Duration longer than beats. Ten seconds of UI time with one blink fills the tail with mush or frozen frames.

Mistake 5: Aspect mismatch. Crop the still to the video lane before upload so the model does not invent side scenery.

Mistake 6: Raising motion strength before a clean micro-pass. Strength amplifies whatever the text already asked for, including drift.

Mistake 7: Changing still, text, and duration in the same retry. Isolate one variable so you learn which line failed.

Mistake 8: Treating this guide as a full SMCD course for blank-frame text-to-video. When you have no upload, you need full subject prose. Use the structure article for that spine; use this page when the still is already locked.

Model notes for image-to-video hosts (2026)

Public names and credit rules shift. Treat the notes below as an August 2026 snapshot. Confirm pickers inside Kling AI, MiniMax Hailuo, Luma Dream Machine, Runway, and Pika before you paste month-old forum prompts. Shared rule: preserve first, one motion beat, one camera path, seconds matched to the slider.

Image-to-video quality tracks still quality more than adjective count. Spend prep time on the hero frame. Spend prompt tokens on delta clarity. Free and trial tiers rotate daily credits; burn them on four-to-five-second tests with honest shells, not on twelve-second wish lists.

When a host adds native audio in 2026 builds, treat sound as optional after motion holds. Write "ambient room tone only" or "no music" so a quiet product clip does not gain a stock score.

Kling AI and Hailuo

Kling AI remains a strong pick for product and character stills on image-to-video as of mid-2026. Separated clauses work well: preserve block, motion block, camera block. Lead with maintain language. Keep camera to one move. Timed beats help: "0-2s static, 2-4s fabric shifts once."

Hailuo from MiniMax rewards clear motion verbs on short cinematic clips. Micro-motion first. Avoid montage lists inside one pass. Chain clips in an editor when you need multiple beats. Both hosts punish conflicting energy words in the same prompt. Pick calm or bold once in the motion line.

Luma Dream Machine, Runway Gen-4, and Pika

Luma Dream Machine favors readable sentences on three-to-five-second image-to-video tests. Subtle motion on uploads lands cleaner than action verbs the still never showed. Locked camera plus smile or blink is a reliable first recipe.

Runway Gen-4 tolerates richer camera vocabulary and longer clips on paid tiers. SMCD-style pillars still apply, but on image-to-video you shrink the subject pillar to preserve language. Use Gen-4 when you need six-to-ten-second moves with explicit lens cues after a clean still export from the same suite.

Pika stays useful for quick social loops and stylized motion. Keep prompts short. One preserve sentence plus one motion sentence plus one camera sentence outperforms a paragraph of film-school adjectives on short Pika runs.

Prep the hero still with PromptMake /image

Video hosts consume pixels. PromptMake /image helps when the still is wrong before motion starts: noisy phone photo, weak Midjourney crop, blown highlights, or a packshot that needs dialect text for another still pass. Soft path: https://promptmake.net/image.

Workflow: upload a reference or iterate until the frame is thumbnail-clear. Copy Midjourney, FLUX, DALL·E, Stable Diffusion, or Leonardo formatted text when you need another still generation. Lock composition with Recreate Exactly when image-to-video identity depends on layout. Adjust Lighting when motion can wait but exposure cannot. Create Variation when you want three hero options that share one motion shell.

Guest accounts get about three /image runs per day. Free registration raises the image quota. Quota is separate from /text. Spend image runs on subject lock. Write your image to video prompt offline so you do not burn still quota hunting synonyms for motion you have not planned.

PromptMake does not render video clips. It shortens the still step so preserve and delta lines attach to a frame worth animating. Paste the finished motion brief into Kling, Luma, Hailuo, Runway, or Pika beside the export.

FAQ

These questions match searches people type after a first image-to-video failure: what an image to video prompt is, how still language differs from text-to-video, which motion verbs to start with, how to stop face drift, where PromptMake /image fits, which host to try first in 2026, and how long a first clip should run. Each answer stays short enough to act on in the same session.

What is an image to video prompt?

An image to video prompt is the text you paste beside an uploaded still to tell a motion model what should change over a few seconds. It focuses on preserve language, timed subject motion, one camera path, and duration that matches the host slider. It does not rebuild the scene from scratch. Kling, Hailuo, Luma, Runway, and Pika all expect that delta-first pattern on image-to-video routes.

How is still → motion language different from text-to-video?

Text-to-video needs full subject nouns because no upload exists. Image-to-video inherits subject pixels from the still, so text carries preserve clauses and deltas. Re-describing wardrobe and lighting on an image-to-video route invites identity drift. Write change, not a second poster caption.

What motion should I try first on a portrait still?

Start with micro-motion and a locked camera: subtle smile growth, one blink, slight breath in the shoulders, three to five seconds. That recipe tests whether the host respects the upload. Add hair wind or a slow push-in only after a clean identity pass. Avoid asks that require unseen geometry, such as a full head turn from a tight profile.

Why does my face drift on Kling or Luma?

Face drift traces back to weak preserve language, motion that exceeds the pose, mismatched aspect ratio, or motion strength set too high. Strengthen maintain clauses, shorten the beat list, crop to output aspect, and lower strength. Change one of those variables per retry so you can see which fix held the face.

When should I use PromptMake /image before video?

Use /image when the hero still is noisy, badly lit, or wrong in composition for motion, or when you need Midjourney or FLUX formatted text from a reference photo before another still pass. Soft path: https://promptmake.net/image. After the still holds, write your image to video prompt with preserve and delta lines, then paste into the video host.

Which image-to-video host should I start with in 2026?

Start with the host you can access today and a four-to-five-second clip. Luma and Kling fit quick portrait and product tests; Hailuo fits short cinematic beats with clear timed verbs. Runway Gen-4 fits longer camera language after the still is locked, and Pika fits fast social loops. Still → motion language travels across all of them; only paste dialect and credits change.

How long should my first image to video prompt clip be?

Start at three to five seconds with one motion beat and one camera path. Longer clips need more temporal budget and fail more often when the still is new to you. Extend duration only after a clean short pass, and add at most one extra beat when you add seconds. Match the UI slider to the seconds you wrote in the prompt.

Ready to generate your own prompts?

Free. No sign-up required. Works with all major AI models.

Related articles