How to Write AI Prompts for Images (Step-by-Step)
How to write AI prompts for images step by step: lock subject, scene, light, and medium, then format for Midjourney, FLUX, GPT Image, and more.
Free AI Prompt Generator — no sign-up required
Works with ChatGPT, Claude, Midjourney, FLUX, Sora, and more.
Learning how to write AI prompts for images means writing a short picture brief before you open Midjourney v7, FLUX, GPT Image, Ideogram, or SDXL. You name the subject, the place, the light, the medium, and one constraint line, then you format that brief for one host. This page walks the craft in order: decide the job, lock nouns, set scene and light, pick one medium, format the paste, score the first grid, and change one line per retry. You leave with a reusable write loop, worked paste examples, mid-2026 dialect notes, mistake fixes, and soft paths to PromptMake /text for blank notes and /image when a reference photo already locks the subject.
Who this step-by-step image prompt craft is for
You need this page if you type one vague wish into an image model and hope the grid looks right. Marketers, freelancers, product designers, and social creators who hop tools in one week get the most value. The skill is the same across hosts: write picture facts first, dialect second.
This guide stays on still image prompts. Timed camera paths and duration language belong on video prompt pages. ChatGPT, Claude, or Gemini text contracts for email and research belong on the beginner how-to-write AI prompts guide in the prompt-engineering cluster. That article teaches Outcome, Audience, Materials, and Shape for language models. Image craft uses a different vocabulary: subject materials, place depth, light direction, and a single medium.
Skip deep SPARC teaching chapters if you want a portable five-label sheet across five models; that lives on the AI prompts for image generation guide. Skip paste-kit libraries if you want ready product and portrait starters; those live on the best AI prompts for images guide. Stay here when you want the write order from a blank page to a scored first generation.
What you decide before you open any image model
Image models invent whatever you leave blank. A one-word subject invites random wardrobe, cluttered rooms, mixed media, and surprise text watermarks. Your job is to close those gaps in plain English before you spend credits. Write the decisions offline in a note. Read them aloud. If you can picture the still, the draft is ready to format.
Think of the note as a shoot brief, not a poem. Concrete nouns beat style adjectives. "One brushed steel water bottle, blue cap, no logo" moves pixels. "Epic cinematic masterpiece bottle" spends tokens and leaves the model free to invent a second label, a neon alley, and a chrome cartoon finish in the same pass.
Hold a reference photo when the subject must match a real packshot, face, or room. Text then adds light changes or medium shifts. Invent from scratch when the brief is a concept still with no source pixels. Both paths use the same write order. Only the source of the subject nouns changes.
Job, subject, and place
Start with the job in one sentence: product hero for a PDP, chest-up documentary portrait, food flat lay for social, landscape keyframe for a deck. The job picks canvas and medium before you name light. Vertical 4:5 fits product and portrait. Square 1:1 fits social flat lays. Wide 16:9 fits landscape and slides.
Subject is the noun stack that must survive thumbnail size. Name material, color, scale, and count. For people, lock hair color, wardrobe tone, and age band once. One primary subject per still unless the brief is a deliberate duo. Crowds need count language: three cyclists mid-ground, one hero rider sharp in a red jersey.
Place grounds the subject. Without Place you get gray void or random clutter. Write location type, time of day or era, and one depth cue: sunlit oak counter, empty aside from the mug, soft window behind. Product catalog jobs can set Place as white seamless backdrop on purpose. That is a Place choice, not silence.
Light, medium, and constraints
Light moves pixels. Direction, quality, and temperature matter. Soft window light from camera left, cool morning fill, gentle shadows works. Beautiful lighting does nothing. Named setups help photographic jobs: Rembrandt, butterfly, golden hour, overcast ambient. Illustration jobs replace physics with shading words: flat even color, hard cel shade, soft pigment wash.
Medium picks one rendering universe. Documentary photograph, editorial fashion, product catalog, oil painting, ink illustration, flat vector poster, clay 3D, anime cel. One medium per pass. Mixed media such as photoreal watercolor chrome cartoon fight each other on Midjourney, FLUX, and GPT Image alike.
Constraints hold canvas and exclusions. Aspect intent belongs here. Must-avoid lists live here: no text, no watermark, no extra limbs, no logo. Host flags live here too when you format later: Midjourney --ar and --style raw, SDXL negatives, Ideogram quoted type, GPT Image plain-language crop. Fill exclusions before you invent Midjourney flags so you do not paste dialect into an empty brief.
Step-by-step: how to write AI prompts for images
Budget twenty minutes the first time. Later runs take five when your note shell exists. Keep a dated row: job, subject, place, light, medium, constraints, host name, winner filename. That row becomes your reusable craft log. The steps below assume a still image job on Midjourney v7, FLUX.1.x / Flux 2 family builds, GPT Image inside ChatGPT, Ideogram v3 or v4, or Stable Diffusion XL as of mid-2026. Confirm public model names in your app picker; badges move.
Do the write work offline first. Chat iteration inside GPT Image can refine a still after you paste a clear brief. Chat without a brief turns into twenty vague edits with no lesson. One scored generation teaches more than a pile of unlogged grids.
Steps 1 to 3: job, nouns, and scene
- Write the job in one plain sentence. Name the use: PDP hero, LinkedIn-safe portrait, Instagram flat lay, slide background. Circle the canvas: 4:5, 1:1, or 16:9.
- Expand Subject. Materials, color, count, and one distinctive mark. "Matte white ceramic mug, blue rim chip on the left, no print" beats "coffee mug." For a person: adult woman, short dark hair, olive jacket, chest-up crop.
- Write Place and Light as separate lines. Place: location, time, depth. Light: direction, quality, temperature. Read both aloud. If Place and Light blur into "nice kitchen vibes," split them until each line names a fact a stranger could check on the PNG.
Stop here if Subject still feels vague. Fix nouns before you touch medium adjectives. Models weight early subject tokens on Midjourney and still favor clear head nouns in FLUX and GPT Image sentences.
Steps 4 to 6: medium, format, and one-line retry
- Lock one Medium. Photoreal catalog, documentary photo, watercolor wash, flat vector poster. Delete any second medium word that sneaked into Subject or Place.
- Format for one host only. Midjourney: compact phrase stack plus trailing
--ar,--style raw,--v 7, and--noexclusions. FLUX: full sentences with the same facts; set aspect in the UI. GPT Image: lead with Create or Keep, then sentences. Ideogram: quote exact lettering in Constraints when type must sit in the frame. SDXL: positive line for picture facts; negative field for must-avoid items. - Generate once. Score five checks: Did Subject hold? Did Place clutter appear? Did Light match direction? Did Medium stay single? Did Constraints exclusions fire? Change one line per retry. Wrong light: edit Light only. Extra props: tighten Place and must-avoid. Wrong medium: swap Medium, leave Subject alone. Finish two clean retries on one host before you reformat for another.
Blank-page writers can draft those labeled lines in PromptMake /text and ask for Midjourney, FLUX, DALL·E, or Stable Diffusion shaped output once the picture facts exist. Soft start: https://promptmake.net/text. Guest quota is about three text runs per day; free registration raises the text quota. Text and image quotas stay separate.
Reference photo path: upload to PromptMake /image when a packshot, portrait, or location still already locks Subject and Place. Pick Midjourney, FLUX, DALL·E, Stable Diffusion, or Leonardo as the paste target. Recreate Exactly when you want near-match language. Change Style when the subject stays but Medium moves. Adjust Lighting when framing is fine and the key is wrong. Create Variation for sibling stills with the same subject shell. Soft start: https://promptmake.net/image. Guest image quota is about three runs per day; registration raises image quota on its own track.
Worked examples you can adapt today
Three teaching copies share one discipline: full picture facts before dialect. Adapt nouns to your product. Keep write order when you edit. Save the winning paste beside the PNG with the job in the filename so the next round does not reopen a park path and break a catalog set.
Example A, product hero. Job: PDP vertical still. Subject: brushed steel water bottle, blue cap, no logo, single object. Place: gray stone surface, empty studio void. Light: soft key from upper left, gentle rim light. Medium: catalog product photo. Constraints: 4:5, no text, no watermark, no logo. Midjourney paste: brushed steel water bottle, blue cap, no logo, gray stone surface, soft key from upper left, gentle rim light, catalog product photo --ar 4:5 --style raw --v 7 --no text, watermark, logo. FLUX paste: A brushed steel water bottle with a blue cap and no logo sits on a gray stone surface. Soft key light from upper left and a gentle rim light. Photoreal catalog product look, shallow depth, no text, no watermark. Set vertical aspect in the FLUX UI.
Example B, documentary portrait. Job: chest-up social portrait. Subject: adult woman, olive jacket, short dark hair. Place: overcast park path. Light: soft cool daylight. Medium: documentary photo. Constraints: 4:5, no text, no watermark. GPT Image create: Create a chest-up documentary photo of an adult woman in an olive jacket with short dark hair on an overcast park path. Soft cool daylight, natural skin, no text, no logo. Keep wardrobe and hair if you upload a reference and ask for an edit instead of a create.
Example C, poster with type. Job: bakery promo still with readable headline. Subject: open bakery storefront at dusk. Place: city sidewalk, warm interior glow. Light: blue hour outside, tungsten inside. Medium: flat graphic poster. Constraints: 16:9, headline reading "FRESH DAILY" in heavy cream letters. Ideogram owns the quoted type. Midjourney versions leave negative space for type in a design tool when lettering must stay perfect.
Weak contrast for study: epic cinematic masterpiece bottle, 8K, trending on artstation, ultra detailed. No materials, no place, no light direction, mixed medium signals, empty exclusions. The six write steps fix that line without folklore adjectives.
Model dialect notes for 2026
Confirm model pickers inside your apps as of mid-2026. Public names shift. The notes below assume Midjourney v7 (plus V8 Alpha where you have access), FLUX.1.x / Flux 2 family builds, GPT Image inside ChatGPT, Ideogram v3 or v4, Stable Diffusion XL, and Gemini image workflows when your Google account exposes them. Picture facts stay fixed. Only Constraints and sentence shape change.
Shared rule: write Subject through Medium first, then format. Do not invent Midjourney flags for FLUX. Do not invent SDXL weight parentheses for GPT Image. Do not ask Midjourney to spell long brand wordmarks when Ideogram owns lettering. Pick the host for the job after the brief is complete.
Midjourney v7 and FLUX
Midjourney v7 rewards compact phrase stacks. Order: Subject, Place, Light, Medium, then Constraints as trailing parameters. Keep adjective count low. Put --sref or --cref in Constraints when you lock a house look or character across a series after the text brief is solid. Use --style raw for photographic jobs.
FLUX family models prefer natural-language sentences with the same fact order. Drop Midjourney flags. Set aspect in the FLUX UI. FLUX handles material and light prose well. Overlong style name lists still hurt. Use Midjourney when aesthetic direction and reference flags matter. Use FLUX when photoreal materials and API-friendly prose matter.
GPT Image, Ideogram, SDXL, and Gemini
GPT Image inside ChatGPT reads conversational create or edit language. Lead with the job verb, then picture facts in sentences. Edit jobs lead with preserve language when you upload a still: Keep the bottle shape and cap; change light to warm tungsten from the right.
Ideogram v3/v4 owns readable type inside the frame. Put exact words in quotes in Constraints when the still needs lettering. Subject and Place still come first. SDXL wants a positive prompt plus a negative prompt field. Encode Subject through Medium in the positive line. Put must-avoid items in the negative: text, watermark, logo, extra fingers, blurry subject.
Gemini image workflows take clear scene sentences and follow-up edits. Keep the same fact order. Ask for one medium only. Language models such as GPT-5.6 Sol, Claude Fable 5, or Gemini 3.1 Pro can draft brief lines from notes; you still approve every noun before you generate pixels.
Common mistakes when you write image prompts
Mistake 1: A wish with no subject nouns. Beautiful scene spends tokens and moves no pixels. Lead with materials and count.
Mistake 2: Mixed medium in one pass. Photoreal plus watercolor plus 3D cartoon collapses on every host in this guide. Pick one.
Mistake 3: Host flags pasted into the wrong model. Midjourney --ar inside ChatGPT Image does nothing useful. SDXL negatives inside Ideogram confuse lettering jobs.
Mistake 4: Rewriting the whole brief when one line failed. Isolate Light if light broke. Isolate Constraints if exclusions failed.
Mistake 5: Asking Midjourney to spell long brand wordmarks. Send quoted type to Ideogram or set type in a design tool on a clean plate.
Mistake 6: Skipping Place on purpose then blaming the model for clutter. Empty seamless is a Place line. Silence is not.
Mistake 7: Video motion verbs in a still prompt. Slow dolly in belongs in video briefs. Stills need light and framing, not camera paths over time.
Mistake 8: Treating text prompt frameworks as image craft. Outcome and audience labels help ChatGPT write email. They do not name materials, light direction, or medium for Midjourney. Use the image write order on this page for pixels.
When PromptMake /text and /image fit
PromptMake splits blank-page work from photo-grounded work. Use /text when your brief starts as rough notes and you want model-shaped prompt text for Midjourney, FLUX, DALL·E, or Stable Diffusion. Use /image when a photo or sketch already locks Subject and Place and you need formatted prompt language plus a goal mode.
CTA both exists because many creators alternate in one day. Morning: /text turns a product brief into labeled lines and a Midjourney paste. Afternoon: /image reads a packshot and returns FLUX prose under Recreate Exactly. Soft links: https://promptmake.net/text and https://promptmake.net/image.
Guest accounts receive about three runs per day on each path. Free registration raises each quota. Quotas do not share a pool. PromptMake outputs prompt text. You still generate pixels in Midjourney, FLUX, ChatGPT Images, Ideogram, Gemini image tools, or your SDXL host. Skip PromptMake when you already hold a perfect paste and need one Light word changed. Open the generator and edit that line.
FAQ
How do I write AI prompts for images if I am new?
Write a short picture brief before you open the generator. Name subject materials and count, place, light direction, one medium, and a must-avoid list. Format that brief for one host. Generate once and change one line if the grid misses. That is how to write AI prompts for images without chasing viral one-liners.
What is the first line I should write?
Start with the job and the subject noun stack. A PDP hero and a brushed steel bottle beat a pile of style adjectives. Place and light come next. Medium and constraints close the brief. Host flags come last, after the picture facts read clean out loud.
Do the same image prompts work on every model?
The same picture facts work. The paste dialect changes. Keep Subject through Medium identical across hosts. Change Constraints and sentence shape per model. Midjourney wants compact phrases and trailing flags. FLUX and GPT Image want natural sentences. Ideogram wants quoted type when lettering matters. SDXL wants a negative field.
How is this different from how to write AI prompts for ChatGPT text?
Text beginner guides teach contracts for language models: outcome, audience, materials, shape. Image craft names subject materials, place depth, light direction, and a single medium. Use the text framework for email and research. Use this step-by-step for Midjourney, FLUX, GPT Image, Ideogram, and SDXL stills.
How long should an AI image prompt be?
Long enough to cover subject, place, light, medium, and constraints, and short enough to avoid medium conflicts. Most Midjourney pastes stay under a few dozen strong tokens plus flags. FLUX and GPT Image tolerate fuller sentences when each sentence adds a picture fact. Cut words that do not change materials, light, medium, or exclusions.
How do I start for free with PromptMake?
Open https://promptmake.net/text with a rough image brief, or https://promptmake.net/image with a reference photo. Guest use is about three runs per day on each tool path. Free registration raises those quotas. Text and image limits stay separate. Copy the formatted prompt into Midjourney, FLUX, GPT Image, Ideogram, Gemini image tools, or SDXL to generate the still.
Which image model should I practice on first?
Pick from the job, not from hype. Midjourney v7 for aesthetic direction and reference flags. FLUX for photoreal materials and natural-language workflows. GPT Image for conversational create and edit inside ChatGPT. Ideogram when readable type must live in the pixels. SDXL when you need negatives and local control. Write the brief once, then format for that host.
Ready to generate your own prompts?
Free. No sign-up required. Works with all major AI models.