AI Image Prompts: How to Write Them
A good image prompt is a clear brief, not a magic spell. Here's the structure that works, copy-paste examples across styles, and how to adapt for Midjourney, DALL-E, and Stable Diffusion.
Last updated June 21, 2026
The anatomy of a strong image prompt
Every effective image prompt answers the same questions a photographer or art director would: what is the subject, what is it doing, where is it, how is it lit, and how should it feel? Think of the prompt as seven slots, subject, action, setting, style/medium, lighting, camera/framing, and mood, and fill the ones that matter for your shot. You don't need all seven every time, but the more concrete decisions you make, the fewer the model makes randomly. Put the subject first so the model anchors on it, then layer modifiers. When you add a word, ask whether it actually changes the picture; if not, cut it. Bloated prompts dilute the model's attention and often produce muddier results than a tight one.
Copy-paste example prompts by style
These are starting points you can paste and then tweak. Swap the subject, adjust the lighting, and change the aspect ratio for your use case.
- Photorealistic portrait: “Close-up portrait of an older fisherman with a weathered face, soft window light from the left, shallow depth of field, 85mm lens, muted color palette, calm and dignified mood, highly detailed.”
- Product shot: “A matte black ceramic coffee mug on a polished concrete surface, studio lighting with a soft key light and gentle rim light, clean minimal background, slight steam rising, commercial product photography, sharp focus.”
- Cinematic landscape: “A lone cabin in a pine valley at blue hour, fog rolling between the trees, warm glow from the cabin windows, wide establishing shot, cinematic color grade, moody and quiet.”
- Flat illustration / logo style: “Flat vector illustration of a fox curled asleep, two-tone orange and cream palette, bold simple shapes, no gradients, centered on a plain background, modern minimalist branding style.”
- Concept art: “Digital concept art of a floating market city above the clouds, airships docking at wooden platforms, warm sunset light, painterly brushwork, sense of scale and wonder, fantasy illustration.”
- Isometric / 3D: “Isometric 3D render of a cozy home office, soft global illumination, pastel color scheme, tiny detailed props, clean studio background, blender-style render.”
Tool-specific tips: Midjourney, DALL-E, and Stable Diffusion
Midjourney leans aesthetic and stylized. Use comma-separated descriptors, control shape with --ar 16:9 (aspect ratio), nudge artistic flourish with --stylize, and exclude elements with --no text. Stable Diffusion gives you the most control: a dedicated negative-prompt field, term weighting, fixed seeds for reproducibility, and ControlNet for poses and composition, at the cost of more tuning. DALL-E and GPT-4o image generation are the most forgiving of plain English, the strongest at following literal instructions, and the best at rendering legible text in signs and posters. The practical rule: write your description once, then adapt the syntax for whichever tool you're using rather than expecting a single prompt to be optimal everywhere.
Iterating: fix problems instead of restarting
The first render is a draft. If hands or text come out wrong, add those terms to a negative prompt or ask for “natural relaxed hands.” If the style is off, name a clearer medium (“watercolor,” “35mm film photo,” “flat vector”). If composition is wrong, specify framing (“wide shot,” “centered,” “rule of thirds”). Generate several variations and keep the cleanest rather than chasing one perfect output. Change one variable at a time so you learn what each word does, that habit turns prompting from guesswork into a repeatable skill.
FAQ
What makes a good AI image prompt?
A good AI image prompt is specific and structured. It names the subject, the action or pose, the setting, the art style or medium, the lighting, the camera or framing, and the mood. Vague prompts like 'a dog' produce generic results; 'a golden retriever puppy sitting in tall grass at golden hour, shallow depth of field, 85mm lens, warm backlight' gives the model concrete decisions to make. Add the subject first, then modifiers, and remove words that don't change the picture.
How long should an AI image prompt be?
Long enough to be specific, short enough that every word earns its place, usually one to three sentences, or a comma-separated list of 8 to 20 descriptors. Midjourney and Stable Diffusion respond well to keyword-style lists, while DALL-E and GPT-4o image generation handle natural-language sentences. If you keep adding adjectives and the image stops changing, you've hit the model's attention limit; trim instead of piling on more.
What is a negative prompt?
A negative prompt tells the model what to leave out, for example 'no text, no watermark, no extra fingers, not blurry.' Stable Diffusion has a dedicated negative-prompt field, which is one of its biggest advantages for fixing recurring artifacts. Midjourney uses the --no parameter (e.g. --no text). DALL-E and GPT-4o image generation don't expose a formal negative field, so you phrase exclusions in plain language inside the prompt instead.
Do the same prompts work in Midjourney, DALL-E, and Stable Diffusion?
The core description transfers, but the syntax and strengths differ. Midjourney favors aesthetic, stylized output and uses parameters like --ar for aspect ratio and --stylize. Stable Diffusion gives the most control (negative prompts, weights, seeds, ControlNet) but needs more tuning. DALL-E and GPT-4o image generation are the most forgiving of plain English and strongest at following literal instructions and rendering readable text. Expect to adapt a prompt slightly per tool rather than paste it unchanged.
Why does AI keep getting hands, text, or faces wrong?
Hands, fine text, and symmetrical faces are historically hard because they have strict structure the model only approximates. Newer models (GPT-4o image generation, Nano Banana Pro, recent Stable Diffusion and Midjourney versions) are much better at text and hands than earlier ones. To improve results, keep hands away from the focal point, request 'natural relaxed hands,' add problem terms to a negative prompt, and generate several variations, then pick the cleanest rather than expecting one perfect render.
Can I use AI-generated images commercially?
Often yes, but it depends on the tool's terms and your jurisdiction. Midjourney, DALL-E, and most hosted Stable Diffusion services grant commercial use to paying users, with conditions. Copyright status of AI images is still unsettled in many countries, and you should avoid prompting for living artists' names or trademarked characters for commercial work. Always check the specific platform's current license and don't assume a free tier grants the same rights as a paid one.
Related: AI image styles.