Pacific Design/ artificial intelligence

Generative Media · entry 02/04

Image generation in practice

Prompt structure, negative prompts, and structural control turn image generation from a slot machine into a directable tool.

Anatomy of an image prompt

An image prompt is a shot brief, not a wish. It rewards the same discipline as any structured prompt: name the subject, then the medium, style, lighting, and composition, in the vocabulary a photographer or art director would use. Vague quality words ("beautiful", "8k") do little; concrete craft words ("backlit", "35mm", "impasto") steer hard.

subject:  lighthouse keeper reading by lamplight
medium:   oil painting, heavy impasto
lighting: single warm source, deep shadows
compose:  wide shot, subject low-left, negative space above
style:    muted palette, Hopper-like stillness
negative: text, watermark, extra fingers
frame:    1536x1024, 3:2

Negative prompts subtract failure classes instead of adding content. Aspect ratio is part of the idea itself: the same words compose differently in a 9:16 portrait than in a 21:9 cinematic frame, so choose the canvas before polishing the words.

Control beyond text

Language underspecifies geometry, so text-only control caps out fast. The working toolkit goes further: image-to-image regenerates an existing frame at a chosen strength; inpainting redraws a masked region while outpainting extends past the canvas; structural control feeds the model a pose skeleton, depth map, or edge sketch it must respect while the prompt fills in the surfaces. Style-reference images pin a look far more reliably than style words, and character-consistency features hold one face or mascot stable across scenes — the capability that actually unlocks storyboards, comics, and brand work.

Generate wide, select, refine

Treat generation as cheap and selection as the job. Produce a dozen candidates, keep the one or two with good bones, then refine: inpaint the broken region, change one prompt clause at a time, and reuse the seed so each change is attributable to its cause. This is a pipeline, not a slot machine — the illustrations on this site are made exactly this way, from a versioned prompts file per page, generated wide and curated down. When iterations refuse to converge, the cause is usually a specification failure — two ideas fighting for one frame — rather than a model limitation.

Failure mode

Asking text to do a controller's job. The stable weak spots are well known: legible text in images (better in 2026, still unreliable past a few words), hands (historically cursed, now mostly fixed, still worth checking), exact counts (ask for five arrows, receive seven), and precise spatial logic ("the red cube left of the blue one"). More adjectives will not fix these, because the model composes statistics, not scene graphs. Route around them instead: set real typography in post, use a depth map or sketch for layout, inpaint the specific error. The expensive version of this mistake is burning fifty generations on wording when one structural reference would have ended the argument.