Control Lighting, Expression and Depth of Field Without Losing the Scene

AI Image Composition Guide

Contents

AI image composition improves when prompts stop behaving like bags of adjectives. Lighting, expression and depth of field are different production variables. Change them deliberately, keep the rest of the scene stable and compare results against a clear objective. The method applies across contemporary image generators; the exact names and availability of controls vary by platform.

This approach is slower than adding “cinematic, dramatic, emotional, beautiful.” It is also much easier to learn from.

Lock the Composition Before Styling It

Start with the spatial decision:

  • subject and action;
  • shot size and camera angle;
  • position in frame;
  • foreground, middle ground and background;
  • important object relationships; and
  • what must remain visible.

Example base prompt:

Eye-level medium close-up of a courier seated beside a train window, body turned slightly toward the aisle, sealed letter held in both hands, station lights visible behind the glass, shoulders and hands in frame.

Generate this neutral version first. If the body, letter or window relationship fails, dramatic lighting will not repair the composition.

Control Lighting as Structure

Lighting should explain where illumination comes from and what it reveals.

Soft window light

Use a broad source and moderate contrast when facial information should remain open and readable:

soft overcast window light from camera left, gentle facial modelling, controlled highlights, readable detail in shadows

Low-key dramatic light

Use a narrow or obstructed source for tension:

narrow warm platform light crossing the face, deep but detailed shadow on the far cheek, dark carriage interior, practical light source visible in reflection

High-key light

High-key lighting reduces shadow contrast. It can suggest openness, commercial clarity or an exposed, almost clinical emotional state.

Avoid combining “high-key,” “deep noir shadows” and “flat light” unless the contradiction is intentional and spatially explained.

Bind Facial Expression to an Event

“Sad expression” names an output but not a performance. Give the expression a cause and a degree of suppression:

She recognizes the handwriting on the sealed letter and tries not to react; lower eyelids tighten, lips remain closed, breath held, gaze fixed on the signature.

The model still interprets the instruction probabilistically, but the prompt now distinguishes restrained recognition from generic sadness.

Use a small expression vocabulary:

  • perception: eyes find or track something;
  • appraisal: brow, focus and pause indicate interpretation;
  • regulation: the character hides or releases a reaction;
  • action tendency: the body prepares to leave, confront or withdraw.

Use Depth of Field to Allocate Attention

Depth of field is not a synonym for quality. It determines how much depth remains legible.

Use shallow depth of field when one plane matters:

shallow depth of field, eyes and letter edge sharp, station lights softly defocused, background identities unreadable

Use deeper focus when spatial relations matter:

deep focus, courier, reflected platform clock and approaching figure all readable, layered composition

If the story depends on a person in the background, extreme bokeh may destroy the information the shot needs.

Change One Variable at a Time

For a controlled comparison:

  1. Lock the model and model version, dimensions, aspect ratio and exposed generation settings.
  2. Record the seed when the platform exposes one, while remembering that software and pipeline differences can still change output.
  3. Keep the base prompt and negative constraints stable.
  4. Change only lighting, expression or focus language.
  5. Compare identity, composition and object state before judging aesthetics.

This is the difference between prompt experimentation and random variation. The result tells you which instruction changed the image.

Add Structural Conditioning When Text Is Not Enough

Text prompts are weak at guaranteeing exact geometry. When the workflow supports it, image-to-image conditioning, masks, pose guidance, depth maps or edge conditioning can constrain the composition more directly. ControlNet-style conditioning is useful precisely because it adds a structural signal beyond text.

Controls are boundary conditions, not a transparent production script. The model still synthesizes the result internally, but the range of acceptable compositions becomes narrower.

A Complete Prompt Template

Subject and event:
[who, what happened, what changes]

Composition:
[shot size, angle, frame position, spatial relations]

Lighting:
[source, direction, contrast, color, shadow detail]

Expression:
[perception, appraisal, restraint, action tendency]

Focus:
[sharp plane, readable secondary plane, background treatment]

Invariants:
[identity, wardrobe, objects, room layout, camera position]

The best composition prompt is not the longest one. It is the one that separates decisions clearly enough to test them.

Sources: