Expression is where AI character photos most often go subtly wrong: faces that are technically her but emotionally vacant, or smiles cranked to toothpaste-commercial intensity with nothing behind them. The fix is vocabulary — expression prompting has its own craft, distinct from scene prompting.
Why "smiling" fails
"Smiling" is to expressions what "beautiful" is to lighting: a flag that produces the averaged, posed version of the thing — symmetrical, full-teeth, eyes uninvolved. Real feeds are made of in-between expressions: the smile arriving or fading, attention caught mid-thought, laughter past its peak. Models render these well — but only when asked with specificity.
The vocabulary that works
Process beats state. Describe the expression as something happening:
- "mid-laugh, eyes closed" / "suppressing a laugh"
- "smile starting to form" / "soft smile fading"
- "caught off guard, about to speak"
- "concentrating, slight frown" / "reading something amusing on her phone"
Attention direction does half the emotion. Where she's looking is the cheapest authenticity lever:
- "looking off-camera at something" → candid
- "looking just past the lens" → editorial intimacy
- "glancing back over her shoulder" → motion and narrative
- "eyes on the task, not the camera" → the documentary register that reads as a phone photo
Physical anchors carry feeling. Emotions live in the body: "shoulders relaxed, head tilted," "hugging the mug with both hands," "leaning into the wind." Prompting the body state produces the facial state for free — and more naturally than naming the emotion.
Matching expression to scene register
A mismatch between expression and context is an uncanny tell even when the face is perfect. Quick mapping: camera-aware scenes (selfies, mirror shots) take direct gaze and performed-but-soft expressions; candid scenes (café, street, kitchen) want attention in the scene, not on the lens; editorial scenes (the fashion register) tolerate neutral intensity that would feel cold in a casual post. When a render feels "off" despite being on-model, audit this axis first.
Consistency: expressions are part of the character
Two pipeline notes. First, the reference library should span expressions — neutral, smiling, candid — because a library of one expression makes every other expression an extrapolation, which is where drift lives. Second, the character should have an expressive range as identity: the voice card extends to the face. A soft, bookish persona whose photos alternate between deadpan and explosive laughter feels miswritten — define her default ("small wry smile, rarely full teeth") and let big expressions be events.
Workflow
Expression is the highest-variance element between renders of the same prompt — generate 3–4 variants and select on expression before anything else; composition flaws crop out, dead eyes don't. At $0.25 a render, the variant habit is the difference between a feed of stock-photo smiles and a character who seems to be having an actual week.
All of it sits on the usual foundation: a locked, verified character whose face you're directing rather than re-rolling.