ChatGPT made image generation conversational: upload a photo, say "now show her at a beach café," iterate in plain English. For one-off edits it's genuinely delightful — which makes its character-consistency limits surprising to people who try to run a persona through it. Those limits are structural, not skill issues. Here's the honest map.
What it's actually good at
Credit first: instruction-following on images is the conversational tools' superpower. "Same person, but profile view, golden hour, looking left" gets understood — the editing intent is parsed correctly far more often than in prompt-syntax tools. For exploring a character concept, mocking up scenes, or producing a handful of images with hands-on art direction, the conversational loop is the fastest tool there is.
Where persistence breaks
Three structural issues for persona work:
- The chat is not a database. Character identity lives in the conversation's context — and context degrades. Twenty images into a session, "her" accumulates small mutations; start a new chat and you're re-uploading references and re-explaining the character. There's no persistent, versioned identity object — the thing a character actually is.
- No verification, conversational confidence. Every result arrives with the same cheerful certainty, but some fraction of "same person" renders are near-miss strangers. Nothing scores outputs against a canonical seed; the quality gate is your tired eyes, image after image. (Why measurement beats eyeballing.)
- Single-anchor coverage. Conditioning rides on whatever image(s) you most recently supplied — usually one. Ask for angles and framings that anchor doesn't contain and the model extrapolates, differently each time. You can manually maintain a folder of references and re-upload the right ones per request… at which point you've become a pipeline, running it by hand in a chat window.
There are also operational frictions for production use: per-image conversational latency, content-policy friction on photorealistic people that varies by request, and no batch workflow — fine for ten images, painful for a monthly content calendar.
The pattern across all the general tools
Same conclusion as Midjourney and the DIY stacks: general image tools optimize for making images; persona operation needs a system that optimizes for maintaining an identity — a locked, verified reference library with coverage, plus per-request reference selection. The test is volume: under ~50 lifetime images, careful manual workflow holds. Posting daily, the missing persistence layer is the product.
Use both, on purpose
The conversational tools earn a permanent slot in the workflow as the concept lab: explore aesthetics, vibe-test character directions, art-direct tricky one-off composites. Then run the persona on a production pipeline — seed → verified 50-image library → scene prompts at $0.25 — where consistency is a property of the system, not of your session hygiene.
Ten minutes, $19, and the character stops living in a chat window.