Short-form video dominates reach on every platform — and it's also where AI personas historically broke, because video generation from text re-invents the character frame by frame. Image-to-video sidesteps the problem elegantly: instead of describing your character to a video model, you hand it a finished photo and describe only the motion. The photo becomes the first frame; identity is pinned by construction.
Why this preserves the face
Text-to-video inherits all the problems of text-to-image identity, multiplied by time. Image-to-video inverts the contract: the model's job is to continue the supplied frame plausibly, not to imagine a person. Your character's face, outfit, and scene are already decided in pixels — the model animates hair, fabric, light, and camera. Used this way, video becomes a post-processing step on your proven photo pipeline rather than a separate identity problem.
The workflow that follows from this: photos first, always. Generate stills, post them, see what the audience rewards — then animate the winners. Never burn video generation on unvalidated concepts; stills are your cheap test layer, video is the amplifier. (On AI CMO, any reference or generated image has an Animate button — costs are shown before you generate.)
Writing motion prompts
The skill transfers from scene-prompting but the rules differ — you're describing change over a few seconds, nothing else:
- Describe only motion. The image already says who, where, and what she's wearing. "She turns her head slightly and smiles, hair drifts in a soft breeze" — not a re-description of the scene. Re-describing invites the model to change things you wanted kept.
- Subtle beats ambitious. Micro-motion — breathing, hair, a head turn, steam rising, fabric sway — reads natural and loops beautifully. Complex actions (standing up, walking through a door, hand interactions) are where current models produce mush and morphs.
- One camera move, named simply. "Slow push in," "gentle handheld sway," "static camera." Camera language does enormous work for the perceived production value.
- Short clips, intended loops. A 3–8 second subtle loop is a premium-feeling Reel; a 15-second action sequence is a risk. Cut your losses in the prompt, not the edit.
What's worth animating
A simple portfolio rule for persona accounts:
- Best-performing stills — your proven content, amplified into the video-favored feed.
- Atmosphere shots — golden hour, rain on windows, café steam: scenes whose whole appeal is gentle motion.
- The character moment — one signature clip (the smile, the hair flip) reused as an intro/outro across content. Repetition builds the character.
Skip talking-head clips for now: lip-synced synthetic speech triggers both uncanny-valley reactions and the strictest platform scrutiny — ambient motion is the high-reward, low-risk lane. And label AI video everywhere; video labeling rules are stricter than image rules on every platform.
Slot it into the calendar
In a 30-day plan: stills carry the daily cadence; one or two animated clips a week ride the video-distribution bonus. That ratio matches both the economics (videos cost meaningfully more than stills — generate them from winners, not guesses) and the audience's appetite: a persona feed that's all motion reads like an ad channel; mostly-stills-plus-occasional-life reads like a person.
Photos prove it, video scales it. Start with the photo engine — first character, $19 — and animate your first winner this week.