Character
Character consistency: so it isn't a different face in every shot
The most common failure in AI video: by the second scene, your lead is somebody else. The fix isn't a longer prompt — it is not working from text at all.
You generate your lead and it comes out perfect. You paste the same prompt for the next scene — and a similar but visibly different person looks back at you. Different nose, different hairline, different age.
That is not a bug; it is how the thing works. A diffusion model starts from scratch every time, pulling an image out of noise. Your prompt gives it a direction, not a description of a person.
It is the order of magnitude that matters here: text on its own is a coin toss, while a reference image is very nearly dependable. Which is why the first move is never to polish the prompt, but to build a reference.
The three layers of a character sheet
A good reference is not a portrait but a data sheet. It has three parts, and — this is the bit most often got wrong — all three go on a SINGLE image, not into separate files.
- Full body, four viewsFront, three-quarter, side and back. Include the clothes and the shoes: give it only a face and the model will invent a new outfit for every shot.
- Close-up, five anglesFront, three-quarter, profile, and fifteen degrees from above and below. With a NEUTRAL expression — if he smiles here, the model takes that as his resting face and he will smile everywhere after that.
- Emotions, eight expressionsNeutral, genuine smile, laughter, serious, surprised, confident, sad, determined. Without these the model produces emotion by redrawing the geometry of the face — which means giving you a different face.
The never-rules — this is what takes 90 to 95
The reference image goes out with a handful of short prohibitions. You are not describing what it should be, but what may not change. That works better because „don't alter this" is a clearer instruction to a model than yet another description.
Never change hair length. Never remove the silver watch. Never add a beard. Never change the shoe style.
Pick the three or four details that make the character recognisable and pin those down. Prohibiting everything is wasted effort — what matters is that the viewer sees the same person.
How closely should it stick to it?
Most generators let you set how tightly it binds itself to the reference. Midjourney, for instance, uses a value between 0 and 100, and the three bands serve three different jobs.
| Value | What it keeps | What it suits |
|---|---|---|
| 100 | everything: face, clothes, pose | product shots, repeated setups |
| 50–75 | face and style; pose is free | scenes where the character moves |
| 0 | the face only | new outfit, new location |
For video the middle band is usually right: your character stays recognisable without standing in the same pose in every scene.
The order that works
- Reference firstBuild the sheet and look at it properly. Anything you dislike here you will dislike later too.
- Then the never-rulesThree or four sentences on what may not change.
- Only then the scenesBy now the scene prompt can be short: location, lighting, camera angle. You never have to describe the character again.
Reference first, prompt second. It is the only order that isn't a coin toss.