All guides

AI Video line · stop 06 of 14 · 28 min · members

Character consistency across shots, when the model has no memory between them

Keeping the same person across a sequence when the model has no memory between generations.

Free with an account

Sign in to read.

Membership is free: an account opens all 86 script pages. The Lab, Studio Canvas and the paid guides need the $99 pass, paid once. Already signed in on this browser? The page opens by itself.

01

The problem

Every generation starts from nothing.

The model does not know it made your character ten minutes ago.

Generate the same described person five times and you get five siblings. Close enough that you notice the resemblance, different enough that cutting between them looks like a continuity error — because it is one.

This is not a limitation you prompt your way out of. No description is specific enough to reconstruct a face. Add more detail and you narrow the range slightly while making the prompt brittle; the fundamental problem is that identity is not something language encodes precisely.

Consistency therefore comes from carrying something visual between generations, and the whole craft is in choosing what to carry and how much of the frame it should control.

02

The three jobs

A reference image does one of three things, and you must decide which.

Identity, style and composition are separate problems that people try to solve with one picture.

When you attach a reference, you are implicitly asking for one of these:

  • Identity — this specific face and body, in a new situation.
  • Style — this treatment, light and palette, applied to different content.
  • Composition — this framing and arrangement, with different subjects.

Using one image for all three is the most common mistake and it produces the characteristic failure: a shot that is nearly right in every dimension and correct in none. The face drifts because the model was also trying to match the lighting; the composition changes because it was also trying to match the face.

Where the tool allows separate reference slots, use them separately. Where it does not, decide which single job matters most for this shot and let the other two be handled by the prompt — which is much better at style and composition than it is at identity.

03

Method one

A locked first frame, carried forward.

The cheapest reliable approach: generate a person once, then never generate them again.

Produce one image of the character that you are genuinely happy with. Frontal or three-quarter, evenly lit, neutral expression, nothing dramatic. This is your master.

Every subsequent shot is an edit or an image-to-video starting from that master, not a fresh generation. You are changing the situation around a fixed person rather than re-describing the person each time.

The limitation is real: you are constrained to angles and expressions reachable from the master. A profile shot from a frontal master will invent the far side of the face. Plan the sequence around what the master supports, or produce two or three masters at different angles and treat each as the source for the shots it covers.

What to fix in the master

Hair, because it changes silhouette and is the first thing to drift. Clothing, for the same reason. And one distinctive, describable feature — a particular jacket, a specific hair length — that you can also reinforce in every prompt, so language and image are pulling in the same direction rather than competing.

04

Method two

A trained identity, when the sequence is long enough to justify it.

Training buys you angles and expressions that a single reference cannot reach.

If you need a character across dozens of shots, over weeks, in situations a single master cannot cover, training an identity is the answer. You supply a set of images and get back something you can call up repeatedly.

The dataset decides everything, and the usual instinct — supply the best photographs — is wrong. What you want is variety:

  • Multiple angles, including profile and three-quarter from both sides.
  • Multiple lighting conditions, including unflattering ones.
  • Neutral expressions predominantly, with a few others.
  • Consistent identity, varying everything else.

Twenty varied images beat sixty similar ones. If every training image is the same flattering three-quarter under the same soft key, you have trained a pose rather than a person, and it will refuse to turn its head.

Training also has a cost the single-master method does not: time before you can start, and a commitment to a look you may want to change. Use it when the sequence is long. For a six-shot job it is rarely worth it.

05

The reinforcement

Say the same words about them every single time.

Prompt drift is a bigger cause of character drift than most people realise.

Whatever method you use, write the character's description once and paste it identically into every prompt. Not a paraphrase. The same words.

People rewrite the description each shot without noticing — 'dark hair' becomes 'black hair' becomes 'short dark hair' — and each variation moves the result slightly. Across ten shots those small movements compound into a visibly different person.

# character block, pasted verbatim into every prompt
A woman in her early thirties, shoulder-length dark brown hair
worn loose, matte black jacket, no jewellery, neutral expression.

Keep this block in a file next to the master image. Together they are the character, and treating them as a single unit is what makes a sequence hold together.

06

Accepting the limit

Design the sequence around what holds.

The best consistency work is invisible because the shots were chosen to be achievable.

The most effective technique is not technical. It is choosing shots that do not stress the weakness.

Wide shots hold better than close-ups, because there is less face to get wrong. Shots where the subject is turned away, backlit, or partially obscured hold almost perfectly. A sequence that opens wide, moves to a three-quarter medium, and cuts away before the close-up is both a normal edit pattern and a technically robust one.

This is the same discipline as shooting with a real constraint — a location you cannot light, an actor available for one day. You design the sequence around it, and the audience never knows the constraint existed.

Fighting the limitation head-on with ten close-ups of the same face is where the work goes wrong. Designing around it is where it starts looking professional.