Generating twenty good scene images, then deciding what the lead's face looks like.
Locking one face reference per character, then generating scenes.
Image models remember nothing between requests. Deciding the face later means those twenty images will not match a single later episode.
