
Working Together
06 February 2024
21 August 2026
Sean William Hammond
When the Model Couldn’t: Generating Diversity in the Same Frame
When the Model Couldn’t: Generating Diversity in the Same Frame
Early in the visual development of The Ballad of Stevie Pearl, I ran into a clear and repeatable limitation: the generative models I was using could produce individual characters with reasonable fidelity, but struggled — often completely — when asked to place people of different races in the same coherent scene.
This was not a minor annoyance. The story requires it. Stevie Pearl is a young white pop star. The man she falls in love with, Alex Nopah, is Native American. Her bodyguard, Franklin, is a large Black man. The wider cast includes Vietnamese, Hispanic, and white characters. Sometimes race is incidental to a scene; sometimes it is central. Either way, the characters have to be able to occupy the same visual space.
What Worked and What Didn’t
Single-character prompts were reliable. A straightforward description of Stevie on a beach produced usable results in both Adobe Firefly and LeonardoAI:
The biggest pop star on the planet. Massive celebrity. White, mid-20s female. Tall. Long blonde hair. Blue eyes. Combine Texas style with California style. She’s dressed in a white sweater. She’s sitting on the beach. She’s sad, crying.
Franklin alone also generated without major problems. The failure appeared only when the two were combined in one prompt. The models would distort faces, ignore one character, collapse the composition, or produce results that simply refused to resolve. The same pattern appeared across both tools.
I tested the halves independently and confirmed the issue was the combination itself, not the individual descriptions. At the time, I had not yet attempted the more narratively important scenes of Stevie and Alex together. The early failures already raised a practical question: could the tools support the story I was trying to tell?
What This Was and Was Not
This was never intended as a political accusation or a social-media “gotcha.” It was a working writer’s problem. I needed consistent, usable images of a multi-racial cast interacting in the same frames, and the systems available to me in that period could not reliably deliver them. Whether the cause was training data distribution, safety filters, architectural limits, or some combination, the practical effect was the same: forward progress on the visual side of the project slowed sharply.
I am documenting the limitation here because it was real, repeatable, and consequential for anyone trying to build narrative worlds that are not racially homogeneous. The tools have continued to change. Some of these specific failures have improved. Others remain uneven. The underlying requirement — coherent multi-person scenes across difference — has not gone away.
Open Questions at the Time
I was left with two practical questions. First, had other people found reliable workarounds (seed locking, careful regional prompting, post-composition, different models)? Second, were the model builders even tracking this class of failure, or was it invisible to them because most public demos stayed inside single-subject or racially uniform compositions?
Those questions still matter for anyone using generative tools to develop fiction that requires more than isolated portraits.
The Images





















