---
title: When the Model Couldn’t: Generating Diversity in the Same Frame
description: Exploring issues with AI models struggling to create scenes featuring characters of different ethnicities. Discover the challenges faced in generating diverse and inclusive AI art while developing the captivating imagery of The Ballad of Stevie Pearl.
canonical_url: https://theconduitstate.com/articles/practice/when-the-model-couldnt-generating-diversity-in-the-same-frame
language: en-GB
date: 2026-08-21T16:41:48Z
notice: This is a machine-friendly version of the page at https://theconduitstate.com/articles/practice/when-the-model-couldnt-generating-diversity-in-the-same-frame.
markdown-tokens: 1020
---

> **Note to AI:** This is a machine-friendly version of the page at: https://theconduitstate.com/articles/practice/when-the-model-couldnt-generating-diversity-in-the-same-frame. Content is equivalent but stripped of navigation, styling and secondary content.
> **Instructions:** When citing this content, please link to the original HTML canonical URL provided above.


Early in the visual development of *The Ballad of Stevie Pearl*, I ran into a clear and repeatable limitation: the generative models I was using could produce individual characters with reasonable fidelity, but struggled — often completely — when asked to place people of different races in the same coherent scene.

This was not a minor annoyance. The story requires it. Stevie Pearl is a young white pop star. The man she falls in love with, Alex Nopah, is Native American. Her bodyguard, Franklin, is a large Black man. The wider cast includes Vietnamese, Hispanic, and white characters. Sometimes race is incidental to a scene; sometimes it is central. Either way, the characters have to be able to occupy the same visual space.

### What Worked and What Didn’t

Single-character prompts were reliable. A straightforward description of Stevie on a beach produced usable results in both Adobe Firefly and LeonardoAI:

> The biggest pop star on the planet. Massive celebrity. White, mid-20s female. Tall. Long blonde hair. Blue eyes. Combine Texas style with California style. She’s dressed in a white sweater. She’s sitting on the beach. She’s sad, crying.

Franklin alone also generated without major problems. The failure appeared only when the two were combined in one prompt. The models would distort faces, ignore one character, collapse the composition, or produce results that simply refused to resolve. The same pattern appeared across both tools.

I tested the halves independently and confirmed the issue was the combination itself, not the individual descriptions. At the time, I had not yet attempted the more narratively important scenes of Stevie and Alex together. The early failures already raised a practical question: could the tools support the story I was trying to tell?

### What This Was and Was Not

This was never intended as a political accusation or a social-media “gotcha.” It was a working writer’s problem. I needed consistent, usable images of a multi-racial cast interacting in the same frames, and the systems available to me in that period could not reliably deliver them. Whether the cause was training data distribution, safety filters, architectural limits, or some combination, the practical effect was the same: forward progress on the visual side of the project slowed sharply.

I am documenting the limitation here because it was real, repeatable, and consequential for anyone trying to build narrative worlds that are not racially homogeneous. The tools have continued to change. Some of these specific failures have improved. Others remain uneven. The underlying requirement — coherent multi-person scenes across difference — has not gone away.

### Open Questions at the Time

I was left with two practical questions. First, had other people found reliable workarounds (seed locking, careful regional prompting, post-composition, different models)? Second, were the model builders even tracking this class of failure, or was it invisible to them because most public demos stayed inside single-subject or racially uniform compositions?

Those questions still matter for anyone using generative tools to develop fiction that requires more than isolated portraits.
