Working Together
14 February 2024
21 August 2026
Sean William Hammond
What I Learned Trying to Make an AI Music Video for The Ballad of Stevie Pearl
What I Learned Trying to Make an AI Music Video for The Ballad of Stevie Pearl
It’s okay to fail in public when the failure teaches you something real. This is a recap of an early attempt to generate a music video for The Ballad of Stevie Pearl — the parts that worked, the parts that didn’t, the practical lessons, and why I still believe the tools will eventually catch up to the ambition.
The Part That Still Feels Amazing
Being able to generate moving images of a character and world that previously existed only in text is genuinely transformative for a storyteller. I can now take characters, settings, and emotional beats far beyond the page and give readers a richer sensory experience. That possibility still excites me. The fact that any of this is possible at all is worth stating first.
The Real Constraints I Hit
I had only been working seriously with generative AI since early December when I attempted this. I am not a developer. I am a content creator, writer, and website builder who has spent twenty years picking up design, photography, video, and editing skills as needed. I brought that background to the work, but the tools and the workflow were still new.
The learning curve is steep and mostly undocumented in any reliable way. Most tutorials amount to “try prompts, adjust settings, see what happens.” Over time the craft becomes the ability to articulate what you want with increasing precision. Until that fluency develops, the process is pure trial and error.
The number of models and interfaces compounds the problem. Each has strengths. Jumping between them means starting over on interface, prompting style, and token economy. I moved from Adobe Firefly to LeonardoAI largely because Firefly could not give me a consistent Stevie Pearl, and because Leonardo had just launched motion features. I burned through a large token allocation quickly.
Character consistency remained the hardest technical problem. I could get close — occasionally very close — but I could not reliably move the same character through different scenes, actions, and emotional registers without visible drift. Training custom datasets produced results that looked markedly worse than the base models. For a project that needs dozens of coherent images and clips, that limitation was decisive.
Lessons That Still Apply
Style is not optional. Every story world has a visual signature. The Ballad of Stevie Pearl lives in saturated pinks, purples, and high-chroma color — what I have called “unicorn puke.” Early in the process I became so focused on getting Stevie’s face right that I abandoned the broader visual language of the project. The finished video does not match the book covers, the promotional materials, or the rest of the intended aesthetic. That was an artistic failure on my part, not a hard limit of the tools.
Voice and register matter as much as likeness. The framing device of the book is a polished, documentary-style interview. The middle of the story is Stevie’s subjective, heightened, sometimes fantastical memory. Those two visual registers should feel different. I did not protect that distinction clearly enough in the video.
Motion generation is still immature for anything that needs to sit beside professional work. To people who already care about AI, the clips can look exciting. To everyone else they still read as low-resolution, slightly uncanny, and unfinished. Publishing them alongside a finished novel would have damaged the brand more than it helped. The same standards I apply to typography, layout, and copyediting apply here: if it undermines the reader’s trust in the world, it is not ready.
I also had almost no directorial control over performance or camera movement. Each clip was a roll of the dice. Acceptable shots often required multiple generations, and the native resolution was low enough that upscaling became mandatory. That workflow is slow and expensive for the quality it currently returns.
Why the Attempt Still Mattered
This version of “Back in ’85” will not appear in The Experience or in any official Stevie Pearl materials. It was still worth making. I learned the practical limits of the tools at that moment, recovered a clearer sense of the visual language the project actually needs, and confirmed that the gap is temporary rather than fundamental. The capability is already impressive. It is not yet invisible.
The novel itself is out. Digital and physical editions are available. The Experience — a more visual, expanded version of the story — is in early development and will continue to use generative tools as they improve.
If you have been trying to generate coherent short-form video or character-consistent motion, I would genuinely like to hear what has worked and what has not. The field moves quickly; shared practical knowledge is still scarce.
The Images





















