Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
AI filmmaking workflowMidjourney srefNano Banana

Inside Gossip Goblin's AI Filmmaking Workflow: Midjourney to Video

How Gossip Goblin builds AI films with Midjourney, Nano Banana, and first-frame image-to-video, skipping omni-reference for tighter camera control.

MindStudio Team RSS
Inside Gossip Goblin's AI Filmmaking Workflow: Midjourney to Video

What is Gossip Goblin’s AI filmmaking workflow?

Gossip Goblin, the AI filmmaker behind creator Zack London, builds scenes by generating characters and environments in Midjourney first, refining them in Nano Banana, then animating with first-frame image-to-video generation instead of relying on omni-reference or multi-reference video tools. The pipeline favors manual control at every stage over the newer, more automated all-in-one generation methods that most AI video tools now push creators toward.

TL;DR

  • Midjourney comes first, not Nano Banana or a video model. All of Gossip Goblin’s shots start as still images generated in Midjourney, using personalization and style-reference (sref) codes to keep a consistent visual identity across scenes.
  • Nano Banana works as the Photoshop step, handling image editing and manipulation after Midjourney generates the base image, rather than being the tool that creates images from scratch.
  • First-frame image-to-video is the animation method, meaning a single finished still image becomes the starting frame for a video generation, giving the filmmaker direct control over camera movement instead of letting an omni-reference model guess it.
  • There’s no single master character sheet. Characters are built modularly and placed into scenes piece by piece rather than locked into one omni-reference asset that a model reuses automatically.
  • The theatrical release of Gossip Goblin’s feature film marks a notable moment for AI filmmaking, with Zack London drawing comparisons to early indie auteurs for building a cult following on YouTube before moving to theaters.
  • The half hour short film “Pomegranate” serves as a public demonstration of this workflow, and it’s built to work as a standalone watch, not a companion piece requiring prior context.

Remy doesn't write the code. It manages the agents who do.

R
Remy
Product Manager Agent
Leading
Design
Engineer
QA
Deploy

Remy runs the project. The specialists do the work. You work with the PM, not the implementers.

How does the Midjourney-first approach work?

Every shot in the Gossip Goblin pipeline begins as a still image in Midjourney, not in a video model and not in Nano Banana. At the time this workflow was documented, Zack London was working primarily in Midjourney version 7 rather than the newer 8.1, leaning on personalization settings and style-reference (sref) codes to keep a recognizable look across different characters and environments.

This matters because a lot of AI video pipelines skip the “pure image generation” step entirely, feeding text prompts straight into a video model and hoping for consistency. Starting in Midjourney gives the filmmaker a stable, art-directed base image to build from. The sref codes act as a style anchor, so different generations across a project still feel like they belong to the same visual world, even when the subject matter or camera angle changes.

Why use Nano Banana as an editing tool instead of a generator?

In this workflow, Nano Banana isn’t where images are born. It’s where they get fixed. Once Midjourney produces a base image, Nano Banana steps in to manipulate it: adjusting details, compositing elements, or cleaning up whatever Midjourney didn’t nail on the first pass.

That division of labor is worth noting because plenty of creators use Nano Banana as a primary generator. Treating it instead as a Photoshop-style refinement layer keeps the creative decisions concentrated in Midjourney, where the sref and personalization settings live, while Nano Banana handles the technical cleanup that turns a good AI image into a shot-ready one.

Why first-frame image-to-video instead of omni-reference?

This is the most distinctive part of the pipeline. Modern video models increasingly support omni-reference or multi-reference generation, where a model takes several reference images (a character, a location, a style) and generates a full video clip while trying to keep everything consistent on its own. It’s convenient, but it hands a lot of creative control to the model.

Gossip Goblin’s workflow avoids that entirely. Instead, every clip starts from a single finished image, the first frame, and the video model animates outward from that frame. The reasoning is straightforward: it preserves camera control. When a model is juggling multiple references and deciding on camera movement itself, the results can drift from what a filmmaker actually wants. Locking in the first frame as a fully art-directed still means the director has already made the composition and framing decisions before any motion is generated.

There’s a broader argument buried in this choice too: that leaning on omni-reference and multi-reference tools can make creators lazy, letting the model make decisions that used to belong to the person behind the camera. Whether or not that’s true across the board, it’s a deliberate trade-off in this pipeline. Slower, more manual, but more controlled.

Why no master character sheet?

Plans first. Then code.

PROJECTYOUR APP
SCREENS12
DB TABLES6
BUILT BYREMY
1280 px · TYP.
yourapp.msagent.ai
A · UI · FRONT END

Remy writes the spec, manages the build, and ships the app.

Most AI production pipelines that use multi-reference or omni-reference models rely on a single character sheet, one locked set of reference images that gets fed into the model every time that character appears. It’s efficient, but it also means the character is generated the same standardized way across every scene.

Gossip Goblin’s approach skips that. Characters are built modularly and placed into scenes individually rather than pulled from one master reference asset. That likely means a given character can be reconstructed slightly differently depending on the needs of a specific shot (lighting, angle, wardrobe) rather than forcing every appearance through the same reference pipeline. It costs more manual effort per shot, but it buys flexibility that a rigid character sheet doesn’t allow.

Is this workflow worth replicating?

For creators chasing speed, probably not. This is a slower, more hands-on pipeline than what most one-click AI video tools promise. Generating a base image in Midjourney, refining it in Nano Banana, and then animating from a single first frame takes more steps than feeding a prompt into an all-in-one video generator.

But for creators chasing consistency and control, particularly over camera work and character presentation, it’s a useful template. The core lesson isn’t “use these specific tools.” It’s the underlying principle: separate the jobs. Let one tool handle art direction and style (Midjourney), another handle refinement (Nano Banana), and treat video generation as animation from a fixed, already-approved image rather than a black box that decides everything at once.

The proof of concept is the output. Gossip Goblin’s short film “Pomegranate,” released as a standalone half-hour piece, doesn’t require any prior context to watch, and it demonstrates the workflow’s ability to produce genuinely coherent long-form scenes. The film’s theatrical release, following a feature announcement that drew comparisons in the Hollywood Reporter to early independent filmmakers, suggests the workflow scales beyond short-form YouTube content into something closer to conventional film production.

Frequently Asked Questions

What tools does Gossip Goblin use to make AI films?

The workflow centers on Midjourney for initial image generation (using sref and personalization codes), Nano Banana for image editing and refinement, and a video generation model for animating finished images using the first-frame image-to-video method.

What is first-frame image-to-video generation?

It’s a technique where a single, fully finished still image serves as the starting frame of a video clip, and the video model animates motion outward from that exact frame. This differs from text-to-video or multi-reference generation, where the model has more freedom to interpret framing and camera movement on its own.

Why avoid omni-reference or multi-reference video models?

Omni-reference and multi-reference models let a video generator combine several inputs (character, setting, style) and produce a clip automatically, but that convenience comes at the cost of camera control. Starting from a single locked first frame keeps composition and camera decisions in the filmmaker’s hands rather than the model’s.

Does Gossip Goblin use one character sheet for consistency?

No. Rather than relying on a single master character reference sheet reused across every scene, characters are built modularly and placed into each scene individually, trading some efficiency for more per-shot flexibility.

Why does this workflow matter for AI filmmaking generally?

Cursor
ChatGPT
Figma
Linear
GitHub
Vercel
Supabase
goremy.ai

Seven tools to build an app. Or just Remy.

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

It shows that a deliberately manual, multi-tool pipeline (still image generation, then editing, then first-frame animation) can produce feature-length, theatrically released work. It’s a counterpoint to the trend of all-in-one video generation tools that automate more of the creative decision-making.

Presented by MindStudio

No spam. Unsubscribe anytime.