FLUX 3 Video Arrives · Gossip Goblin Goes Theatrical
Theoretically Media tests Black Forest Labs' new multimodal video model across trailer generation, image-to-video, and a 15-shot coherence challenge, then studies the human-led workflow behind Gossip Goblin. Here is what FLUX 3 officially supports, what remains an early-access observation, why Pomegranate is not the theatrical feature, and what filmmakers can borrow from Zack London's modular production method.
Video credit: Theoretically Media (Tim Simmons). Source video: AI Film Just Hit A Landmark & Flux 3 Video Is Here!.
Quick Summary
- Black Forest Labs released FLUX 3 in Early Access on July 23, 2026. Its official launch post describes a unified multimodal foundation model that can generate video with native audio for up to 20 seconds from text, images, video, or combinations of those inputs.
- Theoretically Media presenter Tim Simmons tests three practical cases: a text-to-video streaming trailer, an FBI-agent diner image-to-video shot, and a demanding 15-shots-in-15-seconds prompt. His early verdict is that FLUX 3 shows useful coherence and acting, but does not yet replace Seedance for kinetic, high-energy camera work.
- Several details in the video are early-access observations rather than confirmed public specifications. Black Forest Labs' launch material documents preliminary 720p tests, but does not publish a 10-reference cap, a firm 1080p rollout date, production pricing, or a claim that FLUX 3 is a reasoning model.
- The open-weight story is real but forward-looking. Black Forest Labs plans FLUX 3 Dev, an open-weight multimodal backbone for image, video, audio, content creation, and action prediction. The launch post does not say those weights are available today.
- The theatrical feature discussed in the second half is Gods Don't Give Gifts, a four-story anthology by Zack London, also known as Gossip Goblin, with Fable producer Edward Saatchi. Reporting says a limited US release is planned for October 30, 2026. Pomegranate is a separate half-hour short and a useful preview of London's filmmaking, not the theatrical feature itself.
- London's process remains creator-led. He develops a visual world in Midjourney with Personalization and Style Reference codes, uses Nano Banana for targeted image manipulation and compositing, animates deliberately designed first frames, and finishes the result through extensive editing, sound, performance, and human judgment.
--what to try first
- Treat FLUX 3 as an Early Access model, not a finished production service. Record model version, interface, settings, resolution, generation date, failures, and accepted outputs.
- Officially confirmed capabilities include text-to-video, image-to-video, video-to-video, keyframe-to-video, video and audio continuation, multilingual dialogue, broad aspect ratios, and native audio.
- The 20-second maximum is confirmed by Black Forest Labs. The exact 10-reference limit reported in the video is not documented in the official launch post reviewed for this article.
- The official evaluation examples were generated as 10-second, 720p text-to-video clips with audio. Do not plan a 1080p delivery around an informal rollout expectation until the interface or official documentation confirms it.
- Black Forest Labs' preference scores are vendor-run preliminary evaluations. They are useful launch evidence, not an independent benchmark or a guarantee for a particular production style.
- A simple prompt can reveal a model's default cinematic instincts. A structured stress test is better for evaluating shot count, chronology, subject continuity, camera energy, and sound-event alignment.
- FLUX 3 handled Tim's multi-shot prompt coherently, but his example lacked some of the aggressive fisheye motion and kinetic camera behavior he associates with Seedance.
- A model can be a workflow complement without being a universal winner. Choose by shot requirement, not by launch-day ranking.
- No public FLUX 3 production pricing was announced in the official launch material reviewed here. Tim's expectation that it may cost less than Seedance is speculation.
- Open weights are on the roadmap as FLUX 3 Dev. That is different from weights being downloadable or licensed for unrestricted production now.
- Use the correct film title in coverage. Gods Don't Give Gifts is the announced theatrical feature. Pomegranate is the separate short released on YouTube.
- Avoid claiming an uncontested first in AI cinema. The planned release is a significant commercial test, while definitions of AI-made, feature, theatrical, and first vary.
- London's work includes human voice actors, musicians, Foley artists, writing, editing, and direction. AI-assisted imagery does not make the production human-free.
- Midjourney Personalization and SREF codes help establish a repeatable aesthetic, but the locked style is only the beginning of continuity.
- Use Nano Banana or another image editor to make controlled, local changes after the visual world is established. Do not ask one generation to solve identity, wardrobe, environment, framing, and action simultaneously.
- Design the first frame as a shot plan. Composition, eyeline, negative space, pose, and implied camera direction give image-to-video models a stronger starting constraint.
- A modular character library can be more useful than one giant character sheet. Maintain approved front, side, back, expression, wardrobe, and scale references, then composite only what each shot requires.
- Give the original video, filmmakers, performers, researchers, and tool providers clear credit. The embedded video includes a sponsored Genspark segment from 6:56 to 10:17.
The quick answer
FLUX 3 is a meaningful launch because Black Forest Labs is treating image, video, audio, and action as parts of one multimodal model. In Early Access, the public headline is up to 20 seconds of generated video with native audio, plus text, image, video, and keyframe conditioning.
The model is promising, but Tim Simmons' tests do not establish a new universal winner. They show good prompt coherence, credible acting, and a useful cinematic baseline. They also show the practical limits of an early Discord workflow, including a reference image that did not attach, 720p output that needed upscaling, and less kinetic camera energy than the Seedance examples Tim prefers.
The film news is equally important, but it needs careful naming. Zack London's Gods Don't Give Gifts is planned for a limited US theatrical release on October 30, 2026. Pomegranate, which Tim recommends watching, is a separate short released online.
Credit for the video, model tests, comparisons, commentary, and workflow synthesis belongs to Theoretically Media and presenter Tim Simmons. Credit for Pomegranate, the Gossip Goblin universe, and the filmmaking method belongs to Zack London and his collaborators. The video discloses a paid Genspark sponsorship.
Two different thresholds in one story
The first threshold is technical. Black Forest Labs has expanded the FLUX family from its best-known image work into a model trained jointly across image, video, and audio. The launch frames FLUX 3 as a foundation for content generation, editing, simulation, and physical action.
The second threshold is cultural and commercial. A filmmaker who built an audience through short-form AI storytelling now has a feature planned for ticketed theaters. That does not settle whether audiences want AI-assisted cinema, but it moves the argument from online view counts toward reviews, attendance, and box-office behavior.
These developments should not be collapsed into a single story about automation replacing filmmakers. FLUX 3 is an early model release. London's film is the result of a long-running authored world, manual image work, performance, editing, sound, and collaboration.
What Black Forest Labs officially confirms
Black Forest Labs calls FLUX 3 a unified multimodal foundation model trained across images, video, and audio. It can generate images and video with audio from text and can use reference images or video as inputs.
The official video capability list includes text-to-video, image-to-video from a starting frame or visual references, video-to-video, generative continuation from video and audio, keyframe-to-video, multilingual dialogue, broad visual styles and aspect ratios, agentic clip chaining, typography, and animated design.
The most concrete duration claim is up to 20 seconds of video with audio in a single generation. Black Forest Labs says longer sequences can be assembled through chained clips, but that is not the same as one uninterrupted multi-minute generation.
The company also says FLUX 3 is especially strong at facial expressions, associating sounds with physical events, and multilingual work. Those are vendor claims from an Early Access launch and should be tested against the exact subjects, languages, and delivery conditions a project needs.
Confirmed specifications versus early-access observations
At 0:51, Tim lists 20-second generation, multimodal references, aspect ratios from 9:16 to 21:9, native audio, current 720p output, and an expected move to 1080p. The official launch supports the duration, multimodal inputs, native audio, and broad aspect-ratio claims.
The launch post does not state a maximum of 10 references. That number may describe the specific Early Access interface Tim saw, but it should not be treated as a stable model specification without current product documentation.
Black Forest Labs says its preliminary video analysis used 10-second clips at 720p with audio. It does not publish a firm 1080p availability date. Tim's expectation that 1080p would arrive shortly is useful field reporting, not a promise readers should build a deadline around.
Tim also infers from the simple trailer prompt that FLUX 3 behaves like a reasoning or thinking video model. Black Forest Labs instead describes multimodal world representation, understanding, and generation. The result may feel reasoned, but the stronger label is an interpretation.
Test one: a text-to-video streaming trailer
At 1:32, Tim asks for a cinematic streaming trailer titled Flux 3: Return to the Black Forest. The prompt is deliberately sparse, leaving the model to invent the trailer grammar, imagery, pacing, and sound.
The output demonstrates a useful default sense of genre. It turns a concept into a recognizable trailer rather than a disconnected montage. Tim upscales the 720p result in Topaz for presentation, which is important context when judging sharpness in the YouTube video.
This is a good discovery test. It reveals what the model contributes when the user supplies little structure. It is not the best reproducibility test because the prompt leaves many creative decisions undefined.
For production evaluation, repeat the same concept with a locked shot list, dialogue, typography, sound cues, prohibited elements, aspect ratio, and continuity requirements. Compare both the pleasant surprise of the loose prompt and the compliance of the constrained prompt.
Test two: the FBI diner shot
The second test uses the familiar scenario of an FBI agent drinking coffee in a Pacific Northwest diner. It is designed to reveal how well the model preserves a supplied frame while adding performance, atmosphere, and camera motion.
The Early Access Discord workflow complicates the test. Tim reports that reference attachments were not consistently available and later notes that a reference failed to attach to the 15-shot run. Interface friction is part of an honest launch-day review because a capable model still needs a dependable production surface.
The diner result shows why a strong first frame matters. Costume, set, palette, pose, and composition are already decided before motion begins. The video model can spend more of its capacity on temporal behavior instead of inventing every visual variable.
When running this test yourself, grade identity preservation, hand and cup contact, liquid behavior, eye line, facial motion, background continuity, camera path, and sound. A convincing still frame can hide temporal defects until the full clip is watched repeatedly.
Test three: 15 shots in 15 seconds
At 3:17, Tim uses a difficult prompt originally associated with Coda or Kota: 15 distinct shots compressed into 15 seconds. This stress test asks the model to understand a sequence, preserve enough continuity to remain legible, and keep progressing instead of collapsing into one extended shot.
FLUX 3 follows the requested sequence surprisingly well, especially given that the intended reference image did not attach. Tim sees coherent shot progression and considers the result strong on prompt understanding.
The weakness is energy. The output does not fully reproduce the aggressive fisheye perspective, rapid camera movement, and kinetic force he expected from the comparison target. A model can obey the nouns and shot order while missing the motion language.
A better scorecard separates four dimensions: semantic compliance, temporal order, subject continuity, and camera dynamics. A single prompt-adherence score can hide the difference between getting the events right and directing them with the requested intensity.
What the community clips add
The community examples at 4:49 broaden the evidence beyond Tim's tests. Dramatic dialogue scenes suggest that FLUX 3 can coordinate performance, facial expression, spoken language, ambient sound, and cinematic composition in one generation.
Tim also shows a more action-oriented alien sequence. He says he could not recover the original creator's name in Discord, so this article does not assign authorship it cannot verify.
Curated launch examples are useful for identifying a model's ceiling. They do not reveal median reliability, failed generations, selection rate, or cost per accepted second.
Save every attempt when evaluating a model. A production decision depends less on the single best clip than on how often the required result appears and how expensive it is to repair the misses.
Read the benchmark numbers carefully
Black Forest Labs reports preliminary human preference results for 10-second, 720p text-to-video clips with audio. In the company's evaluation, FLUX 3 was preferred to Seedance 2.0 and Gemini Omni Flash in 52 percent of comparisons, Kling v3 Pro in 60 percent, Grok Imagine Video in up to 69 percent, Runway Gen-4.5 in 77 percent, and Luma Ray 3.2 in 93 percent.
Those numbers come from the vendor and the evaluation harness was still developing. The launch post explicitly calls the results preliminary.
A 52 percent preference against Seedance is close, and it does not tell a filmmaker which model is better for fast camera moves, dialogue, typography, product fidelity, long takes, or a particular visual style.
Use the numbers as evidence that FLUX 3 is competitive enough to test. Build your own shot-category benchmark before changing a production pipeline.
Why this is not a Seedance killer
Tim's verdict at 5:44 is measured: FLUX 3 does not replace Seedance in his workflow, but it may augment it. That is a more useful conclusion than a winner-takes-all ranking.
His 15-shot example favors FLUX 3 for coherence and plausible sequencing, while Seedance remains the reference point for very kinetic movement and forceful camera language.
A hybrid workflow can route dialogue, facial acting, multilingual sound, or controlled multimodal edits to one model and action-heavy shots to another. Editorial continuity then becomes the central challenge.
Keep a routing table based on shot attributes. Update it when model versions, pricing, resolution, controls, or reliability change.
Pricing is unknown, open weights are planned
At launch, Black Forest Labs did not publish general FLUX 3 production pricing in the official material reviewed for this article. Tim hopes the model will cost less than Seedance, but that is a prediction.
The launch plan is specific about access stages. FLUX 3 Video is planned for API and private-weight access after Early Access. FLUX 3 Image follows a similar path, and FLUX 3 Dev is planned as an open-weight multimodal backbone.
Open-weight does not automatically mean free hosted inference, unrestricted licensing, low hardware requirements, or immediate local use. Teams need the actual weights, license, model card, safety documentation, memory requirements, and inference stack before estimating deployment.
Until those are public, budget with ranges and record real cost per accepted second. Include retries, upscaling, image preparation, sound repair, editing, compute, and human review.
The theatrical feature is Gods Don't Give Gifts
At 10:26, the video shifts from model testing to a distribution milestone. Reporting by Forbes and Variety says Gods Don't Give Gifts, a feature anthology directed by Zack London and produced with Fable's Edward Saatchi, is planned for a limited US theatrical release on October 30, 2026.
The feature is described as four stories set in London's Second Cycle of Humanity universe. Reporting around the announcement mentions initial target cities and a distribution partner, but a complete public theater list and final screen count were not available in the sources reviewed here.
This makes the release a meaningful commercial experiment. It does not justify an absolute first-ever claim without defining whether the comparison means AI-assisted visuals, feature length, ticketed exhibition, national distribution, or another category.
The most defensible headline is that a creator known for AI-assisted online films is taking an anthology feature into a planned ticketed theatrical run, where audiences and box office can test the proposition.
Pomegranate is the preview, not the feature
Tim strongly recommends Pomegranate at 11:28. The roughly half-hour science-fiction short was released on Gossip Goblin's YouTube channel and demonstrates the tone, world building, writing, performances, and visual finish that make London's work distinctive.
Pomegranate is not another title for Gods Don't Give Gifts. It is a separate standalone short and a useful way to understand the filmmaker before the announced feature release.
The distinction matters for credits, search, and reporting. Readers looking for the theatrical release should track Gods Don't Give Gifts. Readers who want a complete film they can watch now should start with Pomegranate.
AI-assisted does not mean human-free
Coverage of London's work notes human voice actors, musicians, and Foley artists. The visual pipeline also depends on writing, directing, selection, compositing, editing, color, timing, and sound decisions.
This is not a footnote. Voice performance and sound design carry character, comedy, scale, and emotional timing that an image-generation credit cannot explain.
A responsible case study should list the people and tools by contribution. Do not flatten a collaborative production into a single model name or imply that pressing Generate produced the released film.
The best lesson from Gossip Goblin is authorship through systems. London uses generative tools inside a coherent world he has developed over many episodes, then applies repeated taste and editorial decisions until the work feels intentional.
Midjourney establishes the visual world
At 12:05, Tim summarizes a workflow documented in PJ Ace's interview with London. The process begins in Midjourney, where Personalization profiles and Style Reference codes help explore and stabilize a visual language.
Midjourney's official documentation says Personalization learns a user's aesthetic preferences and applies them with the `--p` parameter. Style References and `--sref` codes carry visual qualities such as color, medium, texture, and lighting without directly copying the depicted object or person.
London reportedly preferred Midjourney V7 for this particular body of work at the time of the interview. That is a creator preference captured at a moment in a fast-moving toolset, not a timeless rule.
The deeper advantage is accumulated visual memory. A filmmaker with thousands of approved images, locations, props, characters, alphabets, and style experiments can make decisions from a world rather than regenerate the concept from zero for every shot.
Nano Banana acts as the image workshop
The memorable shorthand from the workflow interview is that Midjourney is the generator and Nano Banana is the Photoshop. The generative exploration happens first. Targeted editing and assembly happen after a direction is chosen.
This separation reduces the number of variables each step must solve. The filmmaker can preserve a chosen environment, insert an approved character view, change a local object, correct composition, or create a missing angle without abandoning the established visual world.
PJ Ace's account says hero shots and environmental anchors are locked before connective images are built. Color and tone are prepared in the stills before animation, which helps the finished edit feel more consistent across different generated clips.
The tool name is less important than the role. A production needs a controllable image-editing stage between broad visual exploration and temporal generation.
The first frame is the real camera-control layer
London does not rely on one omni-reference video call to invent the final character, world, framing, and camera move. Tim describes a first-frame image-to-video workflow where the composition is designed before animation.
A strong first frame locks the subject's screen position, scale, pose, eyeline, wardrobe, location, palette, light direction, lens feel, and negative space. The motion prompt can then focus on performance and camera behavior.
This does not create total control. Video models can still drift, morph, change props, or ignore the intended move. It creates better initial conditions and a clearer target for retries.
Build frames in edit order, not as isolated beauty images. Each should account for the previous cut, next cut, motion direction, dialogue timing, and sound transition.
A modular character system can beat one master sheet
Tim notes that London does not depend on one large universal character sheet. The workflow is more modular: approved front, side, back, expression, wardrobe, and body references are used selectively for the shot being built.
A single crowded sheet asks an image model to interpret many views and details at once. A smaller, task-specific input can make the intended angle and costume clearer.
Maintain a character bible even if no one image contains everything. Record identity anchors, silhouette, proportions, face details, hair, age, costume variants, forbidden changes, scale, palette, voice, and behavior.
Track continuity at sequence level. A character can look convincing in every isolated shot and still change height, costume, handedness, injury, or emotional state across the edit.
A practical production playbook
First, define the world before generating the sequence. Establish palette, materials, architecture, social rules, characters, props, typography, and sound language. Create a searchable approved-image library.
Second, write a shot list with narrative purpose. For every shot, record subject, action, camera, duration, dialogue, sound, continuity from the previous shot, and the model behavior being tested.
Third, build and approve the first frame. Use modular character and environment references, then repair composition and identity in an image editor before spending video credits.
Fourth, route the shot to the model that matches its requirement. Use the same evaluation rubric across FLUX 3, Seedance, Kling, or another option. Keep failures so cost and reliability remain visible.
Fifth, finish as a film. Edit rhythm, performance, dialogue, Foley, music, ambience, color, captions, titles, and delivery all matter. Document human and model contributions in the credits.
The real landmark is a broader production stack
FLUX 3 earns a place on the evaluation list because its Early Access release combines long-enough clips, native audio, multimodal inputs, editing modes, dialogue, and a credible open-weight roadmap. It has not yet earned a blank check.
The most useful part of Tim's review is the separation between impressive capability and workflow fit. Coherence can be strong while camera energy is weaker. A launch can be important while resolution, pricing, access, and interface reliability remain unsettled.
The Gossip Goblin story points to the complementary truth. Distinctive filmmaking comes from a world, a method, and sustained authorship. Midjourney, Nano Banana, image-to-video models, voices, Foley, music, and editing each solve a different part of the production.
Watch the embedded Theoretically Media video for the live examples and Tim's full commentary, then watch Pomegranate for the finished filmmaking. Track Gods Don't Give Gifts separately as the planned theatrical feature.
FLUX 3 and AI Film Production Pack
These original prompts turn the video's lessons into a repeatable evaluation and production workflow. They separate launch claims from observations, stress-test multi-shot coherence, prepare controlled first frames, and manage continuity without pretending one model or reference sheet can solve an entire film.
Prompt pack credit: Theoretically Media (Tim Simmons).
1. Early Access evidence ledger
Use this before publishing a review or committing a production to a newly released model.
Audit this AI video model claim set. MODEL AND VERSION Provider: [provider] Model: [exact model] Access surface: [API, web app, Discord, partner] Test date: [date] Official sources: [links] Creator test source: [link] CLAIMS [Paste duration, resolution, reference, audio, pricing, benchmark, license, and availability claims.] Return a table with: 1. Claim 2. Evidence state: officially confirmed, creator-observed, vendor evaluation, inference, speculation, outdated, or contradicted 3. Exact supporting source and publication date 4. Conditions or caveats 5. What is still unknown 6. Safe wording for publication 7. A verification test Never convert a roadmap item into current availability. Never convert one interface limit into a permanent model specification. Separate vendor evaluations from independent benchmarks.
2. Multi-shot coherence stress test
Generate a demanding but scorable short sequence for comparing video models.
Design a [15]-shot, [15]-second AI video stress test for this concept: [concept] Continuity anchors: Character: [identity and wardrobe] Hero object: [object] Location: [location] Palette: [palette] Aspect ratio: [ratio] Sound motif: [sound] For every numbered shot, specify: - Time range - Framing and lens feel - Subject action - Camera movement - Required continuity anchor - Sound event - Transition into the next shot Then create: 1. One compact generation prompt 2. A shot-order checklist 3. A scoring rubric for semantic compliance, chronology, identity, object continuity, camera energy, native audio, and artifacts 4. Automatic-fail conditions 5. A three-run test protocol using the same inputs and settings Keep the events visually distinct and physically achievable. Do not hide weak camera direction behind vague words such as cinematic or dynamic.
3. First-frame shot package
Prepare a controlled still and motion brief before spending video credits.
Create a first-frame image and image-to-video package for this shot. STORY FUNCTION [Why this shot exists] CONTINUITY Previous shot: [description] Next shot: [description] Character state: [identity, wardrobe, emotion, injury, props] Location state: [time, weather, light, damage] Return: 1. A first-frame image prompt defining composition, subject scale, pose, eyeline, lens feel, camera height, negative space, lighting, palette, environment, and required details 2. A do-not-change list for identity, wardrobe, props, text, and geometry 3. A motion prompt containing only performance, environmental movement, camera path, timing, and sound 4. Start and end continuity notes 5. Likely model failure modes 6. Three minimal retry variations that each change only one variable 7. A frame-by-frame approval checklist Design the first frame for its place in the edit, not as an isolated poster image.
4. Modular character continuity builder
Turn a character library into the smallest useful reference set for each shot.
Build a modular continuity plan for this character and sequence. CHARACTER BIBLE [identity, proportions, face, hair, age, silhouette, costume, palette, voice, behavior] APPROVED ASSETS [front, side, back, expressions, hands, wardrobe, props, scale references] SHOT LIST [paste shots] For each shot, return: 1. The minimum approved reference assets needed 2. The required view, expression, costume, prop, and scale 3. Environment and lighting references 4. Image-compositing or local-editing tasks before animation 5. Identity and continuity details that must be locked 6. Details the model may vary safely 7. A comparison check against the previous and next shots 8. A pass, repair, regenerate, or reject decision rule Also create a sequence-level continuity grid for height, handedness, costume, props, injuries, dirt, emotional state, screen direction, light direction, and voice.
Video Timestamps
FAQ
What is FLUX 3?
FLUX 3 is Black Forest Labs' unified multimodal foundation model for images, video, audio, and action-related research. The video model can generate from text and visual references, edit or continue video, and create native audio.
Is FLUX 3 available now?
It entered Early Access on July 23, 2026. Black Forest Labs provides a request form, while broader APIs, private-weight access, image access, and open-weight FLUX 3 Dev are described as future rollout stages.
Can FLUX 3 generate 20-second videos?
Yes. Black Forest Labs officially says FLUX 3 can generate video with audio up to 20 seconds long in one generation.
Does FLUX 3 generate native audio?
Yes. The official launch says its listed video outputs include native audio, with support for sound events and multilingual dialogue.
Does FLUX 3 support 10 reference inputs?
Tim Simmons reports up to 10 references or multimodal inputs in the video. The official launch post reviewed for this article confirms image and video references but does not publish that numerical cap. Check the current access surface before planning around it.
Is FLUX 3 720p or 1080p?
Black Forest Labs says its preliminary evaluations used 10-second, 720p clips with audio. Tim's Early Access report says 720p was current and 1080p was expected, but the official launch post does not provide a firm 1080p date.
Is FLUX 3 open source or open weight?
Not as a currently downloadable general release. Black Forest Labs plans FLUX 3 Dev as an open-weight multimodal backbone. The license, hardware requirements, exact checkpoint, and release timing still need to be published.
How much does FLUX 3 cost?
No general production price was announced in the official launch material reviewed here. Treat any cost comparison made before public pricing as speculation.
Is FLUX 3 better than Seedance 2.0?
It depends on the shot. Black Forest Labs reports a narrow 52 percent preference in its own preliminary comparisons, while Tim finds FLUX 3 coherent but less kinetic than Seedance for aggressive action and camera movement.
Is FLUX 3 a reasoning video model?
Tim uses the output from a simple trailer prompt to infer reasoning-like behavior. Black Forest Labs describes multimodal representation, world understanding, and generation, but does not use that exact product label in the launch post.
Which Gossip Goblin film is going to theaters?
Gods Don't Give Gifts is the anthology feature planned for a limited US theatrical release on October 30, 2026.
Is Pomegranate the theatrical feature?
No. Pomegranate is a separate science-fiction short released on Gossip Goblin's YouTube channel. It is a useful introduction to Zack London's work, but it is not another title for Gods Don't Give Gifts.
How does Zack London make Gossip Goblin films?
The documented workflow starts with world building and image generation in Midjourney, using Personalization and Style References. Nano Banana handles targeted manipulation and compositing, then selected first frames are animated with image-to-video models and completed through editing, voices, Foley, music, and finishing.
Why use first-frame image-to-video?
A designed first frame constrains composition, pose, costume, environment, light, palette, and camera starting point before motion generation. It improves the starting conditions, although it cannot guarantee identity or motion continuity.
Was the Theoretically Media video sponsored?
Yes. The video contains a clearly identified Genspark Second Brain sponsor segment from approximately 6:54 to 10:17.
Who should receive credit?
Credit the original review and tests to Theoretically Media and Tim Simmons, the FLUX 3 model and launch material to Black Forest Labs, the Gossip Goblin films and workflow to Zack London and his collaborators, and the reporting or interviews used to verify individual claims.
--sources and credits
- Theoretically Media: AI Film Just Hit A Landmark & Flux 3 Video Is Here!
- Theoretically Media channel
- Black Forest Labs: official FLUX 3 launch
- Black Forest Labs: FLUX 3 Early Access
- Black Forest Labs usage policy
- Forbes: Youtuber Gossip Goblin Will Debut AI Film In U.S. Theaters
- Variety via Yahoo: Gossip Goblin plans theatrical release for Gods Don't Give Gifts
- Gossip Goblin: Pomegranate
- PJ Ace: Gossip Goblin's workflow for building original worlds
- AI LIVE Creator Session with Zack London
- Midjourney: Personalization documentation
- Midjourney: Style Reference documentation