← All sessionsHomeSearch
C7 EST | 14 Day AI Sprint·Day 2 | Image & Video Manipulation & clone Generation using AI·4:43:00

Day 2: The Image/Video Model Landscape, Cinematic Prompting, Cost Strategy, and a Commercial Built Live

Jordan Billinkoff Day mentor - 20+ years across media, marketing, product design and creative production; teaches image/video generation as 'the new cameras' and builds a full commercial live · Uthappa Host - logistics, Q&A triage, announces the bonus office hours that start Day 3

The short version

  1. The mental model: generative image and video models are 'the new cameras.' Photographers do not own every brand - pick two or three favorites and build a toolkit. 'You don't need to use every single tool under the sun.'
  2. Timeline: 2022 ChatGPT (text only), 2023 Midjourney + early Runway, 2024 video 'very experimental', 2025 production-ready - and late 2025 (Nano Banana / Nano Banana Pro) added natural-language editing, multi-image reference, camera-angle changes, start/end-frame video, and synchronized dialogue. Character, scene, and product CONSISTENCY is the innovation that makes multi-shot storytelling possible.
  3. Prompting has three grammars: images are comma-separated elements with the most important first (subject, action, scene, atmosphere, then camera/lens/lighting/film stock); videos are described as scenes in sentences; edits are imperative instructions ('replace the man in the first image with the man in the second').
  4. Money: realistic casual spend is '$20 to $50 per month' TOTAL via an aggregator, not per tool. Free trials are for learning, not delivering (watermarks, queues). Open source (Flux, ComfyUI, LoRA fine-tuning) vs closed (Midjourney - no API, must subscribe directly).
  5. The demo is the lesson: pre-production (Figma storyboard) -> production (generate a base image, insert yourself by multi-image reference at 1K, iterate with natural-language edits, upscale to 4K only when happy, chain each output as the next shot's reference) -> video (Kling start/end frames; Flow for the dialogue shot) -> post (Premiere, cut to a pre-made track). 'Tons and tons of iterations... until they get exactly what they want.'
  6. Legal weather: OpenAI tightened celebrity guardrails after backlash; Disney invested in OpenAI while suing Midjourney (with Warner and Universal); his call is that 2026 is when 'the dust is gonna start settling' via lawsuits-turned-licensing, as happened in music.

The concepts

01

'The new cameras': the 2022-2025 timeline and how to read the leaderboard

A wall of forty logos is meant to make you anxious. Treat them like camera brands and the anxiety goes away.

Year by year: 2022 was the ChatGPT moment and the 'AI copywriter' era; 2023 Midjourney (born in Discord) made images mainstream while Runway was an experimental video model; 2024 video was still 'very experimental'; 2025 tools became 'ready for prime time' for commercial use and workflows emerged around them. 'It's still very early.' The landscape at recording - images: Midjourney, Google Nano Banana / Pro, ByteDance Seedream, OpenAI ChatGPT Image 1.5, Black Forest Labs Flux 2 (open), Ideogram (text rendering), Recraft, Qwen; video: Runway, Kling, Google Veo 3.1, Grok, Sora 2 (social, auto multi-shot, less controllable), Moon Valley (licensed data only), LTX Studio, WAN 2.2 Animate (full-body character replacement), Hailuo.

For a neutral ranking he uses Artificial Analysis: two anonymous images, people vote left or right, no brand bias - the day before, ChatGPT Image 1.5 had edged Nano Banana Pro. His own kit: Seedream, Nano Banana Pro, Krea's model, Flux for stills; Seedance, Kling ('love Cling'), Runway, Hailuo for motion.

Why it matters

Model rankings turn over monthly; the durable skill is knowing the categories (open vs closed, image vs video, aggregator vs native) and one neutral place to check.

Go deeper

In one line: Generative image/video models = interchangeable 'cameras' with different aesthetics; pick a small personal kit and re-check Artificial Analysis rather than chasing every release.

2022 text, 2023 Midjourney/Runway, 2024 experimental video, 2025 production-ready with workflows (l3186012 0:31-0:35)

'These generative AI image and video models are the new cameras' (l3186012 0:39)

Artificial Analysis = anonymous left/right crowd votes; ChatGPT Image 1.5 led Nano Banana Pro that week (l3186012 0:40-0:41)

Kling called one of the strongest yet under-recognized video models (l3186012 0:38)

His kit: Seedream, Nano Banana Pro, Krea, Flux (images); Seedance, Kling, Runway, Hailuo (video) (l3186012 0:56-0:57)

▶ Watch this taught:

02

What Nano Banana changed: instructional editing, multi-image reference, consistency

Before mid-2025 you re-rolled the whole image and prayed. Now you say 'remove the mug from his hand' and it does.

The capability cluster that arrived late 2025, credited to Google's Nano Banana and Nano Banana Pro ('about a month or two ago'): natural-language image editing, single- and multi-image references, changing camera angle on an existing image, start/end-frame video generation, and video with synchronized spoken dialogue. Multi-image reference - 'replace the man in the first image with the man in the second image' - is what lets you put a real person or a real product into a generated scene.

This is what makes 'character consistency, scene consistency, product consistency' achievable, which he calls 'a massive innovation in the last 6 months to the past year.' The working technique: never prompt the next shot from scratch; feed the previous generation in as the reference so wardrobe, face, and set carry through the sequence.

Why it matters

Consistency is the difference between a gallery of pretty stills and a story you can cut into a film or an ad.

Go deeper

In one line: Instructional edits + multi-image reference + reference-chaining shot to shot = character/scene/product consistency; the late-2025 unlock for multi-shot AI storytelling.

Nano Banana brought natural-language editing, image references, camera-angle changes (l3186012 0:35-0:36)

Start/end-frame video and synchronized dialogue are late-2025 capabilities (l3186012 0:36)

Multi-image reference (insert a real person into a generated scene) new 'since about June or July' (l3186012 1:45)

Chain each output as the next shot's reference to hold consistency (l3186012 2:17)

Character, scene, and product consistency named as the enabling capabilities (l3186012 1:29-1:30)

▶ Watch this taught:

03

Three prompt grammars: image (comma list), video (scene sentence), edit (instruction)

The same words in a different order make a different picture - whatever is at the front gets the most weight.

Image prompts: Subject -> Action -> Scene/Environment -> Atmosphere -> then camera language (camera body, lens, angle, lighting, film stock, style). 'You're not writing full sentences... separate every single element by commas,' and front-load what matters because 'whatever you have near the front is going to have the most emphasis.' Leave lighting or style out and 'it will just guess for you.'

Cheat sheet: Dutch angle (tilted, unease), low angle (power), close-up / medium / cowboy (thigh-up) / extreme close-up; lighting - natural, volumetric, golden hour, rim, bioluminescent; lens - higher mm is tighter (50mm portrait vs 10-20mm wide); film stock for a retro look; real camera names work (Sony A7S III, ARRI Alexa, RED, Canon 1DX). Video prompts are different - 'just describe the scene that you wanna see' in sentences. Edit prompts are imperative: 'show me, remove, replace, place him a meter back.'

Why it matters

Most bad generations are grammar errors - a video sentence fed to an image model, or a descriptive re-prompt when an instruction was needed.

Go deeper

In one line: Image = comma-separated elements, most important first, ending in camera terms; video = descriptive scene sentences; edit = direct imperative instructions.

Order: subject, action, scene, atmosphere, then camera/lens/lighting/film stock (l3186012 1:00-1:01)

Front of the prompt carries the most emphasis (l3186012 1:00-1:01)

Images: comma-separated elements, not sentences (l3186012 1:43)

Video: describe the scene in sentences; edits: imperative instructions (l3186012 2:38)

Camera vocabulary: Dutch/low angle, close-up/medium/cowboy, golden hour, rim light, 50mm vs 10-20mm (l3186012 1:04-1:06)

▶ Watch this taught:

Check yourself

Answer from memory first — the recall attempt is what makes it stick. Then reveal.

Your kitchen shot came back with two extra coffee mugs. Re-prompt or edit?

Edit: 'remove the two coffee mugs from the countertop' against the existing image. Re-prompting regenerates the whole scene and breaks consistency.

04

$20-50 a month, not per tool: aggregators, trials, open vs closed

Ten subscriptions is the beginner's mistake. One aggregator with credits is how working creatives pay.

Realistic casual spend is '20 to 50 dollars per month' in total. Two aggregator models: monthly subscription with a nice UI and expiring credits (Freepik, Krea), or pay-per-use with a rawer interface (fal.ai, Replicate) - cheaper overall. Free trials are for 'learning and exploring... not good for producing a real piece of content': limited credits, watermarks on video, queues - though you can stack trials (Hailuo, Kling, Higgsfield) and scale a video to 110-120% to crop a corner watermark.

Open vs closed via an Android/Mac analogy: open models (Flux, ComfyUI locally) can be fine-tuned with a LoRA - training on brand colours, product photos, or a person's likeness for consistency; closed models cannot. Midjourney is the odd one out - no public API, so it must be paid for directly. Free tools list: Google ImageFX, ElevenLabs (voice + music), CapCut, DaVinci Resolve, Canva. ComfyUI exists but 'too hard for me to look at'; Flora and Weavy are friendlier node-based tools.

Why it matters

Knowing the aggregator credit costs per model (Krea: Nano Banana Pro 119, ChatGPT Image 1.5 184, Seedream 4.5 32, Krea's own 6) is how you decide which model to iterate on and which to finish on.

Go deeper

In one line: Pay for one aggregator (subscription or pay-per-use) rather than many tools; trials for learning only; LoRA fine-tuning is an open-model privilege; Midjourney alone needs its own subscription.

Realistic total spend $20-50/month via aggregators (l3186012 1:09-1:10)

Subscription camp: Freepik, Krea; pay-per-use camp: fal.ai, Replicate (l3186012 1:18-1:19)

Trials: limited credits, video watermarks, queues - learning only; stack trials (l3186012 1:09, 1:12)

LoRA = train a model on your own dataset for style/likeness consistency (open models only) (l3186012 1:14-1:16)

Midjourney has no public API - subscribe directly (l3186012 1:20)

Krea credit costs: Nano Banana Pro 119, ChatGPT Image 1.5 184, Seedream 4.5 32, Krea model 6 (l3186012 2:19-2:20)

▶ Watch this taught:

05

Pre / production / post: the coffee commercial built live

how-to

A man wakes up a zombie, drinks the coffee, and turns into a suit. Thirty seconds of ad, two hours of iterations - watch where the time actually goes.

Three phases named and used: pre-production (storyboard), production (generate images, then video, then music/voice), post-production (edit). The story: 6 AM alarm, zombie shuffle to the kitchen, brew 'Billin' Coffee - Ultra Strong, Zombie Wake Up', sip, lightning transformation into a suited professional, line: 'That's the best coffee I've ever had.'

What the demo teaches beyond the steps: test the same prompt across models when one keeps failing (the alarm clock rendered wrong until Seedream 4.5); insert yourself by multi-image reference and describe wardrobe explicitly or the model copies the reference's clothes; generate at 1K while iterating and upscale to 4K only for the keeper; chain outputs as references; generate coverage (extreme close-up of the portafilter, overhead) so the edit has options; Kling 2.5 Turbo for speed, 15 credits per 5s clip; Flow/Veo 3.1 for the one shot that needs spoken dialogue and timestamped action; and the storyboard is a reference only - it cannot be imported into any generator.

Do it in this order
Why it matters

'This is really how professionals do it' - the professional part is the iteration discipline and the pipeline, not any single tool.

Go deeper

In one line: Storyboard (Figma) -> base images (Krea) -> insert people/products by multi-image reference -> instructional edits -> upscale keepers -> start/end-frame video (Kling) + dialogue shot (Flow/Veo) -> assemble to a track (Premiere).

Three phases: pre-production, production, post-production (l3186012 1:30-1:31)

Set aspect ratio before generating; swap models when an element keeps failing (l3186012 1:37-1:43)

Insert people by multi-image reference at 1K; upscale to 4K only when happy (l3186012 1:44-1:47)

Describe wardrobe explicitly or the reference's clothing is copied (l3186012 2:00-2:03)

Kling: 15 credits per 5s, 30 per 10s; 2.5 Turbo for speed; 2.6 lacked first/last-frame mode (l3186012 2:35, 2:50)

Flow/Veo 3.1 for dialogue + timestamp prompting in one generation (l3186012 2:45-2:48)

Storyboard cannot be imported into a generator - it is a reference only (l3186012 2:35)

▶ Watch this taught:

06

Lawsuits into licenses: the IP weather for generated media

Every tool will animate your face. None of them will animate Tom Cruise's - and the reason is a courtroom, not a capability.

OpenAI drew backlash for lax guardrails (full South-Park-style episodes), then blocked 'famous people' outright; Disney then invested in OpenAI to license its characters for legal generation. Meanwhile Disney is suing Midjourney (mid-2025), joined by Warner and Universal, and another suit targets a company heard as 'Haylou' (possibly not the video tool). The pattern he expects, borrowed from music (Sony, Universal, Warner vs the music generators): sue, then settle into partnership and investment - 'a big theme of 2026, that the dust is gonna start settling.'

Practical consequences from Q&A: all tools accept your own photo/video but filter well-known faces; starting from your own drawing does not create stronger legal ownership - 'AI law is unsettled' - it just changes the creative input.

Why it matters

Client work with recognizable faces or branded characters is where a cheap generation becomes an expensive letter.

Go deeper

In one line: Celebrity/IP generation is guardrailed by policy after backlash; studios are simultaneously suing (Midjourney) and licensing (OpenAI-Disney); expect 2026 settlements to define what is allowed.

OpenAI added strong celebrity guardrails after backlash; Disney invested to license IP (l3186012 0:51)

Disney v. Midjourney (mid-2025), Warner and Universal joined (l3186012 0:52-0:53)

Music precedent: lawsuits became partnerships/investments (l3186012 0:52)

Prediction: 2026 is the settling year (l3186012 0:53)

Using your own drawings as input does not establish stronger ownership under current law (Q&A) (l3186012 0:52-0:53)

▶ Watch this taught:

Every concept, three clicks deep

The same concepts as a quick reference: the closed row is the glance, open is the study card, and every timestamp jumps into the recording.

01'The new cameras': the 2022-2025 timeline and how to read the leaderboardGenerative image/video models = interchangeable 'cameras' with different aesthetics;

Generative image/video models = interchangeable 'cameras' with different aesthetics; pick a small personal kit and re-check Artificial Analysis rather than chasing every release.

2022 text, 2023 Midjourney/Runway, 2024 experimental video, 2025 production-ready with workflows (l3186012 0:31-0:35)

'These generative AI image and video models are the new cameras' (l3186012 0:39)

Artificial Analysis = anonymous left/right crowd votes; ChatGPT Image 1.5 led Nano Banana Pro that week (l3186012 0:40-0:41)

Kling called one of the strongest yet under-recognized video models (l3186012 0:38)

His kit: Seedream, Nano Banana Pro, Krea, Flux (images); Seedance, Kling, Runway, Hailuo (video) (l3186012 0:56-0:57)

02What Nano Banana changed: instructional editing, multi-image reference, consistencyInstructional edits + multi-image reference + reference-chaining shot to shot = character/scene/product con…

Instructional edits + multi-image reference + reference-chaining shot to shot = character/scene/product consistency; the late-2025 unlock for multi-shot AI storytelling.

Nano Banana brought natural-language editing, image references, camera-angle changes (l3186012 0:35-0:36)

Start/end-frame video and synchronized dialogue are late-2025 capabilities (l3186012 0:36)

Multi-image reference (insert a real person into a generated scene) new 'since about June or July' (l3186012 1:45)

Chain each output as the next shot's reference to hold consistency (l3186012 2:17)

Character, scene, and product consistency named as the enabling capabilities (l3186012 1:29-1:30)

03Three prompt grammars: image (comma list), video (scene sentence), edit (instruction)Image = comma-separated elements, most important first, ending in camera terms;

Image = comma-separated elements, most important first, ending in camera terms; video = descriptive scene sentences; edit = direct imperative instructions.

Order: subject, action, scene, atmosphere, then camera/lens/lighting/film stock (l3186012 1:00-1:01)

Front of the prompt carries the most emphasis (l3186012 1:00-1:01)

Images: comma-separated elements, not sentences (l3186012 1:43)

Video: describe the scene in sentences; edits: imperative instructions (l3186012 2:38)

Camera vocabulary: Dutch/low angle, close-up/medium/cowboy, golden hour, rim light, 50mm vs 10-20mm (l3186012 1:04-1:06)

04$20-50 a month, not per tool: aggregators, trials, open vs closedPay for one aggregator (subscription or pay-per-use) rather than many tools;

Pay for one aggregator (subscription or pay-per-use) rather than many tools; trials for learning only; LoRA fine-tuning is an open-model privilege; Midjourney alone needs its own subscription.

Realistic total spend $20-50/month via aggregators (l3186012 1:09-1:10)

Subscription camp: Freepik, Krea; pay-per-use camp: fal.ai, Replicate (l3186012 1:18-1:19)

Trials: limited credits, video watermarks, queues - learning only; stack trials (l3186012 1:09, 1:12)

LoRA = train a model on your own dataset for style/likeness consistency (open models only) (l3186012 1:14-1:16)

Midjourney has no public API - subscribe directly (l3186012 1:20)

Krea credit costs: Nano Banana Pro 119, ChatGPT Image 1.5 184, Seedream 4.5 32, Krea model 6 (l3186012 2:19-2:20)

05Pre / production / post: the coffee commercial built liveStoryboard (Figma) -> base images (Krea) -> insert people/products by multi-image reference -> instructiona…

Storyboard (Figma) -> base images (Krea) -> insert people/products by multi-image reference -> instructional edits -> upscale keepers -> start/end-frame video (Kling) + dialogue shot (Flow/Veo) -> assemble to a track (Premiere).

Three phases: pre-production, production, post-production (l3186012 1:30-1:31)

Set aspect ratio before generating; swap models when an element keeps failing (l3186012 1:37-1:43)

Insert people by multi-image reference at 1K; upscale to 4K only when happy (l3186012 1:44-1:47)

Describe wardrobe explicitly or the reference's clothing is copied (l3186012 2:00-2:03)

Kling: 15 credits per 5s, 30 per 10s; 2.5 Turbo for speed; 2.6 lacked first/last-frame mode (l3186012 2:35, 2:50)

Flow/Veo 3.1 for dialogue + timestamp prompting in one generation (l3186012 2:45-2:48)

Storyboard cannot be imported into a generator - it is a reference only (l3186012 2:35)

06Lawsuits into licenses: the IP weather for generated mediaCelebrity/IP generation is guardrailed by policy after backlash;

Celebrity/IP generation is guardrailed by policy after backlash; studios are simultaneously suing (Midjourney) and licensing (OpenAI-Disney); expect 2026 settlements to define what is allowed.

OpenAI added strong celebrity guardrails after backlash; Disney invested to license IP (l3186012 0:51)

Disney v. Midjourney (mid-2025), Warner and Universal joined (l3186012 0:52-0:53)

Music precedent: lawsuits became partnerships/investments (l3186012 0:52)

Prediction: 2026 is the settling year (l3186012 0:53)

Using your own drawings as input does not establish stronger ownership under current law (Q&A) (l3186012 0:52-0:53)

Tools referenced

ToolCoverageMomentContext
MidjourneydemonstratedBrief test prompts; closed, no API, Discord-born
KreademonstratedMain aggregator for the build; per-model credit costs shown
Nano Banana ProdemonstratedMulti-image reference inserts and instructional edits
SeedreamdemonstratedByteDance image model (4.5) that fixed the alarm-clock shot; Seedance is the video sibling
KlingdemonstratedStart/end-frame image-to-video; 2.5 Turbo; his go-to
Google FlowdemonstratedVeo 3.1 frames-to-video with timestamps and dialogue
VeodemonstratedVeo 3.1 inside Flow
fal.aidemonstratedPay-per-use re-run of the dialogue shot, cheaper
FigmademonstratedStoryboard template; lock the layer
Adobe Premiere ProdemonstratedFinal assembly to a pre-made track
Artificial AnalysisexplainedAnonymous left/right leaderboard for image and video models
RunwaymentionedEarly video model; 'demos don't totally match the reality'
SoramentionedSora 2 - social-oriented, auto multi-shot, less controllable
FluxmentionedBlack Forest Labs, open source, LoRA-able
IdeogrammentionedHistorically strong on legible text
HailuomentionedMiniMax video; strong motion physics; free trial
HiggsfieldmentionedMulti-angle shot tool; failed to load
ComfyUImentionedNode-based local tool, deliberately not covered
FreepikmentionedSubscription aggregator
ReplicatementionedPay-per-use aggregator
ElevenLabsmentionedVoice + music; not demoed for time
CapCutmentionedFree editor alternative
DaVinci ResolvementionedFree professional editor
CanvamentionedBeginner storyboard + free video editor

Action items

    Resources mentioned

    Resources
    • docDay 2 workbook (free-tools version)
    • docCamera / shot cheat sheet (Google Sheet)
    • docPromised: tool list + prompt screenshots (Notion)

    Extraction notes

    This page was built from an auto-generated transcript, which garbles product and people's names. Those were corrected silently in everything above and logged here for transparency. The warnings flag claims that were true on the recording day but change fast.

    Transcript corrections applied

    The transcript saysThe trainer actually means
    BenkoffBillinkoff (self-confirmed)
    Kria / KoreaKrea (krea.ai)
    Cling / Clang AIKling (Kuaishou)
    Haylou / HeylooHailuo (MiniMax)
    foulfal.ai
    Revpossibly Recraft - unclear
    QuenQwen (Alibaba)
    Vivo's Instagramprobably Vaibhav's Instagram (referenced again at 4:38)

    True on recording day — verify before relying