← AI Creator CourseAll programsHomeSearch
AI Creator Course·AI Video Creation·2:08:55

Sections 4-5: AI Clones and AI Video — the 5-Step Workflow, Image-to-Video Prompting, First/Last Frame, HeyGen, Runway Aleph, ElevenLabs Voice, and Wan 2.2 Replace

Anthony Gallo Instructor - argues for a 'hybrid human plus AI workflow' against one-click video tools, and builds a dinosaur short, a camera commercial and a hand-on-fire effect on screen

The short version

  1. The workflow never changes, AI or not: idea (the 'seed') -> script -> storyboard -> video creation -> editing. One-click whole-video tools 'usually create terrible looking videos' - AI slop. Video creation splits into three: human-AI hybrid, image-to-video, and long-form talking-head tools. The average movie shot is 2.5-4 seconds, so 5-10 s AI clips are not a limit.
  2. Clones: 'if the AI images don't look like you... it's not the tool, it's your reference photos.' One person per photo, rear 1x camera, good light; best of all a CHARACTER SHEET - two full-body angles plus four head close-ups in one PNG (Canva template). Prompt six things: subject/action, environment, outfit, expression, lighting, style.
  3. Image-to-video is the default ('90% of AI videos' use it or text-to-video) because the still already carries subject, style, environment, weather and light - so the prompt only needs subject ACTION, CAMERA MOVEMENT ('no camera movement' / 'tripod shot' or pan, tilt, tracking, push, jib), and MOOD. Text-to-video needs subject, framing, environment and style up front, plus lighting, colour of light, time of day, atmosphere, palette; paste the worksheet into ChatGPT to write it.
  4. First frame / last frame: two Nano Banana stills and a prompt; Veo 3.1 for realism on simple transitions, Kling 2.5 for creative ones (car into Transformer). Chain pairs - each clip's last frame is the next one's first - into a multi-shot commercial. Runway Aleph edits an existing clip: day to night, engulf the dinosaur in flame, snow instead of rain, 'a new wider angle version of this clip.'
  5. HeyGen animates one still to talk for minutes (styles stable/expressive; 720p vs 1080p; 16:9, 9:16, 1:1) from a typed script or an uploaded ElevenLabs clone; keep backgrounds simple, generate in segments, and use HeyGen only for the on-camera minutes with voiceover + B-roll between. ElevenLabs: stability 0.1-0.4 for an expressive voice; clone from an 11-minute clean sample; Voice Changer unifies mismatched clip voices from an MP3 export.
  6. Wan 2.2 Replace: screenshot a frame, redraw the character in Nano Banana ('replace the man with a humanoid robot with glowing blue eyes'), feed video + image; movement and lip-sync survive. Large or 4K files fail - compress or split. The one-person, two-character skit and the low-budget short film are the use cases; it is Gollum's pipeline for the price of a coffee.

At a glance, three clicks deep

Skim here first: the closed row is the glance, open is the study card with the key points and timestamps, and the ↓ link drops to that concept's full write-up below.

01The 5-step AI video workflow and the three ways to create the footageIdea -> script -> storyboard (stills) -> video creation (hybrid / image-to-video / talking-head) -> edit;

Idea -> script -> storyboard (stills) -> video creation (hybrid / image-to-video / talking-head) -> edit; keep humans in the loop; short shots are normal.

Five steps verbatim: idea, script, storyboard, video creation, editing (p2194080911 0:01-0:17)

One-click tools make 'terrible looking videos' (p2194080911 0:00-0:01)

Storyboard as stills because image-to-video is the best strategy (p2194080911 0:06)

Three creation routes: hybrid, image-to-video, talking-head (p2194080911 0:10-0:14)

Average movie shot 2.5-4 s (15,000-film study) (p2194080911 0:13)

↓ Full write-up of this concept

02Clones from reference photos: the character sheet and the six-part promptReference photos (rear camera, one person, good light) or a character sheet PNG;

Reference photos (rear camera, one person, good light) or a character sheet PNG; prompt subject/action + environment + outfit + expression + lighting + style; references can replace any element.

'It's not the tool, it's your reference photos' (p2190797801 0:02)

Up to 14 references; rear 1x camera (p2190797801 0:01-0:03)

Character sheet = 2 body angles + 4 heads in one PNG; Canva template (p2190797801 0:04-0:05)

Six prompt elements; strong vs weak prompt (p2190797801 0:06-0:09)

Reference image can stand in for environment (p2190797801 0:10)

↓ Full write-up of this concept

03Image-to-video: why it wins, and the three-part promptStill first;

Still first; Kling default, Veo fallback; prompt = subject action + explicit camera movement + mood; static shots look professional.

'90% of AI videos' use image-to-video or text-to-video (p2186925504 0:01)

Kling cheaper default; Veo pricier fallback; 5-s clips; Generate Sound (p2186925504 0:04-0:08)

Prompt: action, camera movement, mood (p2186925504 0:10-0:12)

Named moves: pan, tilt, tracking, push, jib; 'no camera movement' (p2186925504 0:11-0:12)

Dinosaur short case study end to end (p2186925504 0:13-0:19)

↓ Full write-up of this concept

04Text-to-video worksheet: four required elements, five bonus elements, ChatGPT fills itSubject + framing + environment + style (required);

Subject + framing + environment + style (required); lighting, light colour, time of day, atmosphere, palette (bonus); ChatGPT drafts from the worksheet.

Required four (p2187033318 0:02-0:04)

Bonus five (p2187033318 0:04-0:06)

Framing vocabulary list (p2187033318 0:03)

Worksheet -> ChatGPT -> edit (p2187033318 0:06-0:07)

↓ Full write-up of this concept

05First frame / last frame, and shot sequencing across clipsStart still + end still + prompt;

Start still + end still + prompt; Veo for realism, Kling for creative bridges; chain last->first frames for multi-shot sequences.

Tools: Veo 3.1, Kling 2.5; frames from Nano Banana Pro (p2193539833 0:00-0:02)

Veo realism vs Kling creative transitions (p2193539833 0:01)

Durations 4/6/8 s; Veo auto audio (p2193539833 0:04-0:05)

Shot-sequencing camera commercial across three pairs (p2193539833 0:10-0:15)

↓ Full write-up of this concept

06Runway Aleph: edit the video you already have - effects, weather, new anglesAleph = prompt-driven edits of an existing clip (effects, weather, relight, new angles) with consistency;

Aleph = prompt-driven edits of an existing clip (effects, weather, relight, new angles) with consistency; 5 s; Runway subscription + credits.

Video-to-video with a text prompt; consistency preserved (p2190276045 0:00-0:07)

Verbatim prompts for night, flame, wider angle, snow, fire, behind-subject (p2190276045 0:01-0:06)

5-s cap; ~60 credits/gen; 1,000 credits/$10; subscription on Runway (p2190276045 0:02, 0:07)

Hand-on-fire composite in Premiere (p2190276045 0:07-0:09)

↓ Full write-up of this concept

07HeyGen talking heads: minutes not seconds, simple backgrounds, segment and hybridizeHeyGen = long-form lip-synced avatar from one still + script/audio;

HeyGen = long-form lip-synced avatar from one still + script/audio; styles, resolutions, ratios; simple backgrounds; segment; hybridize with voiceover + B-roll.

Script + preset voice, or uploaded audio/clone (p2192666653 0:02-0:04)

Styles, 720p/1080p, ratios (p2192666653 0:03)

Background drift - keep it simple (p2192666653 0:08-0:09)

Segment long videos; hybrid with voiceover + B-roll (p2192666653 0:11-0:14)

↓ Full write-up of this concept

08ElevenLabs three ways: voiceover, voice clone, voice swapTTS (stability 0.1-0.4), clone (long clean sample), changer (MP3 in, one voice out);

TTS (stability 0.1-0.4), clone (long clean sample), changer (MP3 in, one voice out); clone drives HeyGen or Seedance 2.0.

Stability low = expressive; 0.1-0.4 preferred (p2197590891 0:04; p2197590892 0:06)

Clone from an 11-minute clean sample (p2197590892 0:04-0:05)

Clone drives HeyGen or Seedance 2.0 (p2197590892 0:01-0:03)

Voice Changer unifies mismatched clip voices (p2197590890 0:01-0:02)

↓ Full write-up of this concept

09Wan 2.2 Replace: swap the character, keep the performanceVideo + Nano-Banana-edited frame -> Wan 2.2 Replace;

Video + Nano-Banana-edited frame -> Wan 2.2 Replace; guidance/steps/quality settings; split large files; performance preserved.

Workflow: screenshot -> Nano Banana edit -> Wan Replace (p2192689182 0:00-0:02)

Settings: guidance, inference steps, quality, fast (p2192689182 0:02)

Large/4K files fail - compress or split (p2192689182 0:04)

Use cases and the Gollum comparison (p2192689182 0:04-0:06)

↓ Full write-up of this concept

The concepts in full

01

The 5-step AI video workflow and the three ways to create the footage

The workflow is older than AI. AI just swapped what sits under each step.

'The fundamental five step workflow to create a video doesn't change whether you use the traditional tools or AI tools': step one an idea, 'the seed'; two the script; three the storyboard - the first visual step, done as stills because 'the best strategy for making detailed AI videos is... image to video'; four video creation; five editing. Framed as a 'hybrid human plus AI workflow' against one-click tools that 'usually create terrible looking videos' and undifferentiated slop. Step four has three routes: human-AI hybrid (you on camera plus AI assets), image-to-video (Kling, Veo), and long-form talking-head tools (HeyGen, Hedra, Lip Sync on fal.ai). A film scholar's analysis of 15,000 movies puts the average shot at 2.5-4 s, so 5-10 s clip caps are the norm, not a constraint; editing tips follow. Demo: a horror short built end to end.

Why it matters

Every later lesson is a plug-in to one of these five steps; the map keeps them from feeling like forty disconnected tools.

The 5-step video workflow (AI Creator Course) and the tools under each step 1 IDEA the seed ChatGPT ideas + titles 2 SCRIPT your voice ChatGPT, Canvas 3 STORYBOARD as stills Nano Banana Pro 4 VIDEO three routes see below 5 EDIT human craft Resolve / CapCut / Premiere HYBRID you on camera + AI assets IMAGE-TO-VIDEO Kling default, Veo fallback first/last frame, Wan Move TALKING HEAD HeyGen for minutes ElevenLabs voice One-click whole-video tools "usually create terrible looking videos" - keep the human in the loop; shots of 2.5-4 s are normal Image first, then animate: the still already carries subject, style, environment, weather and light
Anthony's 5-step workflow with the tools this course assigns to each step, and the three video-creation routes under step 4.
02

Clones from reference photos: the character sheet and the six-part prompt

how-to

Upload a selfie from the front camera and the clone will not look like you. That is the photo, not the model.

Nano Banana Pro accepts up to 14 reference images; describing your appearance in words loses to uploading photos. Photo rules: one person per image, clear facial detail, good light, the phone's rear 1x lens not the selfie camera. Baseline: a close-up plus a full-body shot. Best: a character sheet - one composited PNG with two full-body angles (front, side) and four head close-ups (neutral, smiling, two profiles), the format character designers have always used; a free Canva template is linked. Make a sheet per look (beard vs clean-shaven, outfits). Prompt six elements every time: subject and action, environment, outfit, expression, lighting, style - or substitute a reference image for any of them ('this man at the spooky castle from the reference image'). Strong example: 'create a cinematic ultra realistic image of the man from the reference image standing on top of a snowy mountain at sunrise... dark winter jacket... confident expression. Soft golden sunlight... dramatic clouds.' Weak: 'make me look cool on a mountain.'

Do it in this order
Why it matters

Paul's client headshots and real-face ads (the MEH punch list) start here.

03

Image-to-video: why it wins, and the three-part prompt

'Make a video of a cat jumping on a chair' - which cat, indoors, what time, what weather? The still already answered all of that.

Image-to-video (still first, then animate) is the default: the image 'gives the AI video tool the context on the subject, the time of day, the style' so output lands closer to intent, and stills are cheaper to iterate than 30-second-plus paid generations. Tools: Kling as default (cheaper; versions 2.1/2.6/3.0 at recording - 'the newest version is always going to be the best'), Veo as fallback. Kling settings: model, aspect ratio, 5 vs 10 s (he stays at 5), Generate Sound toggle. Prompt only what the still cannot show: (1) subject action; (2) camera movement - say 'no camera movement' or 'tripod shot', or name pan left/right, tilt up/down, tracking, push in/out, jib up/down, 'the main camera movements that Hollywood directors use for 90% of the shots'; (3) emotion/mood/tone - running to train or running from a threat. Static shots read as more professional. Case study: ChatGPT idea and script -> ElevenLabs voiceover -> Nano Banana character and shots -> Kling ('The dinosaur looks at the lightning and then runs away') -> Premiere; shoot coverage; the same method supplies B-roll for long-form.

Why it matters

The three-part prompt is the reusable asset - it is the same prompt you will write for every clip in Section 6's walkthroughs.

04

Text-to-video worksheet: four required elements, five bonus elements, ChatGPT fills it

No still to lean on, so the prompt has to be the storyboard.

Required: subject; shot composition/framing (close-up, wide, medium waist-up, long, side profile, macro, drone, over-the-shoulder); environment/setting; visual style. Bonus: lighting style (direction and intensity in plain words), colour of light (daylight white vs neon or orange), time-of-day light (sunset, sunrise, blue hour), atmosphere (fog, rain, snow, wind, smoke, dust, embers), colour palette (saturated, noir, or 'like Blade Runner'). Tip: paste the worksheet into ChatGPT with a one-line scene description and let it draft the prompt, then edit rather than accept.

Why it matters

A fill-in-the-blanks brief for any generator - and the same vocabulary Section 3's cinematography prompts use.

05

First frame / last frame, and shot sequencing across clips

how-to

Give it where the shot starts and where it ends. Then make the ending of each shot the beginning of the next.

Two Nano Banana Pro stills plus a prompt; Veo 3.1 or Kling 2.5 on PromptEdit connect them. Veo 3.1: slightly better realism and auto sound effects, but 'can struggle... with complex transitions' - it cuts or fades; Kling 2.5 'is really, really good at creatively transitioning' (car into Transformer, bridging unrelated honeymoon clips) at slightly lower quality. Durations 4/6/8 s; 16:9 used. Examples: barn and horse, a hat transition, a logo animation. Shot sequencing: chain pairs so each clip's first frame is the previous clip's last - the camera commercial floats, deconstructs, reforms in a hand, then the reviewer smiles; each pair made in Nano Banana ('keeping the camera about the same size in the frame, but now it's being held by a blonde girl... field of flowers at sunset'), animated in Kling 2.5, cut in Premiere.

Do it in this order
Why it matters

Sequencing is how you get a 30-second continuous piece out of 6-second generators.

06

Runway Aleph: edit the video you already have - effects, weather, new angles

'Create a new wider angle version of this clip.' Same actor, same motion, a camera you never had.

Video-to-video: an existing clip plus a text change. Examples verbatim: 'change the scene to take place at night with a large full moon' (mixed), 'engulf the dinosaur in flame', 'create a new wider angle version of this clip', 'change the weather outside to snow instead of rain', 'outside of the window a large fire rages and burns', 'create a new angle from directly behind the person... symmetrical position, subtle push in.' Best on additive, localized effects and reframing; full day/night reversals were weaker. 5-second cap. Subscription on Runway directly (not PromptEdit), affordable entry tier plus ~60 credits per generation, 1,000 credits for $10. The hand-on-fire finish: real clip + Aleph 'arm and hand is on fire' clip cut together in Premiere with sound effects.

Why it matters

This is the cheapest way to get coverage and VFX on footage Paul already shot for a client.

07

HeyGen talking heads: minutes not seconds, simple backgrounds, segment and hybridize

Cinematic tools stop at ten seconds. HeyGen will let a still talk for as long as your script runs - as long as nothing behind it has to move.

Input one image (a Nano Banana clone) plus a typed script with a preset voice, or an uploaded audio file - a real recording or an ElevenLabs clone - that HeyGen lip-syncs. Settings: talking style (stable/default/expressive), 720p vs 1080p (cost scales), 16:9 / 9:16 / 1:1. Weakness: background motion drifts ('the smoke will almost start going backwards') - keep backgrounds simple. Practices: split long pieces into per-segment generations so one flub does not re-run five minutes; the cost-saving hybrid - HeyGen for the first 30 s, a middle minute and the last 30 s of a ten-minute video, ElevenLabs voiceover plus B-roll between; for dynamic scenes with a consistent voice, generate in Kling/Veo and run a voice-swap pass. Also on PromptEdit now, where it 'used to be' website-only at a high price.

Why it matters

The faceless-channel and course-narration use cases Paul quotes for clients are exactly this.

08

ElevenLabs three ways: voiceover, voice clone, voice swap

how-to

Every clip came back with a different voice. Export the audio, run it through your clone, and they all speak as one person.

All via PromptEdit's Audio tab. Voiceover: paste the script, pick a voice, set stability (low = expressive, high = monotone; he likes 0.1-0.4), language, generate, download - commercials, faceless channels, audiobooks, course narration. Clone: name it, upload a long clean sample (his: 11 minutes), leave 'remove background noise' unchecked unless the source is noisy, tune stability to your natural delivery; feed the clone to HeyGen or to Seedance 2.0, which - unlike Kling/Veo that 'use whatever voice it wants' - accepts your audio to drive a cinematic clip. Swap: export audio-only MP3 from the Premiere timeline, upload to Voice Changer, choose the target voice (preset or clone), stability, generate one consistent track.

Do it in this order
Why it matters

Voice consistency is the tell in most AI video; this is the fix, and it is three clicks.

09

Wan 2.2 Replace: swap the character, keep the performance

Film yourself once. Replace yourself with a robot, a second character, anyone - the movement and lip-sync stay.

Screenshot a frame of your video; in Nano Banana Pro redraw the subject ('replace the man in the image with a humanoid robot with glowing blue eyes. Everything else... the same'); feed the original video and the new image into Wan 2.2 Replace. Settings: guidance scale, inference steps (raise if results look weird, at time cost), video quality (max), fast mode. Movement, facial motion and lip-sync carry over. Large or 4K files fail - compress, lower resolution, or split. Uses: low-budget shorts, social hooks, two-character skits filmed solo; the Gollum motion-capture pipeline without the suit. Wan 2.2 Move (background swap) is 'the next video' - not captioned here.

Why it matters

A one-person production company's cast extender.

Tools referenced

ToolCoverageMomentContext
Nano Banana ProdemonstratedClones, character sheets, storyboard stills, first/last frames
PromptEditdemonstratedFront end for Kling, Veo, HeyGen, ElevenLabs, Wan
KlingdemonstratedDefault image-to-video; 2.5 for first/last frame
VeodemonstratedVeo 3.1 realism and audio; first/last frame
HeyGendemonstratedLong-form talking head
ElevenLabsdemonstratedTTS, clone, Voice Changer
RunwaydemonstratedAleph video-to-video; subscription + credits
WandemonstratedWan 2.2 Replace character swap
Adobe Premiere ProdemonstratedAssembly, hand-on-fire composite, MP3 export
ChatGPTexplainedIdeas, script, prompt drafting
SeedancementionedSeedance 2.0 accepts your own audio
CanvamentionedCharacter-sheet template
HedramentionedTalking-head alternative
PixversementionedSource robot clip for Aleph demo
fal.aimentionedSource clips; Lip Sync tool

Session materials

Archived locally on V: — click to open. Companion pages link to the LMS.

Action items

    Extraction notes

    This page was built from an auto-generated transcript, which garbles product and people's names. Those were corrected silently in everything above and logged here for transparency. The warnings flag claims that were true on the recording day but change fast.

    Transcript corrections applied

    The transcript saysThe trainer actually means
    Cling / ClangKling
    11 Labs / Eleven labsElevenLabs
    Alif / ALFAleph (Runway)
    Juan 2.2 / 1.2.2Wan 2.2
    Sea DanceSeedance
    hey JenHeyGen
    Fall AIfal.ai
    prompted itPromptEdit.com
    the One X camerathe phone's 1x rear lens

    True on recording day — verify before relying