The raw-video spec the clone trains on
Smile too much in the training video and every serious script comes out grinning.
3-5 minutes; phone camera; static plain background - a white wall in good sunlight beats a bad green screen; tripod or improvised stand; front-facing; soft natural light; neutral expression; natural, generic hand gestures not tied to words; hands and mic never cover the mouth; solid clothing without patterns; no background noise. Avoid overexposure, moving or blurry backgrounds, fast head or hand movement. Test-record 10-15 seconds first; landscape for long-form, portrait for reels. The content is irrelevant - HeyGen analyses hand, eye and lip movement, so read anything (a ChatGPT script will do).
Every flaw in this one recording is replicated in every video the clone ever makes.
