What you’ll be able to do
Right, so here’s the honest version of what’s actually possible. You can write a script with an AI assistant, turn that into a proper shot list, generate or film the clips, bolt on a voice-over, edit it together, and add captions. That’s the whole pipeline. It works, but nobody tells you upfront that you’ll generate ten clips to get two usable ones.
What breaks down is anything that needs to look like a specific real person, anything with a brand name you don’t own, and keeping things consistent across multiple clips. Ask a text-to-video tool for “the same character in scene two” and you’ll often get someone who looks vaguely related but not quite right. Faces drift. Hands do strange things. It’s improved a lot, but it’s not solved.
So the workflow is real. The polish you’re imagining from a big-budget ad is not, not yet anyway.
Before you start
Set your expectations on length first. These tools make clips of a few seconds each, not scenes, not films. Most sit somewhere around 5 to 10 seconds per generation, depending on the tool. A one-minute video means stitching together six, eight, maybe more of these, and that’s before you’ve thrown away the failed attempts.
On usable-clip ratio: expect to generate several takes for each shot before one looks right. This isn’t a fluke of a bad prompt, it’s just how the current generation of tools behaves. Build the time for that into your plan, because it adds up fast.
Commercial licensing matters more than people think. Each tool, Sora, Veo, Runway, Pika, Kling, Luma, has its own terms about what you can do with the output, whether you own it outright, and whether there are restrictions for commercial use. Read the terms for whichever tool you’re using before you put anything in front of a client or a paying audience. Don’t assume they’re all the same, because they’re not.
And disclosure: if your footage could plausibly be mistaken for something real, if it shows people who don’t exist, events that didn’t happen, or anything that could mislead a viewer, label it as AI-generated. This isn’t just good manners, several platforms now expect it.
Get set up
For the script and shot list, use ChatGPT or Claude. Either works fine here, it’s just text generation and structuring, nothing fancy required.
For generating the actual clips, you’re choosing from Sora, Google’s Veo (available through Gemini and Google’s Flow tool), Runway, Pika, Kling, or Luma. They all do roughly the same job, text-to-video generation of short clips, but the style, quality, and quirks differ tool to tool. I’d suggest picking one and getting familiar with its particular flavour of weirdness rather than hopping between all six.
If you want a talking head instead of generated scenes, look at HeyGen or Synthesia. These create avatar-presenter videos from a script, useful for explainer content or anything where a person needs to talk directly to camera without you filming an actual person.
For voice-over, ElevenLabs is the standard choice. You feed it a script, pick or clone a voice, and get audio you can drop into your edit.
For assembly, captions and the final edit, CapCut or Descript both do the job. Descript leans more toward transcript-based editing, which is handy if you’re working from a voice-over script anyway. CapCut is quicker for straightforward cutting and caption styling.
Try it yourself
Start with the script. Here’s a prompt that gets you something usable rather than something generic:
Write a 30-second video script for [describe your topic/product/message].
The tone should be [casual/professional/energetic - pick one].
Include only spoken lines, no scene directions.
Keep it under 75 words so it fits a 30-second read at natural pace.
End with a clear single call to action.
Once you’ve got a script you’re happy with, break it into shots:
Take this script: [paste your script]
Break it into a shot list for a short video. For each shot, give me:
- What's happening on screen
- Roughly how many seconds it should last
- Whether it needs a person speaking to camera, a wide shot, a close-up, or a text overlay
Keep shots to 5-10 seconds each. Aim for [X] total shots to match a [Y]-second video.
Now turn each shot into a generation prompt for whichever text-to-video tool you’re using:
For this shot from my shot list: [paste one shot description]
Write a detailed prompt I can use in a text-to-video generator.
Include camera angle, lighting, setting, mood, and any movement.
Keep it under 40 words and avoid naming real people, real brands, or real locations.
Run that last one through your chosen generator, expect a handful of attempts before you get a clip you like, and repeat for every shot on your list.
Check the result
Once you’ve got clips back, don’t just drop them into your edit and hope. Go through each one against the shot list. Does it actually show what you asked for, or has the tool interpreted “wide shot of a busy street” as something that technically has a street in it but nothing else you wanted?
Watch the transitions once you’ve cut them together. Jarring cuts between clips generated separately are common, because lighting, colour grading and pacing can vary shot to shot even within the same tool. If it looks like six different videos stapled together, that’s normal at first, and it’s what your edit is for.
Check the voice-over sync if you’re using one. Does the pacing of the spoken audio match the visual beats? A voice-over that finishes a sentence three seconds before the clip changes feels off, even to viewers who couldn’t tell you why.
Look closely for visual glitches, extra fingers, faces that shift slightly between frames, text in the background that’s gibberish. These are still common tells, so don’t just skim the footage at speed, actually watch it.
And ask the honest question: does it look like AI, and does that matter for what you’re using it for? For some things, an obviously synthetic look is fine, even part of the appeal. For others, especially anything meant to look documentary or real, it’s a problem, and you need to decide now whether to keep iterating or change your approach.
If it doesn’t work
Don’t feed any of these tools real names, real faces, or real brand names, whether that’s in your script prompt or your shot list. Generating a video of a specific real person without consent, or of a company’s branding without permission, is a fast way to get flagged, blocked, or worse, land yourself in genuinely difficult territory. Keep it generic. “A person in a kitchen” not “my colleague Dave”.
Check each tool’s terms before you use the output commercially. This is worth repeating because it’s the step people skip. Terms differ between Sora, Veo, Runway, Pika, Kling and Luma, and they change over time. Don’t assume last year’s terms still apply.
Label generated footage anywhere it could mislead a viewer into thinking it’s real. This matters more for anything resembling news, testimonials, or documentary style content.
If a shot keeps coming out wrong, try more variations before you give up on the whole script. Sometimes it’s the prompt that’s vague, not the tool that’s broken. Add more specific detail about camera angle, lighting, or movement, and try again a few times before you scrap the idea entirely.
Keep exploring
If you’ve got clips and need to actually put them together properly, there’s a guide on this site called “How to get AI to edit a video” that goes deeper into that part.
If you need a voice reading a script rather than generating footage at all, see “How to get AI to read text out loud”.
And if you just need a still image rather than motion, “How to get AI to create an image” covers that separately.
Sources and review notes
- Text-to-video — Sora: https://openai.com/sora · Veo: https://deepmind.google/technologies/veo/ · Runway: https://runwayml.com/ · Pika: https://pika.art/ · Kling: https://kling.ai/ · Luma: https://lumalabs.ai/
- Avatars and voice — HeyGen: https://www.heygen.com/ · Synthesia: https://www.synthesia.io/ · ElevenLabs: https://elevenlabs.io/
- Editing — CapCut: https://www.capcut.com/ · Descript: https://www.descript.com/
Review date: 15 September 2026 — every source above was opened and checked against the text on that date.