Tuesday, 29 September 2026 | Updating Daily AI insight, written for builders

How to Make AI Videos: The Complete Beginner’s Guide

  • Pick the kind of video first. Text-to-video (type a description), image-to-video (animate a photo you already have), avatar video (a digital presenter reads your script), or AI-assisted editing (captions, voiceover, dubbing added to real footage).
  • Easiest places to start: Canva’s Create a Video Clip makes clips up to 8 seconds with synchronised sound. CapCut’s AI video maker turns a script into a video, offers 100+ avatars, and uses the Seedance 2.5 model.
  • Want longer clips? Seedance 2.5 makes clips up to 30 seconds with sound; MiniMax H3 (Hailuo 3.0) makes up to 15 seconds at 2K with stereo sound.
  • Long videos are stitched, not generated. Make several short clips, join them in an editor, add voiceover and captions. Check current prices on each vendor’s own page, and label AI content.

To make an AI video: choose the type you need, write a short prompt describing the shot, generate a clip (most tools cap between 8 and 30 seconds), then join several clips in a free editor and add voiceover and captions. Canva and CapCut both run in a browser, so you can finish a first video today without installing anything.

The rest of this guide explains the four kinds of AI video, which tool suits which job, what things actually cost, and a first project you can copy.

The four kinds of AI video

Almost every tool you will meet does one of these four things. Knowing which one you want saves hours.

1. Text-to-video

You type a description — “a golden retriever running through tall grass at sunset, handheld camera” — and the tool invents the footage. Good for scenery, product moods, abstract B-roll and short social clips. Weak at anything needing exact detail: readable signs, specific faces, brand logos, or a character who must look identical in ten shots.

2. Image-to-video

You upload a still picture and the tool animates it: the camera drifts, hair moves, steam rises. This is usually the most reliable route for beginners, because you control how the scene looks before any motion is added. Photographers, shop owners and anyone with a product photo should start here. We cover the workflow in detail in our guide to turning an image into a video, and if you need the still first, see how to make AI images.

3. Avatar and talking-head videos

You paste a script and a digital presenter reads it to camera, lip-synced. This is the standard for training videos, explainers, internal announcements and multi-language versions of the same message. Synthesia offers 240+ avatars in 160+ languages; CapCut’s AI video maker includes 100+ avatars. You can usually pick a stock presenter, or create one from consented footage of a real person.

4. AI-assisted editing

The quietest category and often the most useful. Here the video is real — filmed on your phone — and AI only helps with the work around it: auto-captions, background removal, filler-word trimming, text-to-speech narration, and dubbing into another language. CapCut and Canva both do a lot of this. If narration is your main need, our AI voice-over walkthrough goes deeper.

Which tool for which job

What you want Kind Reasonable choice Verified detail
A quick social clip, no install Text-to-video Canva, Create a Video Clip Up to 8 seconds, with synchronised sound
Script straight to finished video Mixed CapCut AI video maker Script to video, 100+ avatars, uses Seedance 2.5
The longest single clip you can get Text- or image-to-video Seedance 2.5 Up to 30 seconds, with sound
Sharper clip with stereo audio Text- or image-to-video MiniMax H3 (Hailuo 3.0) Up to 15 seconds, 2K, stereo sound
Training or explainer with a presenter Avatar Synthesia 240+ avatars, 160+ languages
Captions, dubbing, tidy-up on real footage AI-assisted editing CapCut or Canva Both run in a browser

For a wider shortlist, see our roundups of the best AI video generators and the best free AI video generators.

Your first project: a 20-second video

Plan on 30 to 60 minutes the first time. The trick is to stop thinking “make a video” and start thinking “make four short shots and join them”.

  1. Write four shots on paper. One line each, five seconds each. Example: (1) coffee beans falling into a grinder, close up; (2) steam rising from a cup on a wooden counter; (3) hands sliding the cup across the counter; (4) the shopfront at golden hour.
  2. Turn each line into a prompt. Use the formula below. If writing is the part you dread, ask a chatbot to draft ten prompt variations, then pick the two you like.
  3. Generate one shot at a time. Expect to run each prompt two or three times. Rejecting takes is normal, not failure. Save every usable clip.
  4. Record or generate narration. Your own voice on a phone still sounds better than most synthetic reads. If you use text-to-speech, keep sentences short.
  5. Stitch, caption, export. Drop the clips on a timeline in CapCut or Canva, trim the weak first and last frames of each clip, add music low, auto-generate captions, then read the captions yourself — auto-captions mangle names and numbers.
  6. Watch it once on mute and once on a phone. Most mistakes reveal themselves in those two passes.

A prompt formula that works

Subject, then action, then camera, then light, then style. “A ceramic cup of coffee on a worn wooden counter, steam curling upward, slow push-in, warm morning light from a side window, shot on 35mm film.” Add what you do not want if the tool supports it. Keep one idea per clip — prompts that ask for three events in five seconds produce mush. Our image-to-prompt tool can also reverse-engineer a description from a picture you like.

What AI video costs

Two different pricing worlds exist, and mixing them up is the most common beginner mistake.

Consumer apps (Canva, CapCut, Synthesia and similar) sell subscriptions, usually with credits or a monthly cap on generations. Those limits change often, so check the vendor’s own pricing page — for example Synthesia’s pricing page — rather than trusting a figure in an article.

Developer access to the same underlying models is billed per second of finished video. These are Convly’s published model figures, and they are useful as a sanity check on what a clip is worth:

Model Maker From
Wan 2.5 Alibaba $0.05 per second
Veo 3.1 Google $0.05 per second
Kling 2.5 Turbo Pro Kuaishou $0.07 per second
Seedance 2.5 ByteDance $0.097 per second

At those rates a 20-second video costs roughly one to two dollars of raw generation — before the takes you throw away, which can easily triple it. Google publishes its own per-second video rates on the Gemini API pricing page.

Important: those per-second numbers are developer prices. They are not what a Canva, CapCut, ChatGPT or Gemini subscription gives you, and you cannot use them to work out how many clips a plan includes. For plan-by-plan detail, see our guides to ChatGPT’s Free, Go, Plus and Pro tiers and Google’s AI plans.

If you use a chatbot to write scripts, that side is cheap. Text models are billed per token — a token is roughly three-quarters of a word. GPT-6 Luna runs $0.10 in / $0.50 out per million tokens with a 1.05M-token context window (the amount of text the model can hold in mind at once), and Gemini 3.8 Flash is $0.75 in / $3.75 out per million with a 1M context. A whole afternoon of scriptwriting costs pennies. The video is where the money goes.

Windows, macOS and Linux

Nearly all of this is browser work, so your operating system matters less than your internet connection. Where it does matter:

Windows

Best-supported desktop setup. CapCut offers a Windows desktop app alongside its web editor, and Adobe and DaVinci Resolve both run natively if you outgrow free tools. If your PC has a discrete NVIDIA graphics card you can also experiment with open models locally, though video generation is far heavier than image generation and requirements shift with every release — treat local video as a hobby project, not your first one.

macOS

Also well supported: CapCut has a Mac app, and Canva runs in Safari or Chrome. Apple Silicon Macs handle timeline editing and export comfortably. Cloud generation behaves identically to Windows.

Linux

No official CapCut desktop app. Use the browser versions of Canva, CapCut and the generator sites, which work in Chrome or Firefox. For editing, DaVinci Resolve has a Linux build and Kdenlive is a solid free option. Nothing about cloud AI video is blocked on Linux — only some desktop apps are.

One tool to skip: Sora

Do not build a workflow around OpenAI’s Sora. The Sora app closed on 26 April 2026, and the Sora 2 API stopped on 24 September 2026 with no replacement named. Older tutorials still recommend it; they are out of date. OpenAI lists retirements on its deprecations page. This is a good habit generally: before following any AI video tutorial, check that the tool still exists.

Honesty, consent and labels

A few lines that keep you out of trouble:

  • Do not use a real person’s face or voice without their clear permission. That includes colleagues, celebrities and people in photos you found online. Avatar tools ask for consent footage for a reason. Our piece on deepfakes and how to spot them explains why platforms treat this seriously.
  • No fake reviews or invented testimonials. A synthetic person praising your product is not a customer, and presenting it as one is deceptive. If you are making social ads with AI, read how to make UGC ads without breaking the rules first.
  • Label it. Many platforms require AI-generated or significantly altered video to be disclosed, and several add labels automatically. Rules differ by platform and country, so check the policy of wherever you are posting.
  • Check facts before publishing. Generated footage invents details happily: wrong number of fingers, garbled on-screen text, products that do not exist. For school or work, ask what is permitted before submitting AI-made material.

Common problems and quick fixes

Problem Likely cause Fix
Faces melt or hands look wrong Too much motion, subject too close Move the camera instead of the subject; pull back to a wider shot
On-screen text is gibberish Models are poor at rendering words Add text later in the editor, not in the prompt
Character changes between shots Each clip is generated fresh Use one reference image and image-to-video for every shot
Clip ends mid-action Hit the duration cap Plan shots to fit the limit, or split into two clips
Audio and mouth are out of sync Voiceover added over generated footage Use a purpose-built avatar tool for talking heads

Frequently asked questions

Can I make a full-length AI video?

Not in one generation. Current tools produce short clips — 8 seconds in Canva’s Create a Video Clip, up to 15 seconds for MiniMax H3, up to 30 seconds for Seedance 2.5. Longer videos are made by generating many clips and editing them together, exactly as a human film crew shoots separate takes.

Is there a genuinely free way to do this?

Yes, within limits. Several tools offer free tiers with watermarks, queue waits or a small monthly allowance, and free editors handle the stitching and captions. Allowances change frequently, so check the vendor’s own pricing page rather than any article — including this one. Our free AI video generator roundup is a good starting list.

Do I need a powerful computer?

No. The generation happens on the vendor’s servers, so a basic laptop, Chromebook or even a phone is enough for cloud tools. A stronger machine only helps with the editing and exporting stage. Running video models locally is a different matter and demands serious graphics hardware.

Can I put my own face in an AI video?

Usually yes — several avatar platforms let you create a likeness from footage you record yourself, following their consent process. Using someone else’s face or voice without permission is a different thing entirely and can breach both platform rules and local law.

Which is better, text-to-video or image-to-video?

Image-to-video for anything where the look matters — products, branding, a consistent character — because you approve the frame before it moves. Text-to-video for speed and for scenes you cannot photograph. Most practical projects mix the two.

Why do my clips look worse than the demos?

Demo reels are the best takes out of dozens, often with hand-picked prompts. Assume a two-or-three-to-one reject rate, write one clear idea per prompt, and expect the first version of any project to be a rough draft.

Written by Mustafa Ihsan

Mustafa Ihsan is the founder and editor of Convly.ai. He built and maintains the site's live AI models database, its price-performance index, and its free calculators for VRAM requirements, API costs and self-hosting economics. He writes about model pricing, benchmark results and the hardware needed to run AI models locally, and consistently prefers measured numbers to vendor claims.

Scroll to Top