Saturday, 29 August 2026 | Updating Daily AI insight, written for builders

How to Turn an Image Into a Video With AI (2026): Models, Costs and Free Options

Key takeaways

  • Every current frontier video model accepts an image as the starting frame
    Sora 2, Veo 3.1, Kling 2.5 Turbo Pro and Wan 2.5 all take text or image input.
  • Image-to-video costs the same per second as text-to-video on these models.
    You are billed for output length, not for how the shot was started.
  • The cheapest route is $0.05 per second — Veo 3.1 Lite at 720p or Wan 2.5 —
    so animating a still into a 5-second clip costs about 25 cents per attempt.
  • Free tiers exist but are small and change often. Runway, for example, gives
    125 one-time credits that do not renew.

Turning a still image into a moving clip is now a first-class feature rather than a research
demo, and it is the one video task where the input you already have — a photo, a product shot, a
piece of artwork — does most of the work. This page covers which models accept an image, what
each charges per second of output, and where the free options stop being useful.

Quick answer: how do you turn an image into a video with AI?

You give a video model your image as the first frame and a short prompt describing the motion
you want, then pay per second of output. Every frontier model on the market supports this:
Sora 2, Veo 3.1, Kling 2.5 Turbo Pro and Wan 2.5 all accept text or image input.
The cheapest is $0.05 per second of finished video, which puts a 5-second animation at roughly
$0.25 per generation, before retries. Free browser tools exist and will animate a photo without a
card, but they cap resolution, add watermarks and hand out a fixed allowance rather than a
recurring one.

Which models accept an image as input

This is the part most comparison pages leave out. Image input is not a separate product tier —
it is a mode of the same model, priced identically:

Model Developer Input From (per second)
Veo 3.1 Google Text, Image $0.05 (Lite, 720p)
Wan 2.5 Alibaba Text, Image $0.05 (via fal)
Kling 2.5 Turbo Pro Kuaishou Text, Image $0.07 (via fal)
Sora 2 OpenAI Text, Image $0.05 (batch, 720p)
Sora 2 Pro OpenAI Text, Image $0.15 (batch, 720p)

One timing note that matters if you are choosing today: OpenAI has set
24 September 2026 as the shutdown date for the Videos API and both Sora 2 models,
with no replacement named. If you are starting an image-to-video workflow now, starting it on Sora
is starting it on a countdown.

What it costs to animate one image

Because billing is per second of output, the cost of an image-to-video clip is simply the rate
multiplied by the length you ask for:

Model and tier 5-second clip 10-second clip 30-second clip
Veo 3.1 Lite, 720p $0.25 $0.50 $1.50
Wan 2.5 $0.25 $0.50 $1.50
Kling 2.5 Turbo Pro $0.35 $0.70 $2.10
Veo 3.1 Lite, 1080p $0.40 $0.80 $2.40
Veo 3.1 Fast, 1080p $0.60 $1.20 $3.60
Veo 3.1 Standard, 4K $3.00 $6.00 $18.00

Those are per successful generation. Image-to-video tends to need fewer attempts than
text-to-video, because the first frame is already fixed and you are only directing motion — but a
shot that takes three tries still costs three times the figure above. Budget from your own retry
rate, not from the tariff.

Is there a genuinely free way to do this?

Yes, with limits that are worth understanding before you rely on them. Free tiers in this
category are promotional rather than structural: they exist to get you to the paid plan, and their
terms move.

Runway is the clearest published example. Its pricing page offers
125 one-time credits to try the tools — a fixed deposit that does not renew — with
paid plans starting at $12 a month on annual billing, then $28 and $76 for the higher tiers.

Other tools advertise daily free allowances instead, and those numbers change frequently enough
that any figure quoted on a blog is worth re-checking on the vendor’s own page before you plan
around it. The pattern across free tiers is consistent even when the numbers are not: expect a
lower resolution ceiling, a visible watermark, a shorter maximum clip and a slower queue.

Getting a good result from a still

Image-to-video rewards a different kind of prompt than text-to-video. The model is not inventing
a scene — it already has one — so the prompt should describe movement, not content.

  • Describe the camera, not the subject. “Slow push in, shallow depth of field”
    gives the model a job. “A beautiful mountain landscape” describes what it can already see.
  • Name one motion, not three. Compound instructions across a short clip tend to
    produce a shot that does none of them cleanly.
  • Start from a clean, high-resolution frame. Compression artefacts and heavy
    grain in the source get amplified across every generated frame.
  • Keep clips short and cut between them. Coherence degrades with length on every
    model; four 5-second shots edited together beat one 20-second take.
  • Expect faces and hands to be the failure point. If the still contains either,
    budget more retries than you would for a product or a landscape.

Which should you use?

  • Cheapest workable: Veo 3.1 Lite at 720p or Wan 2.5, both $0.05 per second.
  • Best quality-per-dollar step up: Veo 3.1 Fast at 1080p, $0.12 per second.
  • When output resolution is the requirement: Veo 3.1 Standard at 4K, $0.60 per
    second — the only 4K tier in this group short of Fast at $0.30.
  • Just trying it once: a free browser tier is enough to see whether the idea
    works at all, before any of the above is worth setting up.

Frequently asked questions

Can AI turn any photo into a video?

Technically yes — every model listed here accepts an arbitrary image as the first frame. Results
vary sharply with the source: clean, well-lit, high-resolution stills animate far more reliably
than compressed or heavily textured ones, and images containing faces or hands need more attempts.

Does image-to-video cost more than text-to-video?

No. On all the models above you are billed per second of output regardless of whether the
generation started from a prompt or a picture.

What is the cheapest image-to-video AI?

Veo 3.1 Lite at 720p and Wan 2.5 both cost $0.05 per second of output, which makes a 5-second
clip about $0.25 per generation. Kling 2.5 Turbo Pro is $0.07 per second.

How long can an AI-generated video be?

Practical limits are short. Coherence falls off as clips lengthen on every current model, which
is why most production workflows generate several short shots and edit them together rather than
asking for one long take.

Do free image-to-video tools add a watermark?

Generally yes, alongside a resolution cap and a slower queue. Free allowances are also
promotional and change often — Runway’s, for instance, is a one-time 125-credit deposit rather
than a recurring monthly grant.

Written by Mustafa Ihsan

Mustafa Ihsan is the founder and editor of Convly.ai. He built and maintains the site's live AI models database, its price-performance index, and its free calculators for VRAM requirements, API costs and self-hosting economics. He writes about model pricing, benchmark results and the hardware needed to run AI models locally, and consistently prefers measured numbers to vendor claims.

Scroll to Top
Featured on There's An AI For That