A talking head — a presenter who speaks your script — is a useful thing to put in a course. Plenty of models and services make them. Our video app, Platypus, reaches for the higher-end ones — OmniHuman, VEED Fabric — at roughly twelve to thirteen pence per second of video. Broadcast-grade.
But a single course inside a learning platform needs, on average, about five minutes of video. At those rates that is roughly £35 of animation per course — fine for a hero film, not for a catalogue.
So: what are the cheaper alternatives, and how far down can the price go before the result stops being usable? The target we set ourselves — under 5p per clip.
We ran the same character and the same script through everything we could host. Below is each one: the clip, the model, the time it took, the cost. No commentary — your eyes are as good as ours.
Higher end
Broadcast-grade — what a premium video app uses
Default settings — a diffusion model that animates the whole head and shoulders.
- Render
- 3.5 min
- Cost
- ~£1.30 / clip
≈ 13p per second of video
720p output at default settings.
- Render
- 85 s
- Cost
- ~80p / clip
≈ 12p per second
Their default pipeline — lips re-synced onto the still.
- Render
- ~7 min
- Cost
- ~10p / clip
plus a monthly subscription
The middle
Sonic and Wan — similar quality, similar price
dynamic_scale sets how expressive it is — 0.6 here.
- Render
- 151 s
- Cost
- ~17p / clip
dynamic_scale 0.5 — calmer, and the model's lowest allowed value.
- Render
- 153 s
- Cost
- ~17p / clip
Same settings, a different character.
- Render
- 179 s
- Cost
- ~20p / clip
Full-resolution input image — Wan's cost tracks the input size.
- Render
- 135 s
- Cost
- ~11p / clip
Wan, downscaled
The same model on a smaller input image — Wan is billed by input resolution
Input downscaled to 512px, with a motion prompt and the largest chunk size.
- Render
- 35 s
- Cost
- ~3p / clip
Input downscaled to 512px at the default chunk size.
- Render
- 41 s
- Cost
- ~3p / clip
The cheapest
Free, near-free, or a different mechanism entirely
Full preprocess on a watercolour still.
- Render
- ~8 min
- Cost
- free *
The same model on a photo instead of an illustration.
- Render
- 139 s
- Cost
- free *
cfg_scale drives how much it moves — 5.0, pushed to the top.
- Render
- 18 s
- Cost
- ~1p **
cfg_scale 3.5 — a gentler setting.
- Render
- 23 s
- Cost
- ~1p **
Do we have a solution?
A make-do, perhaps. As of 18 June 2026, on this preliminary investigation, nothing we found is both cheap and convincing. The cost can come down a long way — but the quality comes down with it, in its own way each time.
We have put the most usable of these into our apps as a starting point, so courses can have a talking head today. It is a beginning, not an answer.
The search goes on.
* fal lists SadTalker at ~£0 of compute; to confirm against billing.
** JoyVASA is not available on any hosted API — cost is an estimate for self-hosting.
Costs in pence, approximate, at provider rates June 2026. Replicate figures are render-time × GPU rate (estimated). Subjects and scripts vary between clips.
Talking heads · preliminary investigation · Alt Shift Lab