Lab journal · Talking heads · 18 June 2026

The elusive 3p talking head

A talking head — a presenter who speaks your script — is a useful thing to put in a course. Plenty of models and services make them. Our video app, Platypus, reaches for the higher-end ones — OmniHuman, VEED Fabric — at roughly twelve to thirteen pence per second of video. Broadcast-grade.

But a single course inside a learning platform needs, on average, about five minutes of video. At those rates that is roughly £35 of animation per course — fine for a hero film, not for a catalogue.

So: what are the cheaper alternatives, and how far down can the price go before the result stops being usable? The target we set ourselves — under 5p per clip.

We ran the same character and the same script through everything we could host. Below is each one: the clip, the model, the time it took, the cost. No commentary — your eyes are as good as ours.

Higher end

Broadcast-grade — what a premium video app uses

OmniHuman 1.5
fal · audio-driven

Default settings — a diffusion model that animates the whole head and shoulders.

Render
3.5 min
Cost
~£1.30 / clip

≈ 13p per second of video

VEED Fabric 1.0
fal · 720p

720p output at default settings.

Render
85 s
Cost
~80p / clip

≈ 12p per second

D-ID
own API

Their default pipeline — lips re-synced onto the still.

Render
~7 min
Cost
~10p / clip

plus a monthly subscription

The middle

Sonic and Wan — similar quality, similar price

Sonic
Replicate · dynamic 0.6

dynamic_scale sets how expressive it is — 0.6 here.

Render
151 s
Cost
~17p / clip
Sonic
Replicate · dynamic 0.5

dynamic_scale 0.5 — calmer, and the model's lowest allowed value.

Render
153 s
Cost
~17p / clip
Sonic
Replicate · dynamic 0.6

Same settings, a different character.

Render
179 s
Cost
~20p / clip
Wan 2.2 S2V
Replicate · 832px input

Full-resolution input image — Wan's cost tracks the input size.

Render
135 s
Cost
~11p / clip

Wan, downscaled

The same model on a smaller input image — Wan is billed by input resolution

Wan 2.2 S2V
Replicate · 512px input · tuned

Input downscaled to 512px, with a motion prompt and the largest chunk size.

Render
35 s
Cost
~3p / clip
Wan 2.2 S2V
Replicate · 512px input

Input downscaled to 512px at the default chunk size.

Render
41 s
Cost
~3p / clip

The cheapest

Free, near-free, or a different mechanism entirely

SadTalker
fal · free tier

Full preprocess on a watercolour still.

Render
~8 min
Cost
free *
SadTalker
fal · photo source

The same model on a photo instead of an illustration.

Render
139 s
Cost
free *
JoyVASA
self-hosted · cfg 5.0

cfg_scale drives how much it moves — 5.0, pushed to the top.

Render
18 s
Cost
~1p **
JoyVASA
self-hosted · cfg 3.5

cfg_scale 3.5 — a gentler setting.

Render
23 s
Cost
~1p **

Do we have a solution?

A make-do, perhaps. As of 18 June 2026, on this preliminary investigation, nothing we found is both cheap and convincing. The cost can come down a long way — but the quality comes down with it, in its own way each time.

We have put the most usable of these into our apps as a starting point, so courses can have a talking head today. It is a beginning, not an answer.

The search goes on.

* fal lists SadTalker at ~£0 of compute; to confirm against billing.

** JoyVASA is not available on any hosted API — cost is an estimate for self-hosting.

Costs in pence, approximate, at provider rates June 2026. Replicate figures are render-time × GPU rate (estimated). Subjects and scripts vary between clips.

Talking heads · preliminary investigation · Alt Shift Lab