Video GenerationCloud · ByteDance

OmniHuman 1.5

ByteDance's audio-driven human video model that animates a photo to speak and move in sync with a given audio track.

Selecting OmniHuman 1.5 in the AutorunX model picker

Picking OmniHuman 1.5 as the model for a generation.

DeveloperByteDance
CategoryAudio-driven human video generation (lip sync)
Clip lengthScales with input audio length
InputOne photo/avatar image + one audio track
AccessProprietary, closed API via EvoLink
On AutorunXDefault model for Music Video, Explainer, and Brand Story
Quality tiersStandard and High

What OmniHuman 1.5 does

OmniHuman 1.5 turns a static photo and an audio track into a video of that person appearing to speak or perform the audio naturally — with lip sync, facial expression, and body language that respond to the tone and pacing of the voice, not just its phonemes. It's the model to reach for whenever the goal is a talking presenter, avatar, or character rather than a generated scene.

Because output length tracks the audio input rather than a fixed clip duration, it's well suited to voiceover-driven content — narrated explainers, presenter-style videos, or a character singing along to a track — where the video needs to match a pre-existing audio timeline exactly.

Key features

Audio-driven, not prompt-driven

Generation is controlled by an audio track and a reference photo rather than a text prompt describing a scene — the model's job is to make the photo perform that audio convincingly, not to invent visual content.

Duration matches the audio

Clip length scales with the input audio rather than being capped at a fixed short duration, so a full narrated script or song can drive a single continuous generation.

Emotionally responsive motion

The model reads tone, emphasis, and pacing from the audio and reflects that in facial expression and body language, rather than producing flat mouth-shape-matching lip sync.

Standard and high quality tiers

Two quality tiers let you trade render cost and time for facial detail and motion smoothness depending on whether the output is a draft or a final deliverable.

No filming required

A single reference photo and an audio file are the only inputs needed — there's no need to film a presenter, actor, or avatar performance.

How OmniHuman 1.5 works

OmniHuman 1.5 works fundamentally differently from prompt-driven video models like Veo, Kling, or Seedance. Instead of generating a scene from a text description, it takes two required inputs — a single reference photo or avatar image, and an audio track — and animates the person in the photo to speak and move in sync with that audio. Duration isn't set by a prompt; it's driven by the length of the audio track you provide.

Under the hood, ByteDance's approach fuses a multimodal understanding component with a diffusion-based video generator: the model reads semantic and prosodic cues from the audio (tone, emphasis, pacing) and translates them into lip movement, facial expression, and body language, rather than only matching mouth shapes to phonemes. That's what produces motion that reads as emotionally responsive to the audio rather than mechanically lip-synced.

It offers standard and high quality tiers, trading render time and cost for fidelity in facial detail and motion smoothness. Because it only needs a still image and an audio file as source material, it doesn't require filming a presenter or actor at all.

What people use OmniHuman 1.5 for

Music videos

Animate a character or persona to perform along with a generated or uploaded song, with lip and body movement synced to the track.

Narrated explainer videos

Pair a script's voiceover with a presenter photo so the explainer has a consistent, speaking on-screen host without filming anyone.

Brand story presenters

Give a brand a recurring on-screen spokesperson by animating a chosen photo to deliver narrated brand copy naturally.

Multilingual or dubbed content

Swap in an audio track in a different language against the same reference photo to produce a localized version of a presenter video.

Who built OmniHuman 1.5

ByteDance, the company behind TikTok and Douyin, developed OmniHuman through its digital-human research team as a dedicated audio-driven avatar model, distinct from its general-purpose Seedance video line. It's distributed internationally through BytePlus.

How to use OmniHuman 1.5 on AutorunX

OmniHuman 1.5 is the default model for Music Video in Video Lab.

1

Open Music Video, Explainer, or Brand Story in Video Lab

From the AutorunX dashboard, go to Video Lab and open Music Video, Explainer, or Brand Story — OmniHuman 1.5 is the default model in all three.

2

Provide a reference photo and audio

Upload a clear, front-facing photo of the person or avatar you want animated, plus the audio track (voiceover, song, or narration) it should perform.

3

Confirm OmniHuman 1.5 in the model picker

OmniHuman 1.5 is pre-selected by default for these services. Open the model picker to confirm, or choose the standard or high quality tier.

4

Generate and review

Hit generate — the output length matches your audio track. Review lip sync and expression, then export.

An example generation rendered with OmniHuman 1.5 on AutorunX

An example generation rendered with OmniHuman 1.5.

Credit usage

Billed per clip from your shared AutorunX credit wallet.

101 credits / clip (Music Video/Explainer/Brand Story default)

Tips for better results with OmniHuman 1.5

Start with a clean, front-facing reference photo

A clear, well-lit, front-facing photo gives OmniHuman the most reliable facial landmarks to animate, producing more natural expression than an angled or low-resolution source image.

Use expressive audio for expressive output

Since motion is driven by the audio's tone and pacing, a flat, monotone voiceover will produce flatter facial and body performance than an audio track with natural inflection.

Match quality tier to the deliverable

Use the standard tier for drafts and quick checks, and switch to high quality for the version that actually ships.

Keep audio length intentional

Since output duration tracks the audio track exactly, trim or pace the audio to the length you actually want the video to run before generating.

OmniHuman 1.5 — frequently asked questions

Ready to create with OmniHuman 1.5?

Sign up for AutorunX to get 200 free credits across every lab, including Music Video.