OmniHuman 1.5
ByteDance's audio-driven human video model that animates a photo to speak and move in sync with a given audio track.

Picking OmniHuman 1.5 as the model for a generation.
What OmniHuman 1.5 does
OmniHuman 1.5 turns a static photo and an audio track into a video of that person appearing to speak or perform the audio naturally — with lip sync, facial expression, and body language that respond to the tone and pacing of the voice, not just its phonemes. It's the model to reach for whenever the goal is a talking presenter, avatar, or character rather than a generated scene.
Because output length tracks the audio input rather than a fixed clip duration, it's well suited to voiceover-driven content — narrated explainers, presenter-style videos, or a character singing along to a track — where the video needs to match a pre-existing audio timeline exactly.
Key features
Audio-driven, not prompt-driven
Generation is controlled by an audio track and a reference photo rather than a text prompt describing a scene — the model's job is to make the photo perform that audio convincingly, not to invent visual content.
Duration matches the audio
Clip length scales with the input audio rather than being capped at a fixed short duration, so a full narrated script or song can drive a single continuous generation.
Emotionally responsive motion
The model reads tone, emphasis, and pacing from the audio and reflects that in facial expression and body language, rather than producing flat mouth-shape-matching lip sync.
Standard and high quality tiers
Two quality tiers let you trade render cost and time for facial detail and motion smoothness depending on whether the output is a draft or a final deliverable.
No filming required
A single reference photo and an audio file are the only inputs needed — there's no need to film a presenter, actor, or avatar performance.
How OmniHuman 1.5 works
OmniHuman 1.5 works fundamentally differently from prompt-driven video models like Veo, Kling, or Seedance. Instead of generating a scene from a text description, it takes two required inputs — a single reference photo or avatar image, and an audio track — and animates the person in the photo to speak and move in sync with that audio. Duration isn't set by a prompt; it's driven by the length of the audio track you provide.
Under the hood, ByteDance's approach fuses a multimodal understanding component with a diffusion-based video generator: the model reads semantic and prosodic cues from the audio (tone, emphasis, pacing) and translates them into lip movement, facial expression, and body language, rather than only matching mouth shapes to phonemes. That's what produces motion that reads as emotionally responsive to the audio rather than mechanically lip-synced.
It offers standard and high quality tiers, trading render time and cost for fidelity in facial detail and motion smoothness. Because it only needs a still image and an audio file as source material, it doesn't require filming a presenter or actor at all.
What people use OmniHuman 1.5 for
Music videos
Animate a character or persona to perform along with a generated or uploaded song, with lip and body movement synced to the track.
Narrated explainer videos
Pair a script's voiceover with a presenter photo so the explainer has a consistent, speaking on-screen host without filming anyone.
Brand story presenters
Give a brand a recurring on-screen spokesperson by animating a chosen photo to deliver narrated brand copy naturally.
Multilingual or dubbed content
Swap in an audio track in a different language against the same reference photo to produce a localized version of a presenter video.
Who built OmniHuman 1.5
ByteDance
www.byteplus.comByteDance, the company behind TikTok and Douyin, developed OmniHuman through its digital-human research team as a dedicated audio-driven avatar model, distinct from its general-purpose Seedance video line. It's distributed internationally through BytePlus.
How to use OmniHuman 1.5 on AutorunX
OmniHuman 1.5 is the default model for Music Video in Video Lab.
Open Music Video, Explainer, or Brand Story in Video Lab
From the AutorunX dashboard, go to Video Lab and open Music Video, Explainer, or Brand Story — OmniHuman 1.5 is the default model in all three.
Provide a reference photo and audio
Upload a clear, front-facing photo of the person or avatar you want animated, plus the audio track (voiceover, song, or narration) it should perform.
Confirm OmniHuman 1.5 in the model picker
OmniHuman 1.5 is pre-selected by default for these services. Open the model picker to confirm, or choose the standard or high quality tier.
Generate and review
Hit generate — the output length matches your audio track. Review lip sync and expression, then export.

An example generation rendered with OmniHuman 1.5.
Credit usage
Billed per clip from your shared AutorunX credit wallet.
Tips for better results with OmniHuman 1.5
Start with a clean, front-facing reference photo
A clear, well-lit, front-facing photo gives OmniHuman the most reliable facial landmarks to animate, producing more natural expression than an angled or low-resolution source image.
Use expressive audio for expressive output
Since motion is driven by the audio's tone and pacing, a flat, monotone voiceover will produce flatter facial and body performance than an audio track with natural inflection.
Match quality tier to the deliverable
Use the standard tier for drafts and quick checks, and switch to high quality for the version that actually ships.
Keep audio length intentional
Since output duration tracks the audio track exactly, trim or pace the audio to the length you actually want the video to run before generating.
OmniHuman 1.5 — frequently asked questions
Related models
Seedance 2.0
ByteDance's default AutorunX video engine for Short Film, Movie Maker, and Ad Remake — identity-preserving reference-to-video at up to 4K.
Voice & SpeechSeed Audio 1.0
ByteDance's Doubao-family text-to-speech model, brought to AutorunX as a premium cloud voice lane for audiobooks and podcasts.
Video GenerationKling 3.0
Kuaishou's flagship video model with native multi-lingual audio, in-video editing, and clips up to native 4K.
Ready to create with OmniHuman 1.5?
Sign up for AutorunX to get 200 free credits across every lab, including Music Video.