OmniHuman 1.5
ByteDance's audio-driven human video model that animates a photo to speak and move in sync with a given audio track.
What OmniHuman 1.5 does
OmniHuman 1.5 turns a static photo and an audio track into a video of that person appearing to speak or perform the audio naturally — with lip sync, facial expression, and body language that respond to the tone and pacing of the voice, not just its phonemes. It's the model to reach for whenever the goal is a talking presenter, avatar, or character rather than a generated scene.
Because output length tracks the audio input rather than a fixed clip duration, it's well suited to voiceover-driven content — narrated explainers, presenter-style videos, or a character singing along to a track — where the video needs to match a pre-existing audio timeline exactly.
Capabilities
Key features
- Audio-driven, not prompt-driven
- Generation is controlled by an audio track and a reference photo rather than a text prompt describing a scene — the model's job is to make the photo perform that audio convincingly, not to invent visual content.
- Duration matches the audio
- Clip length scales with the input audio rather than being capped at a fixed short duration, so a full narrated script or song can drive a single continuous generation.
- Emotionally responsive motion
- The model reads tone, emphasis, and pacing from the audio and reflects that in facial expression and body language, rather than producing flat mouth-shape-matching lip sync.
- Standard and high quality tiers
- Two quality tiers let you trade render cost and time for facial detail and motion smoothness depending on whether the output is a draft or a final deliverable.
- No filming required
- A single reference photo and an audio file are the only inputs needed — there's no need to film a presenter, actor, or avatar performance.
How OmniHuman 1.5 works
OmniHuman 1.5 works fundamentally differently from prompt-driven video models like Veo, Kling, or Seedance. Instead of generating a scene from a text description, it takes two required inputs — a single reference photo or avatar image, and an audio track — and animates the person in the photo to speak and move in sync with that audio. Duration isn't set by a prompt; it's driven by the length of the audio track you provide.
Under the hood, ByteDance's approach fuses a multimodal understanding component with a diffusion-based video generator: the model reads semantic and prosodic cues from the audio (tone, emphasis, pacing) and translates them into lip movement, facial expression, and body language, rather than only matching mouth shapes to phonemes. That's what produces motion that reads as emotionally responsive to the audio rather than mechanically lip-synced.
It offers standard and high quality tiers, trading render time and cost for fidelity in facial detail and motion smoothness. Because it only needs a still image and an audio file as source material, it doesn't require filming a presenter or actor at all.
What people use OmniHuman 1.5 for
Music videos
Animate a character or persona to perform along with a generated or uploaded song, with lip and body movement synced to the track.
Narrated explainer videos
Pair a script's voiceover with a presenter photo so the explainer has a consistent, speaking on-screen host without filming anyone.
Brand story presenters
Give a brand a recurring on-screen spokesperson by animating a chosen photo to deliver narrated brand copy naturally.
Multilingual or dubbed content
Swap in an audio track in a different language against the same reference photo to produce a localized version of a presenter video.
How to use OmniHuman 1.5 on AutorunX
OmniHuman 1.5 is the default model for Music Video, in the Video module.
- 01
Open Music Video, Explainer, or Brand Story in Video
From the AutorunX dashboard, go to Video and open Music Video, Explainer, or Brand Story — OmniHuman 1.5 is the default model in all three.
- 02
Provide a reference photo and audio
Upload a clear, front-facing photo of the person or avatar you want animated, plus the audio track (voiceover, song, or narration) it should perform.
- 03
Confirm OmniHuman 1.5 in the model picker
OmniHuman 1.5 is pre-selected by default for these services. Open the model picker to confirm, or choose the standard or high quality tier.
- 04
Generate and review
Hit generate — the output length matches your audio track. Review lip sync and expression, then export.
- Billed per second of generated video from your shared AutorunX credit wallet.
- 168 credits / sec (Music Video/Explainer/Brand Story default)
Tips for better results
- Start with a clean, front-facing reference photo
- A clear, well-lit, front-facing photo gives OmniHuman the most reliable facial landmarks to animate, producing more natural expression than an angled or low-resolution source image.
- Use expressive audio for expressive output
- Since motion is driven by the audio's tone and pacing, a flat, monotone voiceover will produce flatter facial and body performance than an audio track with natural inflection.
- Match quality tier to the deliverable
- Use the standard tier for drafts and quick checks, and switch to high quality for the version that actually ships.
- Keep audio length intentional
- Since output duration tracks the audio track exactly, trim or pace the audio to the length you actually want the video to run before generating.
OmniHuman 1.5 — frequently asked
Related models
Seedance 2.0
ByteDance / BytePlusByteDance's default AutorunX video engine for Short Film, Movie Maker, and Ad Remake — identity-preserving reference-to-video at up to 4K.
Seed Audio 1.0
ByteDance (Doubao / Seed team)ByteDance's Doubao-family text-to-speech model, brought to AutorunX as a premium cloud voice lane for audiobooks and podcasts.
Kling 3.0
Kuaishou TechnologyKuaishou's flagship video model with native multi-lingual audio, in-video editing, and clips up to native 4K.
Ready to create with OmniHuman 1.5?
Open Music Video and generate your first video in a couple of minutes.
Free plan, no card required.