OmniVoice
AutorunX's self-hosted voice engine — TTS and zero-shot cloning on our own GPUs.

Picking OmniVoice as the engine for a voice generation.
What OmniVoice does
OmniVoice converts written text into natural speech, and can clone a specific voice from a short reference recording so that voice can narrate any script going forward.
It's the workhorse behind most of AutorunX's Voice Lab: audiobooks, podcasts, story narration, course narration, and greetings all run on OmniVoice by default, with a cloud model available as an alternative where needed.
Key features
Zero-shot voice cloning
Clone a voice from a single 10-60 second reference clip — no lengthy training run required. The cloned voice can then narrate any new script.
Multi-speaker scripts
Inline speaker tags let one generation produce a conversation between multiple distinct voices, which is what powers AutorunX's multi-host Voice Podcast service.
Pause and duration control
Direct control over pacing and pauses matters for long-form narration — audiobook chapters and course lessons need natural breathing room that a purely automatic TTS pass often gets wrong.
Self-hosted, not per-call billed
Because OmniVoice runs on AutorunX's own GPUs instead of a metered third-party API, it's the platform's lowest-cost voice option, and the one every other voice model is priced relative to.
How OmniVoice works
OmniVoice is a text-to-speech and voice-cloning engine that AutorunX runs itself on its own GPU fleet, rather than routing requests to a third-party API. Feed it a script and it synthesizes natural-sounding speech; feed it a short reference clip and it can clone that voice from as little as 10-60 seconds of audio, without a lengthy training process.
It supports multi-speaker scripts using inline speaker tags, so a single generation can produce a conversation between distinct voices — the basis for AutorunX's Voice Podcast service. It also gives direct control over pauses and pacing, which matters for narration-heavy content like audiobooks and courses.
Because it's self-hosted rather than billed per API call, OmniVoice is the cost baseline AutorunX builds its voice pricing around, and it's tuned specifically for the platform's own voice services rather than being a general-purpose third-party model.
What people use OmniVoice for
Audiobook and story narration
Narrate a full manuscript or short story with one consistent voice, with pacing and pause control tuned for long-form listening.
Multi-host podcasts
Generate a conversation between two or more distinct voices from a single script using multi-speaker tags, for Voice Podcast-style content.
Custom voice cloning
Clone your own voice, or any voice you have rights to, from a short reference clip, and reuse it across every future script.
Course and greeting narration
Bulk-narrate course lessons or generate personalized greetings with merge-field names, drawing on the same underlying engine.
Who built OmniVoice
AutorunX
autorunx.comOmniVoice is AutorunX's own self-hosted voice engine — run on AutorunX's own GPU fleet rather than a third-party API. It's the owned counterpart to the cloud voice models on the platform, built to keep the core voice lane fast and cost-efficient at scale.
How to use OmniVoice on AutorunX
OmniVoice is the default model for Audiobook in Voice Lab.
Open a Voice Lab service
From the AutorunX dashboard, go to Voice Lab and pick a service — Audiobook, Podcast, Story Narrator, Course Pro, or Voice Greeting.
Pick or clone a voice
Choose an existing voice, or clone your own from a short reference clip in the voice picker's clone drawer.
Confirm OmniVoice is selected
OmniVoice is the default engine across voice services. Open the model picker to confirm, or compare against the cloud voice option by rate.
Write your script and generate
Paste or write your script, adjust pacing if needed, and generate. Multi-speaker scripts use inline speaker tags for a conversation.
A real AutorunX generation — narration rendered with OmniVoice.
Credit usage
Billed per 1,000 characters synthesized from your shared AutorunX credit wallet.
Tips for better results with OmniVoice
Use a clean reference clip for cloning
10-60 seconds of clear speech with no background noise or music clones far more reliably than a noisy or short sample.
Use speaker tags for multi-voice scripts
For podcast-style content, tag each speaker's lines explicitly rather than relying on the model to infer who's talking.
Control pacing for long-form narration
For audiobook chapters, use pause and duration controls rather than letting the default pacing run — natural breathing room reads much better over a full chapter.
Only clone voices you have rights to
Voice cloning must have the speaker's consent, whether it's your own voice or someone else's — this applies across every voice service on AutorunX.
OmniVoice — frequently asked questions
Related models
Seed Audio 1.0
ByteDance's Doubao-family text-to-speech model, brought to AutorunX as a premium cloud voice lane for audiobooks and podcasts.
Music GenerationSuno v5
Suno's song-generation model — full arranged tracks with vocals from a prompt.
Music GenerationACE-Step
AutorunX's owned, self-hosted open-source music engine — Apache 2.0 licensed and the cheapest song-generation option available.
Ready to create with OmniVoice?
Sign up for AutorunX to get 200 free credits across every lab, including Audiobook.