Seed Audio 1.0
ByteDance's Doubao-family text-to-speech model, brought to AutorunX as a premium cloud voice lane for audiobooks and podcasts.
What Seed Audio 1.0 does
Seed Audio 1.0 turns a written script into spoken audio — narration for audiobook chapters or dialogue for podcast episodes — with natural-sounding pacing and emotional delivery rather than a monotone read.
On AutorunX it is scoped to text-to-speech narration inside the Audiobook and Voice Podcast services: paste or import a script, and the model returns finished narration ready to drop into a chapter or episode.
Capabilities
Key features
- Human-like prosody
- Built on ByteDance's Seed-TTS lineage, the model renders natural pacing, stress, and emotional inflection instead of a flat, evenly-timed read, which matters most in longer narration passages.
- Long-form narration
- Designed to hold voice consistency across extended passages, making it suited to full audiobook chapters and multi-minute podcast segments rather than short one-off lines.
- Zero-shot voice conditioning (model-level)
- ByteDance's published Seed-TTS research supports conditioning on a short reference clip to match a target timbre; this is a capability of the underlying model family, worth checking against what AutorunX's current picker exposes.
- Cloud-hosted, no infra to manage
- Runs entirely on ByteDance's own cloud — AutorunX doesn't host it, so there's no local pod or queue contention with owned-engine jobs.
- Premium alternative lane
- Sits next to OmniVoice in the picker as an explicit upgrade path — pick it when a project's narration quality bar justifies the higher per-character cost.
How Seed Audio 1.0 works
Seed Audio 1.0 descends from ByteDance's Seed-TTS research line, which combines a speech tokenizer, an autoregressive language model that predicts speech tokens conditioned on the input text, and a diffusion-based renderer that turns those tokens into a waveform. That split — tokenize, predict, render — is what lets the model reproduce natural prosody, pacing, and emotional inflection rather than reading text in a flat, mechanical cadence.
On AutorunX, the model is reached as a hosted cloud model rather than run on AutorunX's own hardware. A script goes in through the Voice, the request is routed to ByteDance's cloud endpoint, and narrated audio comes back — no local inference, no GPU pod to manage on AutorunX's side.
It sits in the Voice picker as a selectable alternative to OmniVoice, AutorunX's owned default voice engine. Choosing it is a deliberate trade: narration from a large ByteDance-trained cloud model at roughly 6x the per-character cost of the owned option, positioned for jobs where the extra spend is worth it.
What people use Seed Audio 1.0 for
Flagship audiobook chapters
Use it for the opening chapter or a hero sample where narration quality is the first thing a listener judges the whole book by.
Polished podcast episodes
Generate broadcast-ready narration for episodes going out under your brand, where a premium cloud voice is worth the extra credits.
Side-by-side quality checks
Run the same script through Seed Audio 1.0 and OmniVoice to compare delivery before committing a whole book or season to one engine.
Short premium samples
Generate trailers or preview clips where the per-job minimum (50 credits) is easily absorbed by a short, high-visibility piece of audio.
How to use Seed Audio 1.0 on AutorunX
Seed Audio 1.0 is available in Audiobook, in the Voice module.
- 01
Open the service
Go to Voice and open the Audiobook or Voice Podcast editor.
- 02
Set up your input
Paste or import your chapter script, or write out your podcast dialogue.
- 03
Pick Seed Audio 1.0
Open the voice engine picker and select Seed Audio 1.0 — it's an alternative lane, not the default, so you choose it explicitly.
- 04
Generate
Run the job, review the narration, and regenerate individual lines if needed.
- Billed per 1,000 characters generated, from your shared AutorunX credit wallet.
- 200 credits / 1K characters (minimum 50 credits per job)
Tips for better results
- Reserve it for hero content
- At roughly 6x OmniVoice's cost, it makes the most sense for chapters, trailers, or episodes that carry outsized weight — not for bulk narration.
- Watch the 50-credit floor
- Every job is billed a 50-credit minimum regardless of length, so batch very short lines together rather than generating them one at a time.
- A/B against OmniVoice first
- Run a short passage through both engines before committing your budget to a full chapter or episode.
- Punctuate cleanly
- TTS models read pacing and emotional cues from punctuation — clear commas, periods, and paragraph breaks give more natural-sounding output.
Seed Audio 1.0 — frequently asked
Related models
OmniVoice
AutorunX ownedAutorunX's self-hosted voice engine — TTS and zero-shot cloning on our own GPUs.
Suno v5
SunoSuno's song-generation model — full arranged tracks with vocals from a prompt.
Musicful
MusicfulA cloud prompt-to-song API that turns text prompts or lyrics into fully produced tracks, selectable in Song Builder.
Ready to create with Seed Audio 1.0?
Open Audiobook and generate your first voice & speech in a couple of minutes.
Free plan, no card required.