Voice & SpeechCloud · ByteDance (Doubao / Seed team)

Seed Audio 1.0

ByteDance's Doubao-family text-to-speech model, brought to AutorunX as a premium cloud voice lane for audiobooks and podcasts.

Selecting Seed Audio 1.0 in the AutorunX model picker

Picking Seed Audio 1.0 as the model for a generation.

DeveloperByteDance (Doubao / Seed team)
Model typeCloud text-to-speech (TTS)
Access on AutorunXVia EvoLink cloud API
TrackCloud — selectable alternative, not the default voice engine
Cost vs. defaultRoughly 6x the cost of AutorunX's owned OmniVoice engine
Recently addedReplaced the older GPT-SoVITS cloud voice lane in the picker
Billing unitPer 1,000 characters, 30-credit minimum per job

What Seed Audio 1.0 does

Seed Audio 1.0 turns a written script into spoken audio — narration for audiobook chapters or dialogue for podcast episodes — with natural-sounding pacing and emotional delivery rather than a monotone read.

On AutorunX it is scoped to text-to-speech narration inside the Audiobook and Voice Podcast services: paste or import a script, and the model returns finished narration ready to drop into a chapter or episode.

Key features

Human-like prosody

Built on ByteDance's Seed-TTS lineage, the model renders natural pacing, stress, and emotional inflection instead of a flat, evenly-timed read, which matters most in longer narration passages.

Long-form narration

Designed to hold voice consistency across extended passages, making it suited to full audiobook chapters and multi-minute podcast segments rather than short one-off lines.

Zero-shot voice conditioning (model-level)

ByteDance's published Seed-TTS research supports conditioning on a short reference clip to match a target timbre; this is a capability of the underlying model family, worth checking against what AutorunX's current picker exposes.

Cloud-hosted, no infra to manage

Runs entirely on ByteDance's cloud via the EvoLink integration — AutorunX doesn't host it, so there's no local pod or queue contention with owned-engine jobs.

Premium alternative lane

Sits next to OmniVoice in the picker as an explicit upgrade path — pick it when a project's narration quality bar justifies the higher per-character cost.

How Seed Audio 1.0 works

Seed Audio 1.0 descends from ByteDance's Seed-TTS research line, which combines a speech tokenizer, an autoregressive language model that predicts speech tokens conditioned on the input text, and a diffusion-based renderer that turns those tokens into a waveform. That split — tokenize, predict, render — is what lets the model reproduce natural prosody, pacing, and emotional inflection rather than reading text in a flat, mechanical cadence.

On AutorunX, the model is reached through the EvoLink cloud integration rather than run on AutorunX's own hardware. A script goes in through the Voice Lab, the request is routed to ByteDance's cloud endpoint, and narrated audio comes back — no local inference, no GPU pod to manage on AutorunX's side.

It sits in the Voice Lab picker as a selectable alternative to OmniVoice, AutorunX's owned default voice engine. Choosing it is a deliberate trade: narration from a large ByteDance-trained cloud model at roughly 6x the per-character cost of the owned option, positioned for jobs where the extra spend is worth it.

What people use Seed Audio 1.0 for

Flagship audiobook chapters

Use it for the opening chapter or a hero sample where narration quality is the first thing a listener judges the whole book by.

Polished podcast episodes

Generate broadcast-ready narration for episodes going out under your brand, where a premium cloud voice is worth the extra credits.

Side-by-side quality checks

Run the same script through Seed Audio 1.0 and OmniVoice to compare delivery before committing a whole book or season to one engine.

Short premium samples

Generate trailers or preview clips where the per-job minimum (30 credits) is easily absorbed by a short, high-visibility piece of audio.

Who built Seed Audio 1.0

ByteDance (Doubao / Seed team)

seed.bytedance.com

ByteDance is the technology company behind TikTok, CapCut, and the Doubao AI assistant. Its Seed research team builds the Seed-TTS and Seed Audio family of speech models, extending ByteDance's earlier work on large-scale, human-like voice generation into full-scene audio production used across CapCut, Jimeng, and Fanqie.

How to use Seed Audio 1.0 on AutorunX

Seed Audio 1.0 is available in Audiobook in Voice Lab.

1

Open the service

Go to Voice Lab and open the Audiobook or Voice Podcast editor.

2

Set up your input

Paste or import your chapter script, or write out your podcast dialogue.

3

Pick Seed Audio 1.0

Open the voice engine picker and select Seed Audio 1.0 — it's an alternative lane, not the default, so you choose it explicitly.

4

Generate

Run the job, review the narration, and regenerate individual lines if needed.

An example generation rendered with Seed Audio 1.0 on AutorunX

An example generation rendered with Seed Audio 1.0.

Credit usage

Billed per 1,000 characters generated, from your shared AutorunX credit wallet.

120 credits / 1K characters (minimum 30 credits per job)

Tips for better results with Seed Audio 1.0

Reserve it for hero content

At roughly 6x OmniVoice's cost, it makes the most sense for chapters, trailers, or episodes that carry outsized weight — not for bulk narration.

Watch the 30-credit floor

Every job is billed a 30-credit minimum regardless of length, so batch very short lines together rather than generating them one at a time.

A/B against OmniVoice first

Run a short passage through both engines before committing your budget to a full chapter or episode.

Punctuate cleanly

TTS models read pacing and emotional cues from punctuation — clear commas, periods, and paragraph breaks give more natural-sounding output.

Seed Audio 1.0 — frequently asked questions

Ready to create with Seed Audio 1.0?

Sign up for AutorunX to get 200 free credits across every lab, including Audiobook.