Voice & SpeechByteDance (Doubao / Seed team)

Seed Audio 1.0

ByteDance's Doubao-family text-to-speech model, brought to AutorunX as a premium cloud voice lane for audiobooks and podcasts.

Cr 200/1k char

What Seed Audio 1.0 does

Seed Audio 1.0 turns a written script into spoken audio — narration for audiobook chapters or dialogue for podcast episodes — with natural-sounding pacing and emotional delivery rather than a monotone read.

On AutorunX it is scoped to text-to-speech narration inside the Audiobook and Voice Podcast services: paste or import a script, and the model returns finished narration ready to drop into a chapter or episode.

Capabilities

Key features

Human-like prosody
Built on ByteDance's Seed-TTS lineage, the model renders natural pacing, stress, and emotional inflection instead of a flat, evenly-timed read, which matters most in longer narration passages.
Long-form narration
Designed to hold voice consistency across extended passages, making it suited to full audiobook chapters and multi-minute podcast segments rather than short one-off lines.
Zero-shot voice conditioning (model-level)
ByteDance's published Seed-TTS research supports conditioning on a short reference clip to match a target timbre; this is a capability of the underlying model family, worth checking against what AutorunX's current picker exposes.
Cloud-hosted, no infra to manage
Runs entirely on ByteDance's own cloud — AutorunX doesn't host it, so there's no local pod or queue contention with owned-engine jobs.
Premium alternative lane
Sits next to OmniVoice in the picker as an explicit upgrade path — pick it when a project's narration quality bar justifies the higher per-character cost.

How Seed Audio 1.0 works

Seed Audio 1.0 descends from ByteDance's Seed-TTS research line, which combines a speech tokenizer, an autoregressive language model that predicts speech tokens conditioned on the input text, and a diffusion-based renderer that turns those tokens into a waveform. That split — tokenize, predict, render — is what lets the model reproduce natural prosody, pacing, and emotional inflection rather than reading text in a flat, mechanical cadence.

On AutorunX, the model is reached as a hosted cloud model rather than run on AutorunX's own hardware. A script goes in through the Voice, the request is routed to ByteDance's cloud endpoint, and narrated audio comes back — no local inference, no GPU pod to manage on AutorunX's side.

It sits in the Voice picker as a selectable alternative to OmniVoice, AutorunX's owned default voice engine. Choosing it is a deliberate trade: narration from a large ByteDance-trained cloud model at roughly 6x the per-character cost of the owned option, positioned for jobs where the extra spend is worth it.

What people use Seed Audio 1.0 for

Flagship audiobook chapters

Use it for the opening chapter or a hero sample where narration quality is the first thing a listener judges the whole book by.

Polished podcast episodes

Generate broadcast-ready narration for episodes going out under your brand, where a premium cloud voice is worth the extra credits.

Side-by-side quality checks

Run the same script through Seed Audio 1.0 and OmniVoice to compare delivery before committing a whole book or season to one engine.

Short premium samples

Generate trailers or preview clips where the per-job minimum (50 credits) is easily absorbed by a short, high-visibility piece of audio.

How to use Seed Audio 1.0 on AutorunX

Seed Audio 1.0 is available in Audiobook, in the Voice module.

  1. 01

    Open the service

    Go to Voice and open the Audiobook or Voice Podcast editor.

  2. 02

    Set up your input

    Paste or import your chapter script, or write out your podcast dialogue.

  3. 03

    Pick Seed Audio 1.0

    Open the voice engine picker and select Seed Audio 1.0 — it's an alternative lane, not the default, so you choose it explicitly.

  4. 04

    Generate

    Run the job, review the narration, and regenerate individual lines if needed.

Billed per 1,000 characters generated, from your shared AutorunX credit wallet.
200 credits / 1K characters (minimum 50 credits per job)

Tips for better results

Reserve it for hero content
At roughly 6x OmniVoice's cost, it makes the most sense for chapters, trailers, or episodes that carry outsized weight — not for bulk narration.
Watch the 50-credit floor
Every job is billed a 50-credit minimum regardless of length, so batch very short lines together rather than generating them one at a time.
A/B against OmniVoice first
Run a short passage through both engines before committing your budget to a full chapter or episode.
Punctuate cleanly
TTS models read pacing and emotional cues from punctuation — clear commas, periods, and paragraph breaks give more natural-sounding output.

Seed Audio 1.0 — frequently asked

Ready to create with Seed Audio 1.0?

Open Audiobook and generate your first voice & speech in a couple of minutes.

Free plan, no card required.