Skip to main content

WaveSpeed adds ByteDance's Seed Audio 1.0 model

By Steven Van ·

The text-to-audio API adds voice cloning, image-guided sound, and dubbing across up to 18 languages.

ByteDance's Seed Audio 1.0 has arrived on WaveSpeedAI. Announced by the platform on June 28, 2026, the launch puts ByteDance's all-in-one audio model behind WaveSpeed's fast inference API alongside 1,000+ other top-tier models. The pitch is simple: type a prompt, tune a few controls, and hear natural speech, sound effects, music, and ambient soundscapes in seconds. WaveSpeed frames it as "Type it. Tune it. Hear it." — and that three-step loop is exactly how the model is exposed on the platform.

Seed Audio 1.0 on WaveSpeed

What Seed Audio 1.0 Brings to WaveSpeed

Seed Audio 1.0 is ByteDance's (Seed team) unified audio-generation model, first shown in late June 2026 and now rolling out across AI platforms including fal, Runware, Runway, Higgsfield — and WaveSpeed. Unlike a traditional text-to-speech engine, it generates the full audio spectrum from a single text prompt: natural human speech, foley-style sound effects, original music, and immersive ambient soundscapes. It also supports multi-speaker dialogue, voice cloning from a reference clip, portrait- and image-guided character voices, and multilingual dubbing across up to 18 languages. On WaveSpeed, that capability set is delivered as a single text-to-audio API endpoint with optional guidance layers.

Preset Voices and Reference-Audio Guidance

WaveSpeed's implementation leans on two complementary voice paths. The first is preset voices: you select a ready-made voice for standard text-to-speech, including multilingual presets that blend English, Chinese, Spanish, Japanese, and Indonesian (the interface lists identifiers such as vivi_mixed_en_zh_ja_es_id). The second is reference-audio guidance — you can upload up to three reference audio clips to steer the voice or overall audio style, which is how you clone a voice or match a specific delivery. Together they cover both quick prototyping with a catalog voice and precise, reference-driven cloning when you need a particular sound.

Image-Guided Audio: Sound From a Picture

The most distinctive control on WaveSpeed is image-guided audio. The API accepts a single reference image URL and uses it to direct the character and mood of the generated audio — letting a portrait inform a character voice, or a scene guide the ambient texture and effects. Paired with reference audio, it gives creators two axes of conditioning: what the voice should sound like, and what the world around it should feel like. For AI video and game workflows, that means a still frame can seed a matching soundscape without hand-authoring every layer.

Fine-Grained Controls and Output Formats

Once the prompt and guidance are set, Seed Audio 1.0 on WaveSpeed exposes granular tuning. Speed and volume are adjustable from 0.5x to 2.0x, and pitch can shift by -12 to +12 semitones for tonal control. Output can be exported as WAV, MP3, PCM, or OGG Opus, at sample rates ranging from 8 kHz up to 48 kHz for high-quality downstream editing. Because it runs on WaveSpeed — an AI media generation platform built for fast image, video, and audio generation at scale via web app, API, and CLI — these audio jobs sit next to the same catalog of 1,000+ models teams already use for the rest of their pipeline, behind one API and one billing account.

Sources: WaveSpeedAI — Seed Audio 1.0, BytePlus (ByteDance) on X, ByteDance Seed — Models.

WaveSpeed
WaveSpeed
AI media generation platform with 1,000+ top-tier models for fast image, video, and audio generation at scale.
View WaveSpeed →

Read WaveSpeed adds ByteDance's Seed Audio 1.0 model on Creators Toolbox