BytePlus unveils Seed Audio 1.0, a single-pass audio model
By Steven Van ·
The model generates voice, music and sound effects in one pass and is now open for enterprise access, with early integrations via Runway and Higgsfield.
BytePlus, ByteDance's AI-native cloud platform, has unveiled Seed Audio 1.0 and opened it for enterprise access applications. Built by ByteDance's Seed team, it is billed as a pioneering non-streaming text-to-speech model that generates voice, music, and sound effects together in a single pass. This is the source announcement from the model's maker — the authoritative post about the model itself, ahead of the integrations now rolling out across the wider AI ecosystem.
Why a single-pass, non-streaming architecture matters
Most production TTS systems are streaming and multi-stage: they emit audio chunk by chunk and often stitch speech, music, and effects together from separate pipelines. Seed Audio 1.0 takes the opposite approach. As a non-streaming model, it generates a complete audio composition in one pass, jointly producing spoken voice, background music, and sound effects rather than layering them in post-production.
The practical payoff is coherence. When speech, score, and ambient sound are planned in the same generation, timing, tone, and acoustic space line up naturally — a narrator's cadence sits inside a soundscape that was built to fit it. That single-pass design trades the low latency of streaming for higher-fidelity, holistically composed output, which is exactly what enterprise media and video workflows tend to want.
The four core capabilities: speech, presets, reference audio, image guidance
BytePlus highlights four headline capabilities. The first is natural speech generation — expressive, human-sounding narration and dialogue from text. The second is a library of preset voices, giving teams ready-made characters without any setup or cloning.
The third capability is reference audio guidance: supply a short reference clip and the model conditions its output on that voice or style, enabling zero-shot personalization from a small sample. The fourth is the most unusual: image-guided audio. Feed the model an image and it can shape the generated sound to match the scene, so a picture of a rainy street or a crowded stadium can steer the mood, effects, and ambience of the audio it produces. Together these controls span everything from clean voiceover to full cinematic soundscapes.
Enterprise access on BytePlus and the Volcano Engine stack
Seed Audio 1.0 is now open for enterprise access application through BytePlus, ByteDance's global cloud platform and the international counterpart to Volcano Engine. BytePlus already hosts a deep catalog of ByteDance's frontier models — the Seedance video family, Seedream image generation, large language models via ModelArk, and recommendation engines — behind enterprise-grade APIs.
Positioning Seed Audio 1.0 inside that stack matters for buyers: an organization already generating video with Seedance or images with Seedream can add jointly composed audio from the same provider, under one set of contracts, quotas, and infrastructure. The enterprise application process signals a controlled, business-first rollout rather than an open consumer launch, aimed at teams building media, gaming, advertising, and localization products at scale.
How Seed Audio 1.0 is rolling out across the AI ecosystem
Beyond BytePlus itself, Seed Audio 1.0 is spreading fast across the platforms creators and developers already use. Serverless inference providers are exposing it through their APIs, and creative suites are wiring it into end-user workflows for dialogue, music, and effects.
Early hosts include fal and WaveSpeed, which offer developer API access to the model, alongside Runware. On the creative side, Runway has added Seed Audio 1.0 to its paid plans so users can generate complete audio scenes — dialogue, music, and sound effects — from a text prompt, and Higgsfield is bringing it into its generation stack. Each of these is an integration that hosts the model; BytePlus remains the authoritative source and the enterprise access point for the model itself.
Sources: BytePlus announcement on X, Seed Audio 1.0 on fal, AlphaSignal: Runway brings Seed Audio 1.0.
