Home
Volcano Engine Launches Doubao Audio Model 1.0: One-Sentence Cinematic Audio, Consistent Character Voices for 10 Mins
Yesterday, Volc Engine officially launched the Doubao Audio Generation Model 1.0 (Doubao-Seed-Audio 1.0), which accepts text or audio input to produce complete audio works end-to-end. Its key breakthrough is that a single prompt can handle dialogue, sound effects, and background music together, eliminating the need for manual multi-track editing.

Turn a sentence into an audio director—no post-production needed.
Previously, creating high-quality audio required generating dialogue, sound effects, and music separately, then manually aligning and mixing tracks—a complex process dependent on post-production expertise. Doubao Audio Generation Model 1.0 consolidates all of this into a single prompt: users can define multiple characters' lines, tone, and emotional rhythm in one instruction, embed details like laughter, sighs, pauses, and dialect accents, and generate background music and ambient sound effects simultaneously, outputting a finished product directly. A creator can type a description and immediately receive a podcast, audiobook, or brand audio ready for release.
Consistent character voices throughout long audio without losing quality.
The biggest challenge in long audio creation is consistency—ensuring a character sounds the same at minute one and minute ten. Doubao Audio Generation Model 1.0 deeply integrates text-to-audio with reference audio, maintaining consistent voice quality across long segments, so creators don't have to compare and revise each part. The current model supports up to 2 minutes of audio per generation, and with the extension feature, it preserves voice consistency for longer outputs, suiting audiobooks, podcasts, and long series.
Additionally, the model offers decoupled control of voice and style, letting the same voice adapt to various emotions and contexts—even enabling a single voice to portray multiple roles with different expressions. This greatly enhances flexibility for character voice acting and creative audio production. Currently, Volc Ark has opened API testing, and individual users can get 30 minutes of creation quota in the Experience Center. Doubao Audio Generation Model 1.0 will also be launched on products like CapCut, Jiemod, and Tomato.
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (1)
0/500
Yesterday, Volc Engine officially launched the Doubao Audio Generation Model 1.0 (Doubao-Seed-Audio 1.0), which accepts text or audio input to produce complete audio works end-to-end. Its key breakthrough is that a single prompt can handle dialogue, sound effects, and background music together, eliminating the need for manual multi-track editing.

Turn a sentence into an audio director—no post-production needed.
Previously, creating high-quality audio required generating dialogue, sound effects, and music separately, then manually aligning and mixing tracks—a complex process dependent on post-production expertise. Doubao Audio Generation Model 1.0 consolidates all of this into a single prompt: users can define multiple characters' lines, tone, and emotional rhythm in one instruction, embed details like laughter, sighs, pauses, and dialect accents, and generate background music and ambient sound effects simultaneously, outputting a finished product directly. A creator can type a description and immediately receive a podcast, audiobook, or brand audio ready for release.
Consistent character voices throughout long audio without losing quality.
The biggest challenge in long audio creation is consistency—ensuring a character sounds the same at minute one and minute ten. Doubao Audio Generation Model 1.0 deeply integrates text-to-audio with reference audio, maintaining consistent voice quality across long segments, so creators don't have to compare and revise each part. The current model supports up to 2 minutes of audio per generation, and with the extension feature, it preserves voice consistency for longer outputs, suiting audiobooks, podcasts, and long series.
Additionally, the model offers decoupled control of voice and style, letting the same voice adapt to various emotions and contexts—even enabling a single voice to portray multiple roles with different expressions. This greatly enhances flexibility for character voice acting and creative audio production. Currently, Volc Ark has opened API testing, and individual users can get 30 minutes of creation quota in the Experience Center. Doubao Audio Generation Model 1.0 will also be launched on products like CapCut, Jiemod, and Tomato.
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur











