Bring a reference
Use up to three audio clips to guide a voice or style. Or add an image and describe the sounds of the scene.
A voice that tells your story. Music that sets the mood. A scene that feels alive. Create it all from a few words.
Powered by ByteDance · 1,000 audio characters per generation
Close-up rain pattering on a window, a low distant thunder roll, then rain easing off. No speech or music.
A taste of its sound effects. Your next creation can include voices and music, too.
Pick a starting point, change a few words, and take your idea into the studio.
Give the model a feeling to follow, a voice to reference, or a scene to bring to life. Keep shaping until it fits your project.
Use up to three audio clips to guide a voice or style. Or add an image and describe the sounds of the scene.
Adjust speed, volume, and pitch. Choose a target length and a sample rate from 8 to 48 kHz. The actual duration may vary.
Listen, save, and download. Your creations stay in your history, ready for the next edit, episode, or idea.
Create narration, dialogue, music, sound effects, or a scene that combines them. Describe the voices, mood, setting, and how the audio should unfold. Seed Audio 1.0 is ByteDance’s audio creation model.
Yes. Attach up to three MP3 or WAV audio clips, each up to 30 seconds, or one JPG, PNG, or WebP image. Each file can be up to 10 MB. Use audio references and image guidance separately. Only upload media you have permission to use.
Set a target duration up to 120 seconds, or choose Auto. Duration is prompt guidance, so the actual length may vary. Download MP3, WAV, OGG Opus, or raw PCM. PCM includes a WAV preview for listening.
Each successful generation uses 1,000 characters from your plan’s audio allowance. You can see your balance and the cost before generating. Listening to the demos on this page does not use your allowance.
Text-to-speech focuses on reading a script in a selected voice. This generator lets you describe a whole audio idea, including voices, music, ambience, and effects. Use a prompt to direct the performance and scene.