Skip to main content
AI Audio Generator turns a description into finished audio: a narrator over music, two characters talking in a café, a jingle, or a whole scene with voices and sound around them. Where text to speech reads your text in one voice, the audio generator directs the whole scene.

Create audio

1

Open AI Audio Generator

Click Audio > AI Audio Generator in the top menu, or find it in All Tools. The number at the top right shows how many text-to-speech characters you have left.
2

Describe your audio

Write it in Describe your audio, up to 2,048 characters. Set the scene, the voices and how they speak, and the music and sounds. Put spoken words in quotes.Not sure where to start? Pick one under Start with an idea, Narration, Music, Dialogue, or Soundscape, to fill in an example description and length.
AI Audio Generator with Seed Audio 1.0 by ByteDance, a description of a calm yoga teacher saying Welcome to Example Yoga Studio over soft ambient pads and a singing bowl, Add media, Output format MP3, Target duration Auto, and Start with an idea on the right
3

Add references (optional)

Click Add media to guide the result with up to 3 audio clips (MP3 or WAV, up to 30 seconds each) or 1 image (JPG, PNG, or WebP), 10 MB each. You can’t mix audio and an image. Refer to the clips in your description as @Audio1, @Audio2, and @Audio3, for example “Use the voice from @Audio1”.
4

Choose the format and length

  • Output format: MP3, WAV, OGG Opus, or PCM.
  • Target duration: Auto, or 5 seconds to 2 minutes. It’s a guide; the actual length can differ.
5

Click Generate audio

The audio is ready in about 20 seconds and appears under Your creations.
Your creations with one clip from Seed Audio 1.0 as MP3, dated Sep 26, a 16-second player, Download MP3, Reuse prompt, and View prompt
Play each creation in the list, click Download to save it, or Reuse prompt to put its description back in the form and try a variation. View prompt shows the full description. Load more shows older creations.

Advanced settings

Advanced settings: Sample rate 24 kHz, Speed 1x, Volume 1x, and Pitch 0 semitones, with the Generate audio button and 1,000 audio characters per generation
Speed, volume, and pitch mainly shape the voices.

What it costs

Each generation uses 1,000 text-to-speech characters, whatever its length, and only when it succeeds. You need at least 1,000 characters left; otherwise the button changes to Upgrade to generate. The Free plan’s 100 characters aren’t enough. See Audio plans and limits. You can make one clip at a time, together with the sound effect generator.

Text to speech

Read a script aloud in one voice.

Sound effects

Make a single sound effect.