Skip to main content
ZeroTwo’s audio Studio provides access to multiple AI audio generation models, each optimized for different types of audio output. This page explains the model types available and how to choose between them.
AI audio generation technology is evolving rapidly. New models are added to ZeroTwo regularly. Check the model dropdown in the audio Studio for the current full list.

Model categories

Audio generation models in ZeroTwo fall into three main categories:

Music generation models

Specialized for generating original music from text descriptions. These models understand genre, mood, instrumentation, tempo, and musical structure. Best for:
  • Background music for videos, presentations, and apps
  • Ambient soundscapes and atmospheric audio
  • Jingles, intros, and branded audio pieces
  • Specific genre requests (jazz, classical, electronic, etc.)
Prompt approach: Describe genre, mood, instruments, tempo, and duration. Example: "Upbeat electronic background music, 120 BPM, synthesizer melody, suitable for a tech product demo, 60 seconds"

Voice synthesis / text-to-speech models

Generate spoken audio from text input. These models produce natural-sounding narration in various voices and styles. Best for:
  • AI narration for videos and presentations
  • Podcast-style spoken content
  • Accessibility audio (screen reader-style narration)
  • Character voices for creative projects
Prompt approach: Provide the text to be spoken and describe the voice characteristics — tone, pace, gender, accent, emotional register. Example: "Narrate the following in a warm, professional female voice at a moderate pace: [text]"

Sound effects models

Generate specific, discrete audio events — clicks, chimes, environment sounds, and other effects. Best for:
  • UI sounds and notification tones
  • Environmental and ambient effects
  • Production sound design
  • Game audio assets
Prompt approach: Describe the specific sound event as precisely as possible. Example: "A single wooden door knock, two knocks, natural reverb, interior setting"

Choosing the right model

For music, describing the genre and mood is the most important part of the prompt. For voice, the most important elements are the text content and the voice tone/style description. For sound effects, precision about the specific sound event produces the best results.

Output formats

Download format options depend on the model selected. MP3 is available from all models; WAV is available from higher-quality models.

Prompting by model type

Each audio model type responds to different prompt elements:

Prompting music generation models

The most important elements for music prompts are genre, mood, and instrumentation: Strong music prompt: Cinematic orchestral piece with rising strings and dramatic percussion, building tension over 30 seconds, suitable for a movie trailer

Prompting voice synthesis models

Voice synthesis prompts focus on the text to be spoken and the voice characteristics: Strong voice prompt: Read the following in a warm, professional female voice at a conversational pace, with natural pauses: [your text here]

Prompting sound effects models

Sound effect prompts should be as specific as possible about the exact sound event: Strong sound effect prompt: A single metallic coin dropped onto a hardwood floor, brief ring and roll, indoor environment with slight room reverb

Model updates

ZeroTwo’s audio model library is updated as new models become available. Check the ZeroTwo changelog for announcements about newly added audio models.

Creating audio

Step-by-step guide and prompt examples for all audio types.

Audio troubleshooting

Fix common issues with audio generation.

Studio overview

Overview of all three Studio sections — images, video, and audio.

Video generation

Generate AI video clips from text descriptions.