Pick music, speech or effects
Those are different entries under MiniMax Audio, not modes of the same one. Pick the job you actually want so the form shows the fields that entry accepts.
AI Audio
MiniMax Audio is an AI audio generator for music, speech (TTS) and short sound effects. Describe a track, paste a script for voice, or name one event for an effect. Length and voice are set before you generate.
Those are different entries under MiniMax Audio, not modes of the same one. Pick the job you actually want so the form shows the fields that entry accepts.
For music, write genre, tempo and mood. For speech, paste the script. For an effect, describe one event and the room it happens in.
Length and voice are chosen before the run. If a take is close, change the wording and run it again rather than starting a new description.
Genre, tempo and mood do more work than instrument lists. Say what the track is for. Some entries under MiniMax Audio write lyrics for you; custom mode is for when the words have to be specific.
Paste the script and pick a voice. This is TTS: the model reads what you typed. A separate voice-clone entry, when listed, copies a short sample rather than inventing a new speaker.
Effects are short single sounds, not songs. Describe the source and the space it happens in, and keep each request to one event.
MiniMax Audio is an AI audio model. Depending on the named entry, it writes music from a description, reads a script as speech (TTS), or makes a short sound effect.
That depends on the named entry, not on the vendor name alone. One entry writes a track, another reads a script, another makes a short sound. Pick the job you want.
Only if you pick a custom or lyrics entry. Inspiration-style music entries write the words from a mood and a genre. Use custom when the words have to be specific.
Music is a full track with structure. A sound effect is one short event: a knock, a whoosh, a room tone. Do not ask an effects entry for a song.