Today, we are introducing MiniMax Music 3.0, our next-generation music generation model. Given a creative concept and optional lyrics, the model composes, arranges, performs, and produces a complete song in a single generation.
Music 3.0 focuses on the aspects of music creation that are hardest to capture with a simple prompt: understanding the creator’s expressive intent; sustaining that intent across a complete song of up to five minutes; rendering instruments with clarity and physical realism; and generating vocals that sound performed rather than synthesized.
To achieve this, we redesigned the entire generation pipeline—from music description and language modeling to audio rendering. Fine-grained temporal descriptions capture the evolution of emotion, instrumentation, and vocal delivery. A global–local Hybrid-LM jointly models long-range structure and acoustic detail. Multi-layer residual vector quantization (RVQ), hidden-state fusion, flow matching, and a Flow-VAE further improve output fidelity. Together, these advances deliver three major upgrades: more accurate interpretation of creative intent, more complete and varied arrangements, and clearer, more natural sound.
Music 3.0: Technical Architecture
Music 3.0 comprises three interconnected core components: the tokenizer, the Hybrid-LM, and the synthesis stack. They address, respectively, how musical information is represented, how long-range structure and local detail are modeled, and how audio is reconstructed at high fidelity. The complete workflow is shown below.
Multi-Layer RVQ: Separating Core Structure from Acoustic Detail
An eight-layer RVQ represents musical information hierarchically. The first layer captures core semantics and structure, while the remaining seven progressively encode acoustic residuals. In staged training, the first layer first learns a stable information backbone; all codebooks are then trained jointly to model fine-grained sound. This design provides a more stable discrete representation for long-sequence prediction and avoids overloading a single token layer with both structural and fidelity-related information.
Hybrid-LM: Joint Modeling of Global Structure and Local Acoustics
The 8B Global LLM is initialized from Qwen3.5-8B and performs frame-by-frame prediction of semantic tokens while modeling global context. A randomly initialized 0.6B Local LLM predicts within-frame acoustic tokens along the depth axis. Training proceeds in two stages: global alignment first establishes the Global LLM’s ability to predict musical semantics, followed by full-parameter joint training of the global and local models. This hierarchical collaboration enables the model to preserve song-level structural stability while resolving acoustic detail within each frame.
Hidden-State Fusion: From Discrete Prediction to Continuous Audio Rendering
Conventional systems typically feed discrete acoustic tokens directly into a decoder. Music 3.0 instead fuses continuous hidden states from the Global LLM and Local LLM and uses them to condition a 2.4B flow-matching module. A 123M Flow-VAE then decodes the audio. The complete path is:
Fused LLM features → Flow matching → VAE hidden states → Flow-VAE decoder → Final audio
This design directly connects the language model’s structural understanding with the acoustic model’s audio reconstruction through continuous representations. It preserves long-range consistency while improving pronunciation accuracy, instrumental coherence, and fine-detail fidelity.。
用户prompt
Demo
Classic Shanghai jazz / soul with gentle lo-fi warmth. Tender, nostalgic verses unfold into a softly radiant chorus, led by an intimate female vocal with restrained phrasing, supported by mellow piano, upright bass, brushed drums, warm brass accents, and a vintage room-like mix.
Futuristic melodic EDM / progressive house. Reflective verses about memory and digital identity build into an uplifting, hook-driven chorus, with an emotive lead vocal, bright layered synths, pulsing bass, crisp four-on-the-floor drums, subtle glitch textures, and a wide polished festival mix.
Progressive house / EDM, 126 BPM, B-flat major. Reflective and nostalgic verses rise into a cathartic, euphoric chorus, with a smooth breathy male tenor, pulsing side-chained synths, punchy club bass, crisp drums, and wide hall reverb.
Bright power pop / pop rock, 112 BPM, E-flat major. Nostalgic, bittersweet verses burst into cathartic choruses, led by a clear female mezzo-soprano with soaring octave hooks, punchy drums, wide layered guitars, subtle synths, and a polished modern mix.
Understanding Creative Intent
AI-generated music can sound like a complete song and still drift away from the original brief. Specified instruments may gradually disappear from the arrangement, the intended emotional character may weaken as the song develops, and a requested vocal style may appear in only one section instead of remaining coherent throughout the piece.
Music 3.0 introduces a more expressive music-description framework. Rather than summarizing an entire piece with a single global label, it uses Structured Captions to describe music at fine temporal granularity. These captions specify genre, tempo, time signature, key, use case, and production character, while also tracking emotional contour; the entry and exit of primary and supporting instruments; the development of groove and low-end energy; and section-level changes in vocal delivery, harmony, and vocal effects.
These descriptions convert subjective listening impressions into a professional arrangement framework that the model can learn and execute. As a result, the model can translate creative language into concrete musical expression more accurately, maintain a coherent musical identity as the song unfolds, and still introduce the necessary dynamic variation.
To make professional music creation more accessible, we also developed a template-based Prompt Enhancement System. It selects appropriate language from a curated library of Structured Caption templates and applies established musical terminology and arrangement principles to expand a simple user description into a detailed, musically coherent instruction. Creators can therefore exercise precise control without mastering specialist vocabulary—whether they want an intimate, restrained unplugged performance; a late-night R&B track driven by rolling hi-hats and deep 808 bass; or a cinematic instrumental that grows from introspection to large-scale intensity.
用户prompt
Demo
Warm Mandarin pop / traditional Chinese ballad, 74 BPM, A-flat major. Tender nostalgia opens into a celebratory chorus, with a mature male baritone-tenor, delicate guzheng, soft strings, gentle acoustic textures, expressive vibrato, and a natural live-room sound.
Anthemic stadium pop rock, 132 BPM, E major. Airy, buoyant verses build into an adrenaline-filled chorus, with a powerful female mezzo-soprano, soaring octave hooks, driving drums, wide guitars, shimmering ambience, and a brief weightless bridge.
Baroque pop / emo rock, 162 BPM, E-flat minor. A stately, mournful opening drives toward a regal wall-of-sound finale, with a gritty theatrical male baritenor, orchestral strings, harpsichord, distorted guitars, aggressive drums, and a defiant breakdown.
More Complete, More Varied Arrangements
The challenge of long-form song generation is not merely generating for a longer duration; it is creating a credible progression across sections. Emotion must build and resolve, instruments must enter, layer, and recede at the right moments, and verses, choruses, bridges, and instrumental passages must all serve a unified direction.
Music 3.0 addresses this challenge at two levels. First, section tags in the lyrics—such as [intro], [verse], [pre-chorus], [chorus], [bridge], [instrumental], [solo], and [outro]—define the song’s macrostructure. Second, the Structured Caption specifies the emotional development, instrumentation changes, vocal delivery, rhythmic foundation, ornamental timbres, and spatial effects at each stage, turning what changes, and where, into an explicit generation condition.
At the model level, Music 3.0 uses a global–local collaborative Hybrid-LM. The 8B Global LLM predicts core semantic and structural tokens frame by frame and maintains full-song context. The 0.6B Local LLM predicts acoustic tokens along the depth axis within each frame, supplying local sonic detail. This division of responsibilities maintains temporal stability across songs of up to five minutes while preserving rich variation within individual sections.
用户prompt
Demo
Uplifting progressive house / EDM, 126 BPM, A-flat major. Focused inspiration grows into a euphoric climax, with a warm slightly gravelly male tenor, soaring pentatonic hooks, punchy club drums, controlled bass, bright synths, and a polished panoramic mix.
Upbeat funk / nu-disco, 112 BPM, E-flat major. Confident, strutting and celebratory, with a smooth soulful male tenor, playful staccato phrasing, falsetto flips, snappy drums, elastic bass, rhythmic guitar, glossy keys, and a wide dance-floor mix.
Dark E-punk hip-hop / industrial rap, 108 BPM, B-flat minor. Ominous sermon-like tension turns into club-ready catharsis, with a gravelly male tenor blending aggressive rap, punk shouting and haunting chants over distorted electronics, hard drums, and saturated bass.
Advancing Audio Quality
Compelling composition requires equally convincing sound. Music 3.0 delivers a substantial improvement in audio quality, producing mixes that are more open, clear, and balanced, with less congestion and muddiness.
Audio modeling begins with multi-layer residual vector quantization (RVQ). The first layer uses a 16,384-entry codebook dedicated to the music’s core semantics and structure. Layers 2–8 each use a 1,024-entry codebook and progressively encode residual acoustic detail. During training, we first train the initial layer independently so that it captures core information as comprehensively as possible; all eight layers are then trained jointly. This hierarchical design balances semantic capacity, generation stability, and detail reconstruction while reducing error accumulation in long-sequence generation.
Music 3.0’s final audio rendering, however, goes beyond discrete tokens. At inference time, the synthesis stack does not load the discrete tokenizer’s decoding path. Instead, it fuses continuous hidden states from the final layers of the 8B Global LLM and 0.6B Local LLM, then generates audio directly through flow matching and the Flow-VAE. Compared with discrete tokens alone, these continuous features retain richer high-dimensional acoustic information, improving vocal pronunciation accuracy and the physical coherence of instrumental sound.
The 2.4B flow-matching module maps the fused language-model features into the VAE latent space. The 123M Flow-VAE inherits the MiniMax Speech architecture and is retrained for music-specific dynamic range and spectral distributions to reconstruct the final waveform. This enables the model to follow more precise instrumental directions and reproduce authentic performance techniques such as glissando and legato. Instrumental roles are more distinct within the arrangement, the low end retains impact without masking the mix, and fine sonic detail remains clear even in densely layered productions.
These improvements extend across the full range of musical expression: string attack and bowing in solo instruments, the impact of drums and bass, source separation in dense electronic arrangements, and the sense of space around the voice. Music 3.0 brings each of these elements closer to the listening experience of a fully produced recording.
用户prompt
Demo
Futuristic melodic EDM / progressive house. Reflective verses about memory and digital identity build into an uplifting, hook-driven chorus, with an emotive lead vocal, bright layered synths, pulsing bass, crisp four-on-the-floor drums, subtle glitch textures, and a wide polished festival mix.
Light bossa nova / Brazilian pop, 88 BPM, D-flat major. Serene and meditative, gradually growing into quiet confidence, with a clear intimate female soprano, relaxed phrasing, gentle rising hooks, soft guitar, piano, restrained bass, and delicate percussion.
Warm healing Chinese Mandopop ballad, deep emotional with light rock elements, middle-aged mature male and female duet, mature parental voices aged around 50-70, mother sings gentle tender lead parts with warm mature soft female vocals, slightly aged timbre, wise and tender tone, father provides deep low steady supportive harmonies, rich mature male vocals, slightly weathered gravelly low register, starts with piano and lush strings opening, chorus builds gradually with drums and electric guitar layers, ends fading back to pure acoustic gentle close, heartfelt true emotion without over-dramatizing, subtle restrained Chinese-style parental love, proud yet tender unspoken encouragement, strong Taiwanese accent in vocals, emotional, sincere, warm parental perspective, builds to uplifting hopeful ending, soft whispered intimate outro like a gentle murmur calling home, nostalgic yet empowering family love song, seasoned experienced voices, no youthful bright tone
More Natural Vocals
Vocals are often where the synthetic character of generated music is most apparent. High-frequency artifacts, rigid phrasing, unclear pronunciation, and unnatural breathing can break the listener’s emotional connection even when the song is otherwise highly polished.
Music 3.0 introduces a new audio-rendering system designed to produce more natural, studio-quality vocal performances. The Structured Caption describes vocal timbre, delivery, techniques such as breathiness and falsetto, harmony arrangement, and effects such as delay and Auto-Tune in fine detail. Continuous hidden states fused from the global and local language models carry this performance information into the flow-matching and Flow-VAE generation process. Together, these mechanisms reduce the high-frequency digital artifacts common in generated vocals while improving control over melody, pronunciation, breathing, and layered harmonies.
Vocals can respond more coherently to changes in emotion and rhythm—moving from restrained verses into expansive choruses, or from tightly articulated rhythmic delivery into sustained, open melodic lines. Harmonies unfold more naturally around the lead vocal, and breathing becomes part of the performance rather than a synthetic artifact layered around it.
用户prompt
Demo
Acoustic bossa nova / folk, 88 BPM, F-sharp major. Serene, pastoral and playful, with buoyant choruses, a breathy honeyed female soprano, relaxed speech-song phrasing, fingerpicked acoustic guitar, warm bass, light percussion, and a soft lingering fade.
Upbeat funk / contemporary R&B with a glossy nu-disco edge. Playful, nocturnal verses glide into a celebratory dance-floor chorus, with a smooth soulful lead vocal, elastic bass, syncopated rhythm guitar, bright keys, crisp drums, neon synth accents, and a warm spacious club mix.
Warm pop rock / soul, 88 BPM, B-flat major. Intimate reflection swells into an empowering anthem of self-worth, led by a seasoned female alto with a velvety light rasp, analog-warm guitars, steady drums, soulful keys, reassuring hooks, and wide cinematic choruses.
From creative intent and song structure to local acoustic detail and final audio rendering, Music 3.0 advances music generation from producing plausible audio to realizing a complete, coherent creative vision. We look forward to hearing what you create.
