Semantic Audio Generative Encoder · ×64 compression
SAGE is a compact VAE that compresses music ×64 into a semantically structured latent, striking the best balance of reconstruction quality, semantic structure and speed.
It runs at the inference cost of Stable Audio Open, matches the listening-test quality of SAME-L, a model 8× larger and 4× slower, holds the best FAD and CLAP on 23 of 25 measurements and leads all 19 semantic probes.
A/B comparison
Same clip, every model. Switch instantly, playback never stops.
Benchmarks
SAGE against state-of-the-art audio VAEs on five evaluation sets, none of them seen in training (only FMA is in domain): reconstruction fidelity and how well its latent space supports downstream music understanding, in and out of domain.
Compress stereo 44.1 kHz music at ×64 and rebuild it with state-of-the-art perceptual quality. Whatever is lost here caps every downstream result.
Distances in the latent must reflect genre, artist, instrumentation: that geometry is what makes it navigable for diffusion and useful as a feature space.
Capacity, training data and inference cost. SAGE runs at the speed of Stable Audio Open and over 4× faster than SAME-L; only SAME-S, a distilled variant, is faster, at a FAD-MERT 3 to 10 times higher.
Reconstruction fidelity and perceptual quality against state-of-the-art audio VAEs, measured on five evaluation sets: FAD on three independent embeddings, CLAP similarity, and signal-level SDR/STFT.
A blind MUSHRA test (ITU-R BS.1534-3) on ten 10 s excerpts of commercial music never seen in training: SAGE against SAME-L, Stable Audio Open and CoDiCodec, with a hidden reference and a 3.5 kHz low-pass anchor. 38 participants, 21 after screening on the hidden reference.
Two groups emerge. SAGE (81.6) and SAME-L (81.8) overlap within their confidence intervals, so listeners cannot tell them apart, although SAME-L has 8× the parameters and 4× the inference cost. Stable Audio Open and CoDiCodec form a second group 15 to 17 points lower: at the same real-time factor as Stable Audio Open, SAGE scores 17 points higher. Music2Latent and SAME-S were left out to keep the test short: the first trails these three on FAD-MERT on every set, the second is rated below Stable Audio Open in its own authors' test.
How well the frozen latent space supports downstream semantic tasks (MAEB), measured on three suites: FMA, MoisesDB, and the MAEB protocol (restricted to music).
At a glance
FAD-MERT vs speed on MoisesDB mixtures, where each bubble's area scales with the parameter count. SAGE has the lowest FAD at the speed of SAO; SAME-S is faster but its FAD is 5× higher, and SAME-L comes close at 8× the parameters and 4× the runtime.
UMAP of frozen SAGE latents on MoisesDB isolated stems: no instrument label ever reaches the encoder, yet the latent organizes by timbre.
Under the hood
Every module trains end-to-end with all losses active: reconstruction, adversarial, and semantic. 500 epochs, 1,536 GPU-hours.
With the encoder frozen and the latent pinned, a residual post-net refines perceptual detail, while the representation downstream models consume stays fixed. 992 epochs, 3,043 GPU-hours.