Semantic Audio Generative Encoder · ×64 compression
SAGE is a transformer variational autoencoder that compresses music ×64 into a compact latent that is both faithful to reconstruct and semantically structured.
It leads every FAD and CLAP fidelity metric on five benchmarks (four never seen in training), sweeps all 19 semantic probes, and runs 2-10× faster than SOTA autoencoders up to 9× larger.
A/B comparison
Same clip, every model. Switch instantly, playback never stops.
Benchmarks
SAGE against state-of-the-art audio VAEs on five evaluation sets, of which only FMA is seen in training: reconstruction fidelity and how well its latent space supports downstream music understanding, in and out of domain.
Compress stereo 44.1 kHz music at ×64 and rebuild it with state-of-the-art perceptual quality. Whatever is lost here caps every downstream result.
Distances in the latent must reflect genre, artist, instrumentation: that geometry is what makes it navigable for diffusion and useful as a feature space.
Capacity, training data and inference cost. SAGE is the lightest to run here: fastest per file and lowest real-time factor.
Reconstruction fidelity and perceptual quality against state-of-the-art audio VAEs, measured on five evaluation sets: FAD on three independent embeddings, CLAP similarity, and signal-level SDR/STFT.
How well the frozen latent space supports downstream semantic tasks (MAEB), measured on three suites: FMA, MoisesDB, and the MAEB protocol (restricted to music).
At a glance
Fidelity vs speed on FMA, where each bubble's area scales with the parameter count. SAGE sits in the winning corner: 2.4× better FAD-MERT than the next best model, at the lowest RTF (47.7 ms per 10 s clip).
UMAP of frozen SAGE latents on MoisesDB isolated stems: no instrument label ever reaches the encoder, yet the latent organizes by timbre.
Under the hood
Every module trains end-to-end with all losses active: reconstruction, adversarial, and semantic. 500 epochs, 912 GPU-hours.
With the encoder frozen and the latent pinned, a residual post-net refines perceptual detail, while the representation downstream models consume stays fixed. 1,500 epochs, 2,544 GPU-hours.
Subjective evaluation
Your ears are the final benchmark: a blind MUSHRA test (ITU-R BS.1534-3), 10 trials, about 12 minutes, SAGE against the state of the art.
Start the test