Semantic Audio Generative Encoder · ×64 compression

Audio compressed into
meaningful latents.

SAGE is a compact VAE that compresses music ×64 into a semantically structured latent, striking the best balance of reconstruction quality, semantic structure and speed.

It runs at the inference cost of Stable Audio Open, matches the listening-test quality of SAME-L, a model 8× larger and 4× slower, holds the best FAD and CLAP on 23 of 25 measurements and leads all 19 semantic probes.

SCROLL TO DISCOVER⌄

A/B comparison

Hear the difference

Same clip, every model. Switch instantly, playback never stops.

Benchmarks

Objective benchmarks

SAGE against state-of-the-art audio VAEs on five evaluation sets, none of them seen in training (only FMA is in domain): reconstruction fidelity and how well its latent space supports downstream music understanding, in and out of domain.

Axis 1 · Reconstruction fidelity

Keep the signal

Compress stereo 44.1 kHz music at ×64 and rebuild it with state-of-the-art perceptual quality. Whatever is lost here caps every downstream result.

Axis 2 · Semantic structure

Organize the signal

Distances in the latent must reflect genre, artist, instrumentation: that geometry is what makes it navigable for diffusion and useful as a feature space.

Model & efficiency

Capacity, training data and inference cost. SAGE runs at the speed of Stable Audio Open and over 4× faster than SAME-L; only SAME-S, a distilled variant, is faster, at a FAD-MERT 3 to 10 times higher.

Reconstruction & perceptual metrics

Reconstruction fidelity and perceptual quality against state-of-the-art audio VAEs, measured on five evaluation sets: FAD on three independent embeddings, CLAP similarity, and signal-level SDR/STFT.

Listening test (MUSHRA)

A blind MUSHRA test (ITU-R BS.1534-3) on ten 10 s excerpts of commercial music never seen in training: SAGE against SAME-L, Stable Audio Open and CoDiCodec, with a hidden reference and a 3.5 kHz low-pass anchor. 38 participants, 21 after screening on the hidden reference.

Two groups emerge. SAGE (81.6) and SAME-L (81.8) overlap within their confidence intervals, so listeners cannot tell them apart, although SAME-L has 8× the parameters and 4× the inference cost. Stable Audio Open and CoDiCodec form a second group 15 to 17 points lower: at the same real-time factor as Stable Audio Open, SAGE scores 17 points higher. Music2Latent and SAME-S were left out to keep the test short: the first trails these three on FAD-MERT on every set, the second is rated below Stable Audio Open in its own authors' test.

Semantic probing metrics

How well the frozen latent space supports downstream semantic tasks (MAEB), measured on three suites: FMA, MoisesDB, and the MAEB protocol (restricted to music).

At a glance

The two axes, one picture

Axis 1 · Reconstruction fidelity

The best fidelity-speed tradeoff

FAD-MERT vs speed on MoisesDB mixtures, where each bubble's area scales with the parameter count. SAGE has the lowest FAD at the speed of SAO; SAME-S is faster but its FAD is 5× higher, and SAME-L comes close at 8× the parameters and 4× the runtime.

0.003 0.005 0.01 0.02 0.05 RTF (log) · ← faster 0.2 0.4 0.6 0.8 1.0 FAD-MERT ↓ · better fidelity SAME-L SAME-S SAO CoDiCodec Music2Latent SAGE
Axis 2 · Semantic structure

Stems cluster by timbre

UMAP of frozen SAGE latents on MoisesDB isolated stems: no instrument label ever reaches the encoder, yet the latent organizes by timbre.

UMAP projection of frozen SAGE latents for isolated MoisesDB stems: points cluster by instrument
piano other keys drums percussion guitar bass vocals

Under the hood

Pin the latent, then polish the decoder

Phase 1 · pre-training
❄ CLAP e 512-d target φ pool + 64→512 ê ℒsem 1 − cos(ê, e) audio E z D reconstruction ℒrec + ℒGAN ❄ frozen loss dataflow

Every module trains end-to-end with all losses active: reconstruction, adversarial, and semantic. 500 epochs, 1,536 GPU-hours.

Phase 2 · decoder fine-tune
audio ❄ E ❄ z D post-net zero-init residual reconstruction ℒrec + ℒGAN ❄ frozen loss dataflow

With the encoder frozen and the latent pinned, a residual post-net refines perceptual detail, while the representation downstream models consume stays fixed. 992 epochs, 3,043 GPU-hours.