Semantic Audio Generative Encoder · ×64 compression

Audio compressed into
meaningful latents.

SAGE is a transformer variational autoencoder that compresses music ×64 into a compact latent that is both faithful to reconstruct and semantically structured.

It leads every FAD and CLAP fidelity metric on five benchmarks (four never seen in training), sweeps all 19 semantic probes, and runs 2-10× faster than SOTA autoencoders up to 9× larger.

SCROLL TO DISCOVER

A/B comparison

Hear the difference

Same clip, every model. Switch instantly, playback never stops.

Benchmarks

Objective benchmarks

SAGE against state-of-the-art audio VAEs on five evaluation sets, of which only FMA is seen in training: reconstruction fidelity and how well its latent space supports downstream music understanding, in and out of domain.

Axis 1 · Reconstruction fidelity

Keep the signal

Compress stereo 44.1 kHz music at ×64 and rebuild it with state-of-the-art perceptual quality. Whatever is lost here caps every downstream result.

Axis 2 · Semantic structure

Organize the signal

Distances in the latent must reflect genre, artist, instrumentation: that geometry is what makes it navigable for diffusion and useful as a feature space.

Model & efficiency

Capacity, training data and inference cost. SAGE is the lightest to run here: fastest per file and lowest real-time factor.

Reconstruction & perceptual metrics

Reconstruction fidelity and perceptual quality against state-of-the-art audio VAEs, measured on five evaluation sets: FAD on three independent embeddings, CLAP similarity, and signal-level SDR/STFT.

Semantic probing metrics

How well the frozen latent space supports downstream semantic tasks (MAEB), measured on three suites: FMA, MoisesDB, and the MAEB protocol (restricted to music).

At a glance

The two axes, one picture

Axis 1 · Reconstruction fidelity

Best fidelity at the lowest runtime

Fidelity vs speed on FMA, where each bubble's area scales with the parameter count. SAGE sits in the winning corner: 2.4× better FAD-MERT than the next best model, at the lowest RTF (47.7 ms per 10 s clip).

0.005 0.01 0.02 0.05 RTF (log) · ← faster 0.10 0.20 0.30 0.40 FAD-MERT ↓ · better fidelity SAME SAO CoDiCodec Music2Latent SAGE
Axis 2 · Semantic structure

Stems cluster by timbre

UMAP of frozen SAGE latents on MoisesDB isolated stems: no instrument label ever reaches the encoder, yet the latent organizes by timbre.

UMAP projection of frozen SAGE latents for isolated MoisesDB stems: points cluster by instrument
bass drums percussion guitar piano other keys vocals

Under the hood

Pin the latent, then polish the decoder

Phase 1 · pre-training
CLAP e 512-d target φ pool + 64→512 ê sem 1 − cos(ê, e) audio E z D reconstruction rec + ℒGAN frozen loss dataflow

Every module trains end-to-end with all losses active: reconstruction, adversarial, and semantic. 500 epochs, 912 GPU-hours.

Phase 2 · decoder fine-tune
audio E z D post-net zero-init residual reconstruction rec + ℒGAN frozen loss dataflow

With the encoder frozen and the latent pinned, a residual post-net refines perceptual detail, while the representation downstream models consume stays fixed. 1,500 epochs, 2,544 GPU-hours.

Subjective evaluation

Take the listening test

Your ears are the final benchmark: a blind MUSHRA test (ITU-R BS.1534-3), 10 trials, about 12 minutes, SAGE against the state of the art.

Start the test