Reserved English test set · real recorded noise

Real noise.
Paired truth.
Listen.

Forty held-out English clips with naturally noisy inputs and aggressive speech-isolated, bandwidth-extended references—compared across Super Sonique and single-speaker Sidon.

Checkpoint Frozen KVAE 64-dim continuous Super Sonique Heun · 8 intervals · 15 NFE · EMA Corruption None added
Super Sonique MOS recovery
Sidon MOS recovery
IPTV · Sonique / Sidon
Podcast · Sonique / Sidon

Episode-disjoint test

Two restorers. One panel.

The noisy recordings are real aligned inputs; no synthetic corruption or target leakage is added. Super Sonique is the 1,030-hour trajectory-supervised residual-EDM model evaluated with its EMA weights and sealed best sampler: eight Heun intervals, 15 network calls, σdata 0.5, σ 0.002→2.0, ρ 7, seed 20260831. Its frozen posterior-mean KVAE decoder uses 250-frame cores with 11-frame halos. The comparison is the official single-speaker Sidon v0.1 model—not DialogueSidon—with its own 0.9-peak, 50 Hz high-pass and W2V-BERT preprocessing. Microsoft DNSMOS P.835 OVRL scores the exact MP3s. Recovery is unclamped; n/a means the clean/noisy MOS gap is under 0.1.

Ground truthAligned clean reference
Real noisyRecording → frozen codec
Super SoniqueEMA · Heun-8 → frozen KVAE decoder
Sidon v0.1Single-speaker restoration model

Loading the real-noise evaluation panel…