Reserved English test set · real recorded noise

Real noise.
Paired truth.
Listen.

Forty held-out English clips with naturally noisy inputs and aggressive speech-isolated, bandwidth-extended references—compared across KVAE residual-EDM and single-speaker Sidon.

Checkpoint Frozen KVAE 64-dim continuous KVAE sampler Heun · 8 intervals · 15 NFE Corruption None added
KVAE MOS recovery
Sidon MOS recovery
IPTV · KVAE / Sidon
Podcast · KVAE / Sidon

Episode-disjoint test

Two restorers. One panel.

The noisy recordings are real aligned inputs; no synthetic corruption or target leakage is added. The KVAE path uses the frozen posterior-mean codec and the sampling guide’s best configuration: Karras residual-EDM, eight Heun intervals, 15 network calls, σdata 0.5, σ 0.002→2.0, ρ 7, seed 20260831. The comparison is the official single-speaker Sidon v0.1 model—not DialogueSidon—with its own 0.9-peak, 50 Hz high-pass and W2V-BERT preprocessing. Microsoft DNSMOS P.835 OVRL scores the exact MP3s. Recovery is 100 × (restored MOS − noisy MOS) ÷ (clean MOS − noisy MOS), unclamped; n/a means the clean/noisy gap is under 0.1.

Ground truthAligned clean reference
Real noisyRecording → frozen codec
KVAE residual-EDMHeun-8 → frozen KVAE decoder
Sidon v0.1Single-speaker restoration model

Loading the real-noise evaluation panel…