These excerpts come from the test set of MUSDB18, the corpus the predictor was evaluated on. Each one is 2.5 seconds of full-band music, band-limited to a given input sampling rate, encoded, and then completed inside the codec.
Pick an input sampling rate and a codec: every player and spectrogram on the page follows, so you can move across rates while staying on the same excerpt. A single predictor covers all four rates — nothing about the model changes between them.
- Band-limited input
- The input's low band with nothing above the cutoff — the part the decoder already holds exactly. Every system below keeps this band unchanged, so what you hear between them is entirely the band each one recovered.
- A2SB and UniverSR
- Two recent published bandwidth-extension systems, run on the same excerpts. Their output is spliced onto the same low band as ours, so only the recovered band differs.
- Ours
- The band-limited codes completed over residual-quantiser depth, decoded, and spliced onto the low band. No extra bits are sent.
- Codec ceiling
- The same excerpt encoded from the full-band original and decoded again. It is what the codec can reproduce at this bitrate, so it bounds what any predictor inside it can reach.
- Reference
- The original 48 kHz excerpt, untouched.
Spectrograms run from 0 to 24 kHz; the vertical axis is marked in kHz.
Everything above 4 kHz is missing from the input and has to be recovered, inside the SpectroStream codec.
Bobby Nobody — Stitch Up
MUSDB18 test set
No 32 kHz model is provided by the authors.
Arise — Run Run Run
MUSDB18 test set
No 32 kHz model is provided by the authors.
The Doppler Shift — Atrophy
MUSDB18 test set
No 32 kHz model is provided by the authors.
Punkdisco — Oral Hygiene
MUSDB18 test set
No 32 kHz model is provided by the authors.
Out-of-domain results on solo orchestral instruments are on the OrchideaSOL page. Code and trained models: github.com/Multi-Rate-BWE-by-Token-Completion/Multi-rate-BWE.