Listen to the difference.
Then read it.
Every enhancement demo on the internet ends at “doesn’t that sound better?” That question is not the one your product depends on. This one switches channel without moving the playhead, and puts the transcript the speech-to-text engine produced from each channel side by side, scored.
speech + 5-talker babble + pink bed + transients @ 5 dB SNR
| Measurement | Raw | Enhanced | Δ |
|---|---|---|---|
| Programme level | —dBFS | —dBFS | — |
| Noise floor | —dBFS | —dBFS | — |
| Crest factor | —dB | —dB | — |
| Usable bandwidth | —Hz | —Hz | — |
The signal got cleaner. Did the transcript get better?
Reference: “I need to move my appointment from Thursday afternoon to sometime next week, preferably in the morning if you have anything open.”
i need to move my appointment from thursday after (inserted) noon (substituted for afternoon) to some (inserted) time (substituted for sometime) next week for (inserted) a (inserted) family (substituted for preferably) in the morning if you have any (inserted) think (substituted for anything) open
i need to move my appointment from thursday afternoon to sometime next week preferable (substituted for preferably) in the morning if (deleted) you have anything open
Enhancement helped · Broadband babble is the case suppression was designed for. Substitutions collapse and the residual errors are compounding, not semantic.
- Substitution
- Insertion
- Deletion
- WER = (S + D + I) / N
Engine under test: Stand-in suppressor · 32.0 ms algorithmic latency · not a measured result
Placeholder · The audio and transcripts on this page are synthetic and reproducible, not recordings or live STT output. Nothing here is a published result for Chamber. Measured numbers land on the Null Test leaderboard.
Point it at your own room
Both channels are analysed continuously; monitoring is off until you pick one. Use headphones — monitoring through speakers will feed back.
- What is measured here
- Waveforms, spectrograms, level, noise floor, crest factor and usable bandwidth are computed in your browser from the decoded PCM. The word error rate and the S/I/D counts come from a Levenshtein alignment against the reference, run client-side — the diff you see and the number beside it cannot disagree.
- What is a stand-in
- The speech is text-to-speech, not human recordings. The degradations are synthetic and reproducible. The enhancement is an oracle-SNR Wiener filter, not the shipping model, and it carries a deliberate over-suppression bias so the failure mode is visible rather than hidden. The hypothesis transcripts reproduce the error patterns each condition produces; they are not the output of a live STT run.
- Why the telephony case is here
- Because it is the case where nothing we tested helped. On our band-limited condition, no enhancer we measured produced a meaningful gain — the best landed within a single reference word of doing nothing, and several cost seven to ten points of word error rate. Note what that condition is and is not: every engine scored on it is a wideband model handed narrowband material, and no natively-8 kHz engine has been benchmarked yet — including ours. The benchmark methodology page states exactly what the condition contains, and any telephony number quoted from us should carry that with it.
The full matrix
Four conditions is a demo. Null Test runs every condition against every engine against every transcriber and publishes the whole grid.
Open the leaderboardWhy we split the stream
The telephony case above is the whole argument for sending clean audio to Onset and raw audio to your transcriber.
Read the guideRun it on your own audio
Instant key, no sales call. Anecho is priced per minute with a free tier you can actually finish a prototype on.
Quickstart