Skip to content
anecho.ai
Bench · A/B comparator

Listen to the difference.
Then read it.

Every enhancement demo on the internet ends at “doesn’t that sound better?” That question is not the one your product depends on. This one switches channel without moving the playhead, and puts the transcript the speech-to-text engine produced from each channel side by side, scored.

Signal chain

speech + 5-talker babble + pink bed + transients @ 5 dB SNR

16 kHz808000 HzChamber
Decoding…
Time domain · both channels drawn, monitored one litSpace play · A switch
0.00 / 0.00 s
Level · raw over enhanced
Frequency domain · raw channel512-pt FFT · Hann · 50% overlap
Computing FFT…
Signal measurementsComputed in-browser from the decoded PCM of both channels.
Signal measurements computed from the decoded audio of both channels
MeasurementRawEnhancedΔ
Programme leveldBFSdBFS
Noise floordBFSdBFS
Crest factordBdB
Usable bandwidthHzHz
Downstream effect

The signal got cleaner. Did the transcript get better?

Reference: I need to move my appointment from Thursday afternoon to sometime next week, preferably in the morning if you have anything open.

Raw → STTS4 · I5 · D0

i need to move my appointment from thursday after (inserted) noon (substituted for afternoon) to some (inserted) time (substituted for sometime) next week for (inserted) a (inserted) family (substituted for preferably) in the morning if you have any (inserted) think (substituted for anything) open

40.9% WER
22 ref words
Enhanced → STTS1 · I0 · D1

i need to move my appointment from thursday afternoon to sometime next week preferable (substituted for preferably) in the morning if (deleted) you have anything open

9.1% WER
22 ref words
−31.8pts ΔWER

Enhancement helped · Broadband babble is the case suppression was designed for. Substitutions collapse and the residual errors are compounding, not semantic.

  • Substitution
  • Insertion
  • Deletion
  • WER = (S + D + I) / N

Engine under test: Stand-in suppressor · 32.0 ms algorithmic latency · not a measured result

Placeholder · The audio and transcripts on this page are synthetic and reproducible, not recordings or live STT output. Nothing here is a published result for Chamber. Measured numbers land on the Null Test leaderboard.

Live input

Point it at your own room

Both channels are analysed continuously; monitoring is off until you pick one. Use headphones — monitoring through speakers will feed back.

RawEnhanced
Monitor
Onset · silenceStand-in suppressor · stand-in engine
Reading the console
What is measured here
Waveforms, spectrograms, level, noise floor, crest factor and usable bandwidth are computed in your browser from the decoded PCM. The word error rate and the S/I/D counts come from a Levenshtein alignment against the reference, run client-side — the diff you see and the number beside it cannot disagree.
What is a stand-in
The speech is text-to-speech, not human recordings. The degradations are synthetic and reproducible. The enhancement is an oracle-SNR Wiener filter, not the shipping model, and it carries a deliberate over-suppression bias so the failure mode is visible rather than hidden. The hypothesis transcripts reproduce the error patterns each condition produces; they are not the output of a live STT run.
Why the telephony case is here
Because it is the case where nothing we tested helped. On our band-limited condition, no enhancer we measured produced a meaningful gain — the best landed within a single reference word of doing nothing, and several cost seven to ten points of word error rate. Note what that condition is and is not: every engine scored on it is a wideband model handed narrowband material, and no natively-8 kHz engine has been benchmarked yet — including ours. The benchmark methodology page states exactly what the condition contains, and any telephony number quoted from us should carry that with it.
Next

The full matrix

Four conditions is a demo. Null Test runs every condition against every engine against every transcriber and publishes the whole grid.

Open the leaderboard
Architecture

Why we split the stream

The telephony case above is the whole argument for sending clean audio to Onset and raw audio to your transcriber.

Read the guide
Build

Run it on your own audio

Instant key, no sales call. Anecho is priced per minute with a free tier you can actually finish a prototype on.

Quickstart