Null TestEleven enhancers. One control. Nothing won on average.
One corpus, 19 conditions, one harness, every engine scored the same way. On pooled word error rate, no enhancer beat the unprocessed control. The null hypothesis won.
The vendor numbers cannot all be true.
Four public answers to the same question. No shared corpus among them.
| Source | Claim | What shipped as evidence |
|---|---|---|
| Krisp | ~46% WER reduction | Marketing figure. Corpus, conditions and transcriber unstated. |
| ai-coustics | ~43% WER reduction | Marketing figure. Same shape, different model family. |
| Deepgram | Enhancement hurts STT accuracy | Published position. No dataset we can find. |
| arXiv:2512.17562 | Raw beat enhanced in all 40 configurations | Independent — but one research model, medical speech, semantic WER. |
Here is a shared corpus. None of the four survives it. The reduction claims: the best engine on this bench moved pooled WER +0.9 points — a rise, not the advertised fall. The never-helps claims: at least one engine beat raw in 11 of 19 conditions. A number without a corpus is not evidence. This one you can download.
The best enhancer, ai-coustics Quail L: 15.2% pooled against the control’s 14.3%. The best you can legally ship, anecho.ai_focus_model_16khz_v8_2: 16.6%. Insertions — invented words — rose from 103 raw to 381 on DeepFilterNet3.
Pooled is not the whole result: at least one engine beat raw in 11 of 19 conditions, by up to −10.5 points on Competing speaker 0 dB. Enhancement pays where the audio is genuinely bad, costs accuracy where it is not; pooling hides the trade.
| Licence | Better than raw | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Raw (no processing) | — | 14.32 | ref | 103 | 60 | — | 0.948 | 19.3 | — | 0 ms |
| Passthrough (control) | — | 14.32 | ±0.00 | 103 | 60 | 0/19 | 0.948 | 19.3 | 0.000 | 0 ms |
| ai-coustics Quail L | proprietary | 15.18 | +0.86 | 128 | 59 | 3/19 | 0.943 | 20.6 | 0.123 | 30 ms |
| anecho.ai_focus_model_16khz_v8_2 | proprietary — ours | 16.59 | +2.27 | 119 | 73 | 4/19 | 0.941 | 18.5 | 0.067 | 15 ms |
| ai-coustics Quail VF 2.2 L | proprietary | 17.59 | +3.27 | 163 | 76 | 5/19 | 0.945 | 16.1 | 0.075 | 30 ms |
| GTCRN | MIT | 20.14 | +5.82 | 223 | 41 | 1/19 | 0.945 | 17.9 | 0.050 | 16 ms |
| FastEnhancer-S | MIT | 20.33 | +6.01 | 173 | 72 | 3/19 | 0.943 | 17.8 | 0.034 | 16 ms |
| FastEnhancer-L | MIT | 20.36 | +6.04 | 171 | 112 | 6/19 | 0.941 | 15.6 | 0.377 | 26 ms |
| FastEnhancer-B | MIT | 20.64 | +6.32 | 163 | 80 | 3/19 | 0.943 | 17.7 | 0.021 | 16 ms |
| FastEnhancer-T | MIT | 21.66 | +7.34 | 146 | 110 | 3/19 | 0.940 | 17.5 | 0.013 | 16 ms |
| FastEnhancer-M | MIT | 22.13 | +7.81 | 234 | 84 | 5/19 | 0.940 | 16.1 | 0.110 | 22 ms |
| DeepFilterNet3 (96-frame warm-up) | MIT OR Apache-2.0 | 30.69 | +16.37 | 375 | 130 | 0/19 | 0.914 | 24.3 | 0.226 | 100 ms |
| DeepFilterNet3 | MIT OR Apache-2.0 | 34.04 | +19.72 | 381 | 200 | 0/19 | 0.901 | 23.9 | 0.133 | 100 ms |
Blue row is the unprocessed reference · red ΔWER means the engine made transcription worse · green VAD FA means it cut false triggers
Generated 2026-08-14T02:16:50Z · faster-whisper:base.en · silero-vad · Apple M1 Max · onnxruntime 1.28.0 · 1 thread
Enhancement is not useless. It is pointed at the wrong consumer.
Wire it upVoice activity F1 barely moves — but in the conditions that break barge-in, the false-alarm rate collapses.
| Engine | Transcription | Turn-taking (Onset) | ||||
|---|---|---|---|---|---|---|
| WER raw → enh | ΔWER | Ins raw → enh | F1 raw → enh | FA Babble 5 dB | FA Competing 5 dB | |
| ai-coustics Quail L | 14.3→15.2 | +0.86 | 103→128 | 0.948→0.943 | 98→96 | 58→58 |
| anecho.ai_focus_model_16khz_v8_2 | 14.3→16.6 | +2.27 | 103→119 | 0.948→0.941 | 98→92 | 58→53 |
| ai-coustics Quail VF 2.2 L | 14.3→17.6 | +3.27 | 103→163 | 0.948→0.945 | 98→75 | 58→37 |
| Resample-only 16→8→16 (control) | 14.8→18.6 | +3.86 | 101→173 | 0.950→0.948 | 98→97 | 58→55 |
| GTCRN | 14.3→20.1 | +5.82 | 103→223 | 0.948→0.945 | 98→92 | 58→57 |
| FastEnhancer-S | 14.3→20.3 | +6.01 | 103→173 | 0.948→0.943 | 98→85 | 58→56 |
| FastEnhancer-L | 14.3→20.4 | +6.04 | 103→171 | 0.948→0.941 | 98→68 | 58→39 |
| FastEnhancer-B | 14.3→20.6 | +6.32 | 103→163 | 0.948→0.943 | 98→87 | 58→56 |
| FastEnhancer-T | 14.3→21.7 | +7.34 | 103→146 | 0.948→0.940 | 98→89 | 58→57 |
| FastEnhancer-M | 14.3→22.1 | +7.81 | 103→234 | 0.948→0.940 | 98→75 | 58→48 |
| DeepFilterNet3 (96-frame warm-up) | 14.3→30.7 | +16.37 | 103→375 | 0.948→0.914 | 98→93 | 58→57 |
| DeepFilterNet3 | 14.3→34.0 | +19.72 | 103→381 | 0.948→0.901 | 98→92 | 58→57 |
Every engine we tested raised word error rate. Several of them cut the VAD false-alarm rate by twenty to thirty points in the conditions that actually break turn-taking. That is not a contradiction — it is two different consumers with two different requirements, and it is why the SDK emits two channels instead of picking a winner on your behalf.
Raw → your transcriberEnhanced → Onset
One machine, one recogniser, one test set.
Exactly which, and where that limits the conclusion.
| Recogniser | CTranslate2 int8 CPU, greedy, model=base.en |
|---|---|
| Voice activity | silero-vad · threshold 0.5 · 512-sample blocks |
| Test set | v1 · 10 speakers · 190 clips · 1485 s |
| Conditions | Clean, Noise -5 dB, Noise 0 dB, Noise 5 dB, Noise 10 dB, Noise 20 dB, Babble 5 dB, Competing speaker 0 dB, Competing speaker 5 dB, Reverb, Reverb + noise 10 dB, Telephony 8 kHz, Telephony + noise 10 dB, telephony_alaw, telephony_g722, telephony_opus12k, telephony_loss3, telephony_loss10, Competing speaker 10 dB |
| Seed | 20260101 |
| Host | Apple M1 Max · macOS-26.5.2-arm64-arm-64bit |
| Runtime | onnxruntime 1.28.0 · 1 intra-op thread |
- DeepFilterNet3 is measured through a block-online adapter: every public ONNX export of it lacks recurrent state tensors, so per-frame streaming is impossible with the published weights. Its RTF here is an upper bound and its latency (100 ms) is far above the ~40 ms a native (Rust/tract) runtime achieves. Read the DFN3 rows as a lower bound on quality and an upper bound on cost.
- ai-coustics rows are a proprietary baseline run through the licensed aic-sdk. They are a reference point, never a shipping code path.
- WER is measured with faster-whisper:base.en (CTranslate2 int8 CPU, greedy, model=base.en). A different recogniser will give different absolute numbers; the harness supports Deepgram and AssemblyAI so the *ordering* can be checked on another engine.
- DeepFilterNet3 runs at 48 kHz on 16 kHz source material upsampled to 48 kHz — there is no real content above 8 kHz for it to work with. That is the honest telephony/VoIP situation, but it is not the condition DFN3 was designed for.
- The test set is 10 speakers / ~190 reference words per condition. That is enough to separate large effects and not enough to resolve differences of a point or two of WER.
- Run-to-run repeatability of the recogniser was measured, not assumed — and the standing warning did not reproduce here. Elsewhere in this project faster-whisper on CPU has been seen to return different transcripts for byte-identical input (three runs of one aggregate gave 121.65 / 122.15 / 121.40%), and the standing guidance from that is to treat differences below roughly 6 pp on a small aggregate as decoder noise. On this host and this harness it did not happen: 25/25 scopes were bit-identical across repeated decodes (raw on 18 condition(s), 180 clips x 3 decodes; clearline on 5 condition(s), 50 clips x 3 decodes), including the arms where the recogniser is already hallucinating past 100% WER. Independently, re-measuring `raw` and `gtcrn` in a fresh process two days after the published run reproduced all 54 rows exactly (max ΔWER 0.0000 pp). Three repeats cannot prove determinism and a different host, venv or thread count may well behave differently, so the conservative 6 pp guidance still governs anything quoted from another machine — but on this table the resolution limit is set by the size of the test set, not by the decoder: one reference word is 0.53 pp in a single condition (190 words), and about 0.03 pp pooled across all 18. Differences of a point or two between adjacent rows are still not results; differences of tens of points are.
- onnxruntime is pinned to 1 intra-op thread with sequential execution: the measured cost is one stream on one core, not a fan-out across the whole machine.
- The `telephony*` conditions in this table are stored at 16 kHz after a full 8 kHz carrier leg (band-limit, codec, packet loss), so every backend sees them at its own native rate. The genuinely narrowband experiment — where nothing is resampled — is a separate artefact (`bench/out/telephony.json`).
- Packet loss is modelled as bursty (2-state Gilbert-Elliott) erasure of 20 ms frames with attenuated-repeat concealment, not i.i.d. loss with silence fill. i.i.d. loss flatters every decoder and silence fill overstates the damage.
What is on the bench, and what you are allowed to ship.
Licence is a first-class column: the engine that comes closest cannot legally go into your product — the closest one you can ship is ours.
| Engine | Licence | Shippable | RTF | Latency | Pooled WER |
|---|---|---|---|---|---|
| Raw (no processing) | — | permitted | — | 0 ms | 14.32% |
| Passthrough (control) | — | permitted | 0.000 | 0 ms | 14.32% |
| ai-coustics Quail L | proprietary | benchmark only | 0.123 | 30 ms | 15.18% |
| anecho.ai_focus_model_16khz_v8_2 | proprietary — ours | permitted | 0.067 | 15 ms | 16.59% |
| ai-coustics Quail VF 2.2 L | proprietary | benchmark only | 0.075 | 30 ms | 17.59% |
| GTCRN | MIT | permitted | 0.050 | 16 ms | 20.14% |
| FastEnhancer-S | MIT | permitted | 0.034 | 16 ms | 20.33% |
| FastEnhancer-L | MIT | permitted | 0.377 | 26 ms | 20.36% |
| FastEnhancer-B | MIT | permitted | 0.021 | 16 ms | 20.64% |
| FastEnhancer-T | MIT | permitted | 0.013 | 16 ms | 21.66% |
| FastEnhancer-M | MIT | permitted | 0.110 | 22 ms | 22.13% |
| DeepFilterNet3 (96-frame warm-up) | MIT OR Apache-2.0 | permitted | 0.226 | 100 ms | 30.69% |
| DeepFilterNet3 | MIT OR Apache-2.0 | permitted | 0.133 | 100 ms | 34.04% |
What happens when the audio never leaves the phone line.
Wideband engines are measured here on narrowband input, which is what running one on a call actually requires.
Retracted · we used to say something else
We used to lead with a community report attributing a large fixed loss of every caller's audio to the default frame-processor resample pattern. Then we built the control, and our own measurement does not support it: 8→16→8 kHz with a deliberately broken sample-repeat resampler scores 150 dB SI-SDR against its input — at an exact 2:1 ratio it is algebraically the identity — while a correct polyphase resampler scores 38 dB, lower, because it actually filters. We are not making that claim any more.
If being the referee means anything, it means publishing the call that goes against you.
Ignore pooled claims. Including any we might publish.
A pooled number is a property of the condition mix as much as of the engine, and no vendor shares a mix. Judge on your own voice.
The complete results: one JSON document at a stable, versioned URL, CC-BY-4.0./data/nulltest/v1/results.json
The harness is private because it embeds a licensed competitor SDK. Think a row is wrong, or your engine belongs on the bench? Write to [email protected] — we run it and publish the row, whichever way it lands.
- The control is a row
- Unprocessed audio is the first row of every table and every ΔWER is measured against it — here the null hypothesis won.
- The same code scores everyone
- One scoring path, nothing special-cased, and the result is published even when it reads “nothing won”.
- Pooled and per-condition are reported separately
- Both are published and neither may stand in for the other, whichever makes the better headline.
- Error decomposition, not just a total
- Substitutions, insertions and deletions are reported separately — inserted chatter and a dropped “not” are different failures.
- A one-word difference is not a result
- Any delta inside a single reference word is reported as a number and described as noise, never as a win.
- Cost is reported next to accuracy
- Real-time factor and latency sit in the same table as WER; an engine that misses the frame deadline has not won.
- The proprietary baseline never ships
- ai-coustics rows are a licensed measuring stick, never a code path in anything we sell.