Skip to content
anecho.ai
Proof · nulltest.dev

Null TestEleven enhancers. One control. Nothing won on average.

One corpus, 19 conditions, one harness, every engine scored the same way. On pooled word error rate, no enhancer beat the unprocessed control. The null hypothesis won.

14.3% WER
Unprocessed control3,610 reference words across 19 conditions
15.2% WER
Best enhancerai-coustics Quail L — still worse than doing nothing
0of 11
Engines that beat raw, pooledBut at least one engine beat raw in 11 of 19 individual conditions
103→381
Insertions, raw → worstEnhancement produced more invented words, not fewer
Contrast · the published claims

The vendor numbers cannot all be true.

Four public answers to the same question. No shared corpus among them.

Published enhancement claims and their evidence
SourceClaimWhat shipped as evidence
Krisp~46% WER reductionMarketing figure. Corpus, conditions and transcriber unstated.
ai-coustics~43% WER reductionMarketing figure. Same shape, different model family.
DeepgramEnhancement hurts STT accuracyPublished position. No dataset we can find.
arXiv:2512.17562Raw beat enhanced in all 40 configurationsIndependent — but one research model, medical speech, semantic WER.

Here is a shared corpus. None of the four survives it. The reduction claims: the best engine on this bench moved pooled WER +0.9 points — a rise, not the advertised fall. The never-helps claims: at least one engine beat raw in 11 of 19 conditions. A number without a corpus is not evidence. This one you can download.

The best enhancer, ai-coustics Quail L: 15.2% pooled against the control’s 14.3%. The best you can legally ship, anecho.ai_focus_model_16khz_v8_2: 16.6%. Insertions — invented words — rose from 103 raw to 381 on DeepFilterNet3.

Pooled is not the whole result: at least one engine beat raw in 11 of 19 conditions, by up to −10.5 points on Competing speaker 0 dB. Enhancement pays where the audio is genuinely bad, costs accuracy where it is not; pooling hides the trade.

13 engines
Pooled results across every acoustic condition, by enhancement engine. The unprocessed control is the first row and every delta is measured against it.
LicenceBetter than raw
Raw (no processing)—14.32ref10360—0.94819.3—0 ms
Passthrough (control)—14.32±0.00103600/190.94819.30.0000 ms
ai-coustics Quail Lproprietary15.18+0.86128593/190.94320.60.12330 ms
anecho.ai_focus_model_16khz_v8_2proprietary — ours16.59+2.27119734/190.94118.50.06715 ms
ai-coustics Quail VF 2.2 Lproprietary17.59+3.27163765/190.94516.10.07530 ms
GTCRNMIT20.14+5.82223411/190.94517.90.05016 ms
FastEnhancer-SMIT20.33+6.01173723/190.94317.80.03416 ms
FastEnhancer-LMIT20.36+6.041711126/190.94115.60.37726 ms
FastEnhancer-BMIT20.64+6.32163803/190.94317.70.02116 ms
FastEnhancer-TMIT21.66+7.341461103/190.94017.50.01316 ms
FastEnhancer-MMIT22.13+7.81234845/190.94016.10.11022 ms
DeepFilterNet3 (96-frame warm-up)MIT OR Apache-2.030.69+16.373751300/190.91424.30.226100 ms
DeepFilterNet3MIT OR Apache-2.034.04+19.723812000/190.90123.90.133100 ms

Blue row is the unprocessed reference · red ΔWER means the engine made transcription worse · green VAD FA means it cut false triggers

Generated 2026-08-14T02:16:50Z · faster-whisper:base.en · silero-vad · Apple M1 Max · onnxruntime 1.28.0 · 1 thread

Route · what the data changed our mind about

Enhancement is not useless. It is pointed at the wrong consumer.

Wire it up

Voice activity F1 barely moves — but in the conditions that break barge-in, the false-alarm rate collapses.

Split pipeline · measuredSame audio, two consumers, opposite verdicts
Effect of enhancement on transcription versus on voice activity detection, per engine.
EngineTranscriptionTurn-taking (Onset)
WER raw → enhΔWERIns raw → enhF1 raw → enhFA Babble 5 dBFA Competing 5 dB
ai-coustics Quail L14.3→15.2+0.86103→1280.948→0.94398→9658→58
anecho.ai_focus_model_16khz_v8_214.3→16.6+2.27103→1190.948→0.94198→9258→53
ai-coustics Quail VF 2.2 L14.3→17.6+3.27103→1630.948→0.94598→7558→37
Resample-only 16→8→16 (control)14.8→18.6+3.86101→1730.950→0.94898→9758→55
GTCRN14.3→20.1+5.82103→2230.948→0.94598→9258→57
FastEnhancer-S14.3→20.3+6.01103→1730.948→0.94398→8558→56
FastEnhancer-L14.3→20.4+6.04103→1710.948→0.94198→6858→39
FastEnhancer-B14.3→20.6+6.32103→1630.948→0.94398→8758→56
FastEnhancer-T14.3→21.7+7.34103→1460.948→0.94098→8958→57
FastEnhancer-M14.3→22.1+7.81103→2340.948→0.94098→7558→48
DeepFilterNet3 (96-frame warm-up)14.3→30.7+16.37103→3750.948→0.91498→9358→57
DeepFilterNet314.3→34.0+19.72103→3810.948→0.90198→9258→57

Every engine we tested raised word error rate. Several of them cut the VAD false-alarm rate by twenty to thirty points in the conditions that actually break turn-taking. That is not a contradiction — it is two different consumers with two different requirements, and it is why the SDK emits two channels instead of picking a winner on your behalf.

Raw → your transcriberEnhanced → Onset

Meas · provenance

One machine, one recogniser, one test set.

Exactly which, and where that limits the conclusion.

Run configuration
Harness configuration for this run
RecogniserCTranslate2 int8 CPU, greedy, model=base.en
Voice activitysilero-vad · threshold 0.5 · 512-sample blocks
Test setv1 · 10 speakers · 190 clips · 1485 s
ConditionsClean, Noise -5 dB, Noise 0 dB, Noise 5 dB, Noise 10 dB, Noise 20 dB, Babble 5 dB, Competing speaker 0 dB, Competing speaker 5 dB, Reverb, Reverb + noise 10 dB, Telephony 8 kHz, Telephony + noise 10 dB, telephony_alaw, telephony_g722, telephony_opus12k, telephony_loss3, telephony_loss10, Competing speaker 10 dB
Seed20260101
HostApple M1 Max · macOS-26.5.2-arm64-arm-64bit
Runtimeonnxruntime 1.28.0 · 1 intra-op thread
The harness’s own caveats
  • DeepFilterNet3 is measured through a block-online adapter: every public ONNX export of it lacks recurrent state tensors, so per-frame streaming is impossible with the published weights. Its RTF here is an upper bound and its latency (100 ms) is far above the ~40 ms a native (Rust/tract) runtime achieves. Read the DFN3 rows as a lower bound on quality and an upper bound on cost.
  • ai-coustics rows are a proprietary baseline run through the licensed aic-sdk. They are a reference point, never a shipping code path.
  • WER is measured with faster-whisper:base.en (CTranslate2 int8 CPU, greedy, model=base.en). A different recogniser will give different absolute numbers; the harness supports Deepgram and AssemblyAI so the *ordering* can be checked on another engine.
  • DeepFilterNet3 runs at 48 kHz on 16 kHz source material upsampled to 48 kHz — there is no real content above 8 kHz for it to work with. That is the honest telephony/VoIP situation, but it is not the condition DFN3 was designed for.
  • The test set is 10 speakers / ~190 reference words per condition. That is enough to separate large effects and not enough to resolve differences of a point or two of WER.
  • Run-to-run repeatability of the recogniser was measured, not assumed — and the standing warning did not reproduce here. Elsewhere in this project faster-whisper on CPU has been seen to return different transcripts for byte-identical input (three runs of one aggregate gave 121.65 / 122.15 / 121.40%), and the standing guidance from that is to treat differences below roughly 6 pp on a small aggregate as decoder noise. On this host and this harness it did not happen: 25/25 scopes were bit-identical across repeated decodes (raw on 18 condition(s), 180 clips x 3 decodes; clearline on 5 condition(s), 50 clips x 3 decodes), including the arms where the recogniser is already hallucinating past 100% WER. Independently, re-measuring `raw` and `gtcrn` in a fresh process two days after the published run reproduced all 54 rows exactly (max ΔWER 0.0000 pp). Three repeats cannot prove determinism and a different host, venv or thread count may well behave differently, so the conservative 6 pp guidance still governs anything quoted from another machine — but on this table the resolution limit is set by the size of the test set, not by the decoder: one reference word is 0.53 pp in a single condition (190 words), and about 0.03 pp pooled across all 18. Differences of a point or two between adjacent rows are still not results; differences of tens of points are.
  • onnxruntime is pinned to 1 intra-op thread with sequential execution: the measured cost is one stream on one core, not a fan-out across the whole machine.
  • The `telephony*` conditions in this table are stored at 16 kHz after a full 8 kHz carrier leg (band-limit, codec, packet loss), so every backend sees them at its own native rate. The genuinely narrowband experiment — where nothing is resampled — is a separate artefact (`bench/out/telephony.json`).
  • Packet loss is modelled as bursty (2-state Gilbert-Elliott) erasure of 20 ms frames with attenuated-repeat concealment, not i.i.d. loss with silence fill. i.i.d. loss flatters every decoder and silence fill overstates the damage.
Rack · engines under test

What is on the bench, and what you are allowed to ship.

Licence is a first-class column: the engine that comes closest cannot legally go into your product — the closest one you can ship is ours.

Engines under test and their licences
EngineLicenceShippableRTFLatencyPooled WER
Raw (no processing)—permitted—0 ms14.32%
Passthrough (control)—permitted0.0000 ms14.32%
ai-coustics Quail Lproprietarybenchmark only0.12330 ms15.18%
anecho.ai_focus_model_16khz_v8_2proprietary — ourspermitted0.06715 ms16.59%
ai-coustics Quail VF 2.2 Lproprietarybenchmark only0.07530 ms17.59%
GTCRNMITpermitted0.05016 ms20.14%
FastEnhancer-SMITpermitted0.03416 ms20.33%
FastEnhancer-LMITpermitted0.37726 ms20.36%
FastEnhancer-BMITpermitted0.02116 ms20.64%
FastEnhancer-TMITpermitted0.01316 ms21.66%
FastEnhancer-MMITpermitted0.11022 ms22.13%
DeepFilterNet3 (96-frame warm-up)MIT OR Apache-2.0permitted0.226100 ms30.69%
DeepFilterNet3MIT OR Apache-2.0permitted0.133100 ms34.04%
Narrowband run · 8 kHz end to end · G.711 µ-law, 300–3400 Hz

What happens when the audio never leaves the phone line.

Wideband engines are measured here on narrowband input, which is what running one on a call actually requires.

11.1% WER
Untouched 8 kHz caller audioThe null hypothesis · 60 clips, 1,140 reference words · faster-whisper:base.en
−0.18pts
Best path that exists todayresample only (soxr_vhq) — 2 words out of 1,140. A tie, not a win.
+13.7pts
Worst path we measuredfastenhancer-t via soxr_vhq — a wideband model handed a 3.4 kHz signal through a resample wrapper.

Retracted · we used to say something else

We used to lead with a community report attributing a large fixed loss of every caller's audio to the default frame-processor resample pattern. Then we built the control, and our own measurement does not support it: 8→16→8 kHz with a deliberately broken sample-repeat resampler scores 150 dB SI-SDR against its input — at an exact 2:1 ratio it is algebraically the identity — while a correct polyphase resampler scores 38 dB, lower, because it actually filters. We are not making that claim any more.

If being the referee means anything, it means publishing the call that goes against you.

Verdict · what to do with this

Ignore pooled claims. Including any we might publish.

A pooled number is a property of the condition mix as much as of the engine, and no vendor shares a mix. Judge on your own voice.

Take the data

The complete results: one JSON document at a stable, versioned URL, CC-BY-4.0./data/nulltest/v1/results.json

Challenge it

The harness is private because it embeds a licensed competitor SDK. Think a row is wrong, or your engine belongs on the bench? Write to [email protected] — we run it and publish the row, whichever way it lands.

Rules we hold ourselves to
The control is a row
Unprocessed audio is the first row of every table and every ΔWER is measured against it — here the null hypothesis won.
The same code scores everyone
One scoring path, nothing special-cased, and the result is published even when it reads “nothing won”.
Pooled and per-condition are reported separately
Both are published and neither may stand in for the other, whichever makes the better headline.
Error decomposition, not just a total
Substitutions, insertions and deletions are reported separately — inserted chatter and a dropped “not” are different failures.
A one-word difference is not a result
Any delta inside a single reference word is reported as a number and described as noise, never as a win.
Cost is reported next to accuracy
Real-time factor and latency sit in the same table as WER; an engine that misses the frame deadline has not won.
The proprietary baseline never ships
ai-coustics rows are a licensed measuring stick, never a code path in anything we sell.