Null TestTen enhancers. One control. Nothing won on average.
This is not the product — Clearline is. This is how we prove what we say about it. Audio condition × enhancement engine × transcriber, one harness, one open dataset, every vendor scored the same way, results published whether or not they suit us.
No speech enhancer we tested beat doing nothing on pooled word error rate. The best of them, ai-coustics Quail L, scored 15.6% against the control’s 14.8%. The best engine with a licence you can actually ship, GTCRN, scored 20.3%.
That pooled number is the headline and it is not the whole result. Per condition, at least one engine did beat raw in 11 of 18 cases — up to −5.8 points on Competing speaker 5 dB — and the only conditions where nothing helped were the two reverberant ones. The honest summary is that enhancement pays where the audio is genuinely bad and costs you accuracy where it is not, and pooling across a mix that is mostly not-that-bad hides the trade.
Insertions — words the recogniser invented that were never spoken — went up, from 101 on raw audio to 410 on Clearline (8 kHz, resampled). That is the direct opposite of the vendor claim that enhancement removes hallucinated words. On the band-limited condition, no engine produced a meaningful improvement — the best, FastEnhancer-S, landed −0.5 points from the control, which is inside a single reference word — while the worst, DeepFilterNet3 (96-frame warm-up), cost +10.0 points.
| Licence | Better than raw | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Raw (no processing) | — | 14.77 | ref | 101 | 60 | — | 0.950 | 17.2 | — | 0 ms |
| Passthrough (control) | — | 14.77 | ±0.00 | 101 | 60 | 0/18 | 0.950 | 17.2 | 0.000 | 0 ms |
| ai-coustics Quail L | proprietary | 15.56 | +0.79 | 126 | 57 | 3/18 | 0.945 | 18.7 | 0.124 | 30 ms |
| ai-coustics Quail VF 2.2 L | proprietary | 18.10 | +3.33 | 161 | 75 | 5/18 | 0.945 | 15.3 | 0.076 | 30 ms |
| GTCRN | MIT | 20.26 | +5.50 | 210 | 40 | 1/18 | 0.947 | 16.4 | 0.050 | 16 ms |
| FastEnhancer-L | MIT | 20.44 | +5.67 | 159 | 107 | 6/18 | 0.942 | 14.7 | 0.376 | 26 ms |
| FastEnhancer-S | MIT | 20.64 | +5.88 | 159 | 72 | 3/18 | 0.945 | 16.2 | 0.034 | 16 ms |
| FastEnhancer-B | MIT | 21.14 | +6.37 | 159 | 79 | 3/18 | 0.944 | 16.2 | 0.021 | 16 ms |
| FastEnhancer-M | MIT | 22.40 | +7.63 | 217 | 83 | 5/18 | 0.941 | 14.9 | 0.110 | 22 ms |
| FastEnhancer-T | MIT | 22.46 | +7.69 | 145 | 109 | 3/18 | 0.941 | 15.8 | 0.013 | 16 ms |
| DeepFilterNet3 (96-frame warm-up) | MIT OR Apache-2.0 | 31.64 | +16.87 | 368 | 129 | 0/18 | 0.914 | 22.8 | 0.226 | 100 ms |
| DeepFilterNet3 | MIT OR Apache-2.0 | 35.18 | +20.41 | 380 | 197 | 0/18 | 0.901 | 22.2 | 0.133 | 100 ms |
| Clearline (8 kHz, resampled) | internal | 39.65 | +24.88 | 410 | 170 | 0/18 | 0.835 | 15.1 | 0.048 | 16 ms |
Blue row is the unprocessed reference · red ΔWER means the engine made transcription worse · green VAD FA means it cut false triggers
Generated 2026-08-14T02:16:50Z · faster-whisper:base.en · silero-vad · Apple M1 Max · onnxruntime 1.28.0 · 1 thread
Enhancement is not useless. It is pointed at the wrong consumer.
Wire it upWe expected the transcription result. We did not expect the turn-taking result to be this clean. Voice activity F1 barely moves, but in the conditions that actually break barge-in, the false-alarm rate collapses.
| Engine | Transcription | Turn-taking (Onset) | ||||
|---|---|---|---|---|---|---|
| WER raw → enh | ΔWER | Ins raw → enh | F1 raw → enh | FA Babble 5 dB | FA Competing 5 dB | |
| ai-coustics Quail L | 14.8→15.6 | +0.79 | 101→126 | 0.950→0.945 | 98→96 | 58→58 |
| ai-coustics Quail VF 2.2 L | 14.8→18.1 | +3.33 | 101→161 | 0.950→0.945 | 98→75 | 58→37 |
| Resample-only 16→8→16 (control) | 14.8→18.6 | +3.86 | 101→173 | 0.950→0.948 | 98→97 | 58→55 |
| GTCRN | 14.8→20.3 | +5.50 | 101→210 | 0.950→0.947 | 98→92 | 58→57 |
| FastEnhancer-L | 14.8→20.4 | +5.67 | 101→159 | 0.950→0.942 | 98→68 | 58→39 |
| FastEnhancer-S | 14.8→20.6 | +5.88 | 101→159 | 0.950→0.945 | 98→85 | 58→56 |
| FastEnhancer-B | 14.8→21.1 | +6.37 | 101→159 | 0.950→0.944 | 98→87 | 58→56 |
| FastEnhancer-M | 14.8→22.4 | +7.63 | 101→217 | 0.950→0.941 | 98→75 | 58→48 |
| FastEnhancer-T | 14.8→22.5 | +7.69 | 101→145 | 0.950→0.941 | 98→89 | 58→57 |
| DeepFilterNet3 (96-frame warm-up) | 14.8→31.6 | +16.87 | 101→368 | 0.950→0.914 | 98→93 | 58→57 |
| DeepFilterNet3 | 14.8→35.2 | +20.41 | 101→380 | 0.950→0.901 | 98→92 | 58→57 |
| Clearline (8 kHz, resampled) | 14.8→39.6 | +24.88 | 101→410 | 0.950→0.835 | 98→82 | 58→51 |
Every engine we tested raised word error rate. Several of them cut the VAD false-alarm rate by twenty to thirty points in the conditions that actually break turn-taking. That is not a contradiction — it is two different consumers with two different requirements, and it is why the SDK emits two channels instead of picking a winner on your behalf.
Raw → your transcriberEnhanced → Onset
What this run does and does not prove.
Every number above came off one machine, one recogniser and one test set. Here is exactly which, and where that limits the conclusion.
| Recogniser | CTranslate2 int8 CPU, greedy, model=base.en |
|---|---|
| Voice activity | silero-vad · threshold 0.5 · 512-sample blocks |
| Test set | v1 · 10 speakers · 180 clips · 1406 s |
| Conditions | Clean, Noise -5 dB, Noise 0 dB, Noise 5 dB, Noise 10 dB, Noise 20 dB, Babble 5 dB, Competing speaker 0 dB, Competing speaker 5 dB, Reverb, Reverb + noise 10 dB, Telephony 8 kHz, Telephony + noise 10 dB, telephony_alaw, telephony_g722, telephony_opus12k, telephony_loss3, telephony_loss10 |
| Seed | 20260101 |
| Host | Apple M1 Max · macOS-26.5.2-arm64-arm-64bit |
| Runtime | onnxruntime 1.28.0 · 1 intra-op thread |
- DeepFilterNet3 is measured through a block-online adapter: every public ONNX export of it lacks recurrent state tensors, so per-frame streaming is impossible with the published weights. Its RTF here is an upper bound and its latency (100 ms) is far above the ~40 ms a native (Rust/tract) runtime achieves. Read the DFN3 rows as a lower bound on quality and an upper bound on cost.
- ai-coustics rows are a proprietary baseline run through the licensed aic-sdk. They are a reference point, never a shipping code path.
- WER is measured with faster-whisper:base.en (CTranslate2 int8 CPU, greedy, model=base.en). A different recogniser will give different absolute numbers; the harness supports Deepgram and AssemblyAI so the *ordering* can be checked on another engine.
- Clearline is an 8 kHz model and these conditions are stored at 16 kHz, so its row is measured through the harness's resample path — 16 k -> 8 k -> model -> 8 k -> 16 k (soxr), the same path the native-8 kHz ai-coustics baselines take. On `clean`, `noise_*`, `babble_snr5`, `competing_*` and `reverb*` that is not its operating point: the 4-8 kHz band is discarded before the model ever sees it. The `resample8k` control row is that round trip with no model in it, so the band limit and the model can be told apart; and the telephony sub-aggregate below reports the regime the model was actually built for.
- The Clearline checkpoint benchmarked here is not shippable, and this row is not a claim that it is. Two gates it does not pass, both measured on its own eval (`runs/t4_300h_200ep_solo`, 1,800-clip test_v2) and neither of them visible in this table: (1) a solo regression — on clips with no competing speaker the model makes the audio worse, -0.874 dB SI-SDR with 37.1% of solo clips degraded (-0.595 dB / 33.0% on val_v2); (2) speaker conditioning contributes approximately nothing — the matched ablation scores FiLM +2.815 dB against no-conditioning +2.690 dB, a 0.125 dB difference against a 0.3 dB noise floor, so what is being measured here is a plain denoiser and not the speaker isolation the architecture is for. The row exists so Clearline can be *placed* next to the competition on identical audio, not to argue that it is ready.
- Clearline's own eval (`clearline/evaluate.py`, 1,800 clips, difficulty buckets) and this leaderboard are not comparable and were never meant to be: different clips, different conditions, different pooling. Only the recogniser is shared. This row exists precisely so that a comparable number exists.
- DeepFilterNet3 runs at 48 kHz on 16 kHz source material upsampled to 48 kHz — there is no real content above 8 kHz for it to work with. That is the honest telephony/VoIP situation, but it is not the condition DFN3 was designed for.
- The test set is 10 speakers / ~190 reference words per condition. That is enough to separate large effects and not enough to resolve differences of a point or two of WER.
- Run-to-run repeatability of the recogniser was measured, not assumed — and the standing warning did not reproduce here. Elsewhere in this project faster-whisper on CPU has been seen to return different transcripts for byte-identical input (three runs of one aggregate gave 121.65 / 122.15 / 121.40%), and the standing guidance from that is to treat differences below roughly 6 pp on a small aggregate as decoder noise. On this host and this harness it did not happen: 25/25 scopes were bit-identical across repeated decodes (raw on 18 condition(s), 180 clips x 3 decodes; clearline on 5 condition(s), 50 clips x 3 decodes), including the arms where the recogniser is already hallucinating past 100% WER. Independently, re-measuring `raw` and `gtcrn` in a fresh process two days after the published run reproduced all 54 rows exactly (max ΔWER 0.0000 pp). Three repeats cannot prove determinism and a different host, venv or thread count may well behave differently, so the conservative 6 pp guidance still governs anything quoted from another machine — but on this table the resolution limit is set by the size of the test set, not by the decoder: one reference word is 0.53 pp in a single condition (190 words), and about 0.03 pp pooled across all 18. Differences of a point or two between adjacent rows are still not results; differences of tens of points are.
- onnxruntime is pinned to 1 intra-op thread with sequential execution: the measured cost is one stream on one core, not a fan-out across the whole machine.
- The `telephony*` conditions in this table are stored at 16 kHz after a full 8 kHz carrier leg (band-limit, codec, packet loss), so every backend sees them at its own native rate. The genuinely narrowband experiment — where nothing is resampled — is a separate artefact (`bench/out/telephony.json`).
- Packet loss is modelled as bursty (2-state Gilbert-Elliott) erasure of 20 ms frames with attenuated-repeat concealment, not i.i.d. loss with silence fill. i.i.d. loss flatters every decoder and silence fill overstates the damage.
- The control is a row
- Unprocessed audio is the first row of every table, and every ΔWER is measured against it in the same condition. A benchmark without a null hypothesis is a brochure — and in this case the null hypothesis won.
- We publish results that embarrass us
- This benchmark exists to sell an engine. It currently says that no engine on the market — including the one we would most like to beat — lowered pooled word error rate against doing nothing, and it does not yet contain a single row for ours. When those rows land they will be produced by this same code, and they will stay up whatever they say.
- Pooled and per-condition are reported separately
- A pooled average across a fixed condition mix is a property of the mix as much as of the engine. So the pooled number and the per-condition number are both published, and neither is allowed to stand in for the other — including when the pooled one makes a better headline.
- Error decomposition, not just a total
- Substitutions, insertions and deletions are reported separately, because a pipeline whose errors are inserted background chatter is in different trouble from one that is dropping the word “not”.
- A one-word difference is not a result
- Each condition carries roughly 190 reference words, so a single word is about half a point of WER. Deltas smaller than that are reported as numbers and described as noise, never as wins — including when they would flatter us.
- Cost is reported next to accuracy
- Real-time factor and algorithmic latency sit in the same table as WER, with the host machine recorded. An engine that wins on accuracy and misses the frame deadline has not won.
- The proprietary baseline never ships
- ai-coustics rows are run through their licensed SDK as a reference point. They are a measuring stick, never a code path in anything we sell.
git clone https://github.com/anecho-official/nulltest
cd nulltest
python3 -m pip install -e .
# fetch the open dataset + model weights
python3 scripts/download_models.py
# one cell
python3 -m engine.bench.run \
--engine gtcrn \
--condition telephony \
--stt faster-whisper:base.en
# the whole matrix
python3 -m engine.bench.run --matrix full \
--out apps/web/public/benchmark.jsonAdding a competitor is one file: implement the four-method Enhancer protocol and the harness picks it up. If you think we have measured your engine unfairly, the fastest way to prove it is a pull request — and we will publish the corrected row.
Or skip the harness and take the numbers: the complete results are one JSON document at a stable, versioned URL, published CC-BY-4.0./data/nulltest/v1/results.json
What is on the bench, and what you are allowed to ship.
Licence is a first-class column because the engines that come closest cannot legally go into your product.
| Engine | Licence | Shippable | RTF | Latency | Pooled WER |
|---|---|---|---|---|---|
| Raw (no processing) | — | permitted | — | 0 ms | 14.77% |
| Passthrough (control) | — | permitted | 0.000 | 0 ms | 14.77% |
| ai-coustics Quail L | proprietary | benchmark only | 0.124 | 30 ms | 15.56% |
| ai-coustics Quail VF 2.2 L | proprietary | benchmark only | 0.076 | 30 ms | 18.10% |
| GTCRN | MIT | permitted | 0.050 | 16 ms | 20.26% |
| FastEnhancer-L | MIT | permitted | 0.376 | 26 ms | 20.44% |
| FastEnhancer-S | MIT | permitted | 0.034 | 16 ms | 20.64% |
| FastEnhancer-B | MIT | permitted | 0.021 | 16 ms | 21.14% |
| FastEnhancer-M | MIT | permitted | 0.110 | 22 ms | 22.40% |
| FastEnhancer-T | MIT | permitted | 0.013 | 16 ms | 22.46% |
| DeepFilterNet3 (96-frame warm-up) | MIT OR Apache-2.0 | permitted | 0.226 | 100 ms | 31.64% |
| DeepFilterNet3 | MIT OR Apache-2.0 | permitted | 0.133 | 100 ms | 35.18% |
| Clearline (8 kHz, resampled) | internal | benchmark only | 0.048 | 16 ms | 39.65% |
The telephony column
No enhancer meaningfully improved 8 kHz audio and most made it worse. Nothing on the market handles the phone line properly — which is exactly the gap Clearline exists to close.
Read itWhy the published claims disagree
The five hidden variables that flip the sign of the result, and why every vendor figure so far leaves at least three of them unstated.
Read itSee one cell up close
The comparator shows a single condition end to end: audio, spectrum, transcript diff, error decomposition.
Open the comparator