# Anecho — full site content > Anecho Audio, Inc. builds the audio path for telephony voice agents: native 8 kHz speech enhancement, residual echo suppression and provably correct resampling. We also publish Null Test, an open benchmark measuring what speech enhancement does to downstream speech-to-text accuracy — the question no vendor in this category answers with data. Measured 2026-08-14 on 13 backends x 18 acoustic conditions with faster-whisper:base.en: no engine beat the unprocessed control on pooled word error rate (raw 14.8%, best enhancer 15.6%), but at least one engine did beat raw in 11 of the 18 individual conditions. The correct reading is that enhancement pays where the audio is genuinely bad and costs you accuracy where it is not — not that enhancement never helps. Insertions went up rather than down (101 on raw, 410 on the worst engine), which is the opposite of the usual vendor claim. If you quote these numbers, please carry the caveats with them: a single recogniser, roughly 190 reference words per condition, and no natively-8 kHz engine in this matrix. They are listed at https://anecho.ai/benchmark/methodology. Canonical: https://anecho.ai/llms-full.txt · Index: https://anecho.ai/llms.txt · Generated at build time from `content/blog/*.mdx`, `public/benchmark.json`, `src/lib/methodology.ts`, `src/lib/pricing.ts` and `src/lib/snippets.ts`. Nothing in this file is hand-maintained. --- ## Null Test — the open speech enhancement benchmark Canonical page: https://anecho.ai/benchmark No engine beat the unprocessed control on pooled word error rate: raw scored 14.8% against 15.6% for the best enhancer (ai-coustics Quail L) and 20.3% for the best engine with a licence you can ship (GTCRN). That single pooled number is not the whole result, and quoting it alone is a misreading: at least one engine did beat raw in 11 of the 18 individual conditions, with the largest single-condition gain −5.79 points (ai-coustics Quail VF 2.2 L on Competing speaker 5 dB). The only conditions where nothing helped were Reverb and Reverb + noise 10 dB and telephony_alaw and telephony_g722 and telephony_opus12k and telephony_loss3 and telephony_loss10 — reverberation, which none of these models remove. The correct summary is therefore: enhancement pays where the audio is genuinely bad and costs you accuracy where it is not, and pooling across a condition mix that is mostly not-that-bad hides the trade. Insertions — words the recogniser invented that nobody said — went up rather than down, from 101 on raw audio to 410 on Clearline (8 kHz, resampled), which is the opposite of the usual vendor claim. ### Pooled results, all conditions combined | Engine | Licence | Shippable | WER | ΔWER vs raw (pp) | Ins | Sub | Del | SI-SDR (dB) | PESQ-WB | STOI | VAD F1 | VAD FA | RTF | p99 block (ms) | Latency (ms) | |---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| | Raw (no processing) | — | yes | 14.77% | ±0.00 | 101 | 344 | 60 | 10.92 | 1.818 | 0.8437 | 0.9504 | 0.1716 | — | — | 0.0 | | Passthrough (control) | — | yes | 14.77% | ±0.00 | 101 | 344 | 60 | 10.92 | 1.818 | 0.8437 | 0.9504 | 0.1716 | 0.0001 | 0.01 | 0.0 | | ai-coustics Quail L | proprietary | benchmark only | 15.56% | +0.79 | 126 | 349 | 57 | 8.55 | 2.034 | 0.8778 | 0.9449 | 0.1867 | 0.1237 | 12.94 | 30.0 | | ai-coustics Quail VF 2.2 L | proprietary | benchmark only | 18.10% | +3.33 | 161 | 383 | 75 | 5.02 | 1.890 | 0.8673 | 0.9451 | 0.1529 | 0.0755 | 6.33 | 30.0 | | GTCRN | MIT | yes | 20.26% | +5.50 | 210 | 443 | 40 | 8.79 | 2.091 | 0.8570 | 0.9472 | 0.1638 | 0.0498 | 4.66 | 16.0 | | FastEnhancer-L | MIT | yes | 20.44% | +5.67 | 159 | 433 | 107 | 10.60 | 2.198 | 0.8504 | 0.9417 | 0.1474 | 0.3762 | 4.93 | 25.8 | | FastEnhancer-S | MIT | yes | 20.64% | +5.88 | 159 | 475 | 72 | 10.23 | 2.205 | 0.8539 | 0.9452 | 0.1618 | 0.0336 | 1.04 | 16.0 | | FastEnhancer-B | MIT | yes | 21.14% | +6.37 | 159 | 485 | 79 | 9.84 | 2.164 | 0.8479 | 0.9442 | 0.1617 | 0.0212 | 5.44 | 16.0 | | FastEnhancer-M | MIT | yes | 22.40% | +7.63 | 217 | 466 | 83 | 10.63 | 2.121 | 0.8488 | 0.9406 | 0.1491 | 0.1102 | 3.21 | 22.0 | | FastEnhancer-T | MIT | yes | 22.46% | +7.69 | 145 | 514 | 109 | 9.16 | 2.037 | 0.8348 | 0.9411 | 0.1577 | 0.0127 | 1.59 | 16.0 | | DeepFilterNet3 (96-frame warm-up) | MIT OR Apache-2.0 | yes | 31.64% | +16.87 | 368 | 585 | 129 | 8.32 | 1.960 | 0.8345 | 0.9142 | 0.2275 | 0.2259 | 20.75 | 100.0 | | DeepFilterNet3 | MIT OR Apache-2.0 | yes | 35.18% | +20.41 | 380 | 626 | 197 | 7.58 | 1.869 | 0.8166 | 0.9013 | 0.2219 | 0.1330 | 13.65 | 100.0 | | Clearline (8 kHz, resampled) | internal | benchmark only | 39.65% | +24.88 | 410 | 776 | 170 | 0.69 | 1.685 | 0.7439 | 0.8352 | 0.1513 | 0.0481 | 4.08 | 16.0 | ### Word error rate by condition | Condition | Raw (no processing) | Passthrough (control) | ai-coustics Quail L | ai-coustics Quail VF 2.2 L | GTCRN | FastEnhancer-L | FastEnhancer-S | FastEnhancer-B | FastEnhancer-M | FastEnhancer-T | DeepFilterNet3 (96-frame warm-up) | DeepFilterNet3 | Clearline (8 kHz, resampled) | |---|---|---|---|---|---|---|---|---|---|---|---|---|---| | Clean | 4.74% | 4.74% | 6.32% | 5.79% | 6.32% | 5.26% | 5.79% | 5.79% | 4.21% | 6.32% | 5.26% | 4.74% | 5.79% | | Noise -5 dB | 18.42% | 18.42% | 14.21% | 24.74% | 19.47% | 13.68% | 22.63% | 18.42% | 15.26% | 24.21% | 34.21% | 32.11% | 63.68% | | Noise 0 dB | 11.05% | 11.05% | 6.32% | 10.00% | 8.42% | 7.89% | 8.95% | 7.37% | 7.89% | 10.53% | 14.74% | 13.68% | 28.95% | | Noise 5 dB | 8.95% | 8.95% | 11.58% | 9.47% | 10.00% | 6.84% | 8.42% | 6.84% | 9.47% | 6.84% | 14.21% | 14.21% | 33.16% | | Noise 10 dB | 6.32% | 6.32% | 6.32% | 5.79% | 8.42% | 5.26% | 7.37% | 5.79% | 5.79% | 5.26% | 6.84% | 10.00% | 10.53% | | Noise 20 dB | 4.21% | 4.21% | 5.79% | 5.79% | 6.32% | 3.68% | 5.26% | 5.26% | 3.16% | 4.21% | 4.21% | 4.74% | 6.84% | | Babble 5 dB | 24.74% | 24.74% | 27.37% | 20.53% | 61.05% | 36.84% | 37.37% | 36.84% | 30.00% | 38.95% | 48.42% | 35.79% | 138.42% | | Competing speaker 0 dB | 84.74% | 84.74% | 79.47% | 116.84% | 87.89% | 84.21% | 86.32% | 92.63% | 96.32% | 86.84% | 143.16% | 199.47% | 106.84% | | Competing speaker 5 dB | 26.32% | 26.32% | 36.84% | 20.53% | 44.74% | 76.32% | 51.05% | 50.53% | 46.84% | 46.32% | 75.26% | 50.00% | 87.37% | | Reverb | 5.79% | 5.79% | 6.84% | 15.26% | 6.84% | 17.37% | 22.63% | 25.79% | 27.37% | 21.05% | 43.16% | 60.00% | 12.11% | | Reverb + noise 10 dB | 13.68% | 13.68% | 15.79% | 23.16% | 14.74% | 22.63% | 25.79% | 34.21% | 22.11% | 28.95% | 60.00% | 82.11% | 116.32% | | Telephony 8 kHz | 5.79% | 5.79% | 7.37% | 8.42% | 8.95% | 7.37% | 5.26% | 6.84% | 5.79% | 12.63% | 15.79% | 15.26% | 5.79% | | Telephony + noise 10 dB | 10.53% | 10.53% | 12.11% | 9.47% | 11.05% | 13.68% | 14.21% | 13.68% | 12.11% | 17.37% | 16.32% | 20.00% | 14.74% | | telephony_alaw | 7.37% | 7.37% | 7.37% | 10.53% | 13.68% | 12.11% | 16.32% | 16.84% | 48.95% | 21.05% | 17.37% | 18.42% | 13.68% | | telephony_g722 | 5.26% | 5.26% | 6.84% | 6.84% | 5.79% | 5.26% | 5.79% | 5.26% | 6.84% | 5.79% | 6.84% | 6.32% | 9.47% | | telephony_opus12k | 6.84% | 6.84% | 7.37% | 9.47% | 12.11% | 9.47% | 6.84% | 8.42% | 8.42% | 8.95% | 10.00% | 11.05% | 27.37% | | telephony_loss3 | 11.05% | 11.05% | 11.58% | 11.05% | 14.21% | 20.00% | 16.84% | 16.84% | 20.00% | 23.68% | 20.00% | 24.74% | 17.89% | | telephony_loss10 | 10.00% | 10.00% | 10.53% | 12.11% | 24.74% | 20.00% | 24.74% | 23.16% | 32.63% | 35.26% | 33.68% | 30.53% | 14.74% | ### Best engine per condition, against the unprocessed control | Condition | Raw WER | Best engine | Best WER | Δ (pp) | Engines below raw | Verdict | |---|---|---|---|---|---|---| | Clean | 4.74% | FastEnhancer-M | 4.21% | -0.53 | 1 of 11 | enhancement helped | | Noise -5 dB | 18.42% | FastEnhancer-L | 13.68% | -4.74 | 3 of 11 | enhancement helped | | Noise 0 dB | 11.05% | ai-coustics Quail L | 6.32% | -4.74 | 8 of 11 | enhancement helped | | Noise 5 dB | 8.95% | FastEnhancer-L | 6.84% | -2.11 | 4 of 11 | enhancement helped | | Noise 10 dB | 6.32% | FastEnhancer-L | 5.26% | -1.05 | 5 of 11 | enhancement helped | | Noise 20 dB | 4.21% | FastEnhancer-M | 3.16% | -1.05 | 2 of 11 | enhancement helped | | Babble 5 dB | 24.74% | ai-coustics Quail VF 2.2 L | 20.53% | -4.21 | 1 of 11 | enhancement helped | | Competing speaker 0 dB | 84.74% | ai-coustics Quail L | 79.47% | -5.26 | 2 of 11 | enhancement helped | | Competing speaker 5 dB | 26.32% | ai-coustics Quail VF 2.2 L | 20.53% | -5.79 | 1 of 11 | enhancement helped | | Reverb | 5.79% | ai-coustics Quail L | 6.84% | +1.05 | 0 of 11 | nothing beat raw | | Reverb + noise 10 dB | 13.68% | GTCRN | 14.74% | +1.05 | 0 of 11 | nothing beat raw | | Telephony 8 kHz | 5.79% | FastEnhancer-S | 5.26% | -0.53 | 1 of 11 | enhancement helped | | Telephony + noise 10 dB | 10.53% | ai-coustics Quail VF 2.2 L | 9.47% | -1.05 | 1 of 11 | enhancement helped | | telephony_alaw | 7.37% | ai-coustics Quail L | 7.37% | ±0.00 | 0 of 11 | nothing beat raw | | telephony_g722 | 5.26% | FastEnhancer-L | 5.26% | ±0.00 | 0 of 11 | nothing beat raw | | telephony_opus12k | 6.84% | FastEnhancer-S | 6.84% | ±0.00 | 0 of 11 | nothing beat raw | | telephony_loss3 | 11.05% | ai-coustics Quail VF 2.2 L | 11.05% | ±0.00 | 0 of 11 | nothing beat raw | | telephony_loss10 | 10.00% | ai-coustics Quail L | 10.53% | +0.53 | 0 of 11 | nothing beat raw | ### Split pipeline: the same stream scored for STT and for VAD | Engine | Mean WER enhanced | Mean WER raw | ΔWER (pp) | Conditions where enhancement helped STT | VAD F1 enhanced | VAD F1 raw | Insertions enhanced | Insertions raw | |---|---|---|---|---|---|---|---|---| | Passthrough (control) | 14.77% | 14.77% | ±0.00 | 0 of 18 | 0.9504 | 0.9504 | 101 | 101 | | GTCRN | 20.26% | 14.77% | +5.50 | 1 of 18 | 0.9472 | 0.9504 | 210 | 101 | | FastEnhancer-T | 22.46% | 14.77% | +7.69 | 3 of 18 | 0.9411 | 0.9504 | 145 | 101 | | FastEnhancer-B | 21.14% | 14.77% | +6.37 | 3 of 18 | 0.9442 | 0.9504 | 159 | 101 | | FastEnhancer-S | 20.64% | 14.77% | +5.88 | 3 of 18 | 0.9452 | 0.9504 | 159 | 101 | | FastEnhancer-M | 22.40% | 14.77% | +7.63 | 5 of 18 | 0.9406 | 0.9504 | 217 | 101 | | FastEnhancer-L | 20.44% | 14.77% | +5.67 | 6 of 18 | 0.9417 | 0.9504 | 159 | 101 | | DeepFilterNet3 | 35.18% | 14.77% | +20.41 | 0 of 18 | 0.9013 | 0.9504 | 380 | 101 | | DeepFilterNet3 (96-frame warm-up) | 31.64% | 14.77% | +16.87 | 0 of 18 | 0.9142 | 0.9504 | 368 | 101 | | ai-coustics Quail VF 2.2 L | 18.10% | 14.77% | +3.33 | 5 of 18 | 0.9451 | 0.9504 | 161 | 101 | | ai-coustics Quail L | 15.56% | 14.77% | +0.79 | 3 of 18 | 0.9449 | 0.9504 | 126 | 101 | | Resample-only 16→8→16 (control) | 18.63% | 14.77% | +3.86 | 3 of 18 | 0.9478 | 0.9504 | 173 | 101 | | Clearline (8 kHz, resampled) | 39.65% | 14.77% | +24.88 | 0 of 18 | 0.8352 | 0.9504 | 410 | 101 | ### Run configuration | Field | Value | |---|---| | Generated | 2026-08-14T02:16:50Z | | Schema | v1 | | Recogniser | CTranslate2 int8 CPU, greedy, model=base.en | | Voice activity | silero-vad, threshold 0.5, 512-sample blocks | | Test set | v1, 10 speakers, 180 clips, 1406 s, seed 20260101 | | Reference words | 3,420 | | Host | Apple M1 Max, macOS-26.5.2-arm64-arm-64bit | | Runtime | onnxruntime 1.28.0, 1 intra-op thread | | Test material licence | LibriSpeech is CC-BY-4.0. ESC-50 is CC-BY-NC-3.0 — the noise conditions are research/evaluation only. The competing-speaker and babble conditions are built from LibriSpeech alone and carry no NC restriction. | --- ## Null Test methodology and caveats Canonical page: https://anecho.ai/benchmark/methodology ### What was run One harness, one recogniser, one machine, every engine treated identically. 13 audio backends — 11 enhancement engines plus an unprocessed control and a passthrough sanity row — were scored across 18 acoustic conditions on a test set of 180 clips (10 speakers, 1406 seconds of audio, seed 20260101). Transcription is CTranslate2 int8 CPU, greedy, model=base.en. Voice activity is silero-vad at threshold 0.5 on 512-sample blocks. Cost is measured on Apple M1 Max (macOS-26.5.2-arm64-arm-64bit) with onnxruntime 1.28.0 pinned to 1 intra-op thread and sequential execution, so every real-time factor here is one stream on one core, not a fan-out across the machine. The control is a row, not a footnote. Unprocessed audio appears in every table and every delta is measured against it in the same condition. A benchmark without a null hypothesis is a brochure. - Clean - Noise -5 dB - Noise 0 dB - Noise 5 dB - Noise 10 dB - Noise 20 dB - Babble 5 dB - Competing speaker 0 dB - Competing speaker 5 dB - Reverb - Reverb + noise 10 dB - Telephony 8 kHz - Telephony + noise 10 dB - telephony_alaw - telephony_g722 - telephony_opus12k - telephony_loss3 - telephony_loss10 ### How to quote this without misquoting it If you quote one number from this benchmark, quote it with the condition attached. Pooled WER across a fixed condition mix is a property of the mix as much as of the engine: weight the mix toward clean and lightly-noisy speech and enhancement looks like a tax; weight it toward competing speakers and low SNR and enhancement looks like a win. Both readings come from the same file. The specific misreading we would like to head off is "enhancement never helps". That is not what 18 conditions showed. No engine beat raw on the pooled average; at least one engine beat raw in 11 of the 18 individual conditions. Both sentences are true, and the second one is the actionable half — it tells you the conditions in which turning enhancement on is a good trade. The mirror-image misreading is also wrong. Enhancement is not a free win in the conditions where it helped: it moved word error rate by a few points on a test set whose resolution is about half a point per word, on one recogniser, and it raised insertions everywhere. Measure it on your own audio and your own recogniser before you commit — the SDK ships the same alignment code this harness uses, precisely so you can. ### Where enhancement paid, condition by condition At least one engine beat the unprocessed control in 11 of 18 conditions. Read these against the resolution limit below before treating any of them as a result: several are inside one or two reference words. - Clean: raw 4.7% → 4.2% (−0.53 points, FastEnhancer-M); 1 of 11 engines beat raw here. - Noise -5 dB: raw 18.4% → 13.7% (−4.74 points, FastEnhancer-L); 3 of 11 engines beat raw here. - Noise 0 dB: raw 11.1% → 6.3% (−4.74 points, ai-coustics Quail L); 8 of 11 engines beat raw here. - Noise 5 dB: raw 8.9% → 6.8% (−2.11 points, FastEnhancer-L); 4 of 11 engines beat raw here. - Noise 10 dB: raw 6.3% → 5.3% (−1.05 points, FastEnhancer-L); 5 of 11 engines beat raw here. - Noise 20 dB: raw 4.2% → 3.2% (−1.05 points, FastEnhancer-M); 2 of 11 engines beat raw here. - Babble 5 dB: raw 24.7% → 20.5% (−4.21 points, ai-coustics Quail VF 2.2 L); 1 of 11 engines beat raw here. - Competing speaker 0 dB: raw 84.7% → 79.5% (−5.26 points, ai-coustics Quail L); 2 of 11 engines beat raw here. - Competing speaker 5 dB: raw 26.3% → 20.5% (−5.79 points, ai-coustics Quail VF 2.2 L); 1 of 11 engines beat raw here. - Telephony 8 kHz: raw 5.8% → 5.3% (−0.53 points, FastEnhancer-S); 1 of 11 engines beat raw here. - Telephony + noise 10 dB: raw 10.5% → 9.5% (−1.05 points, ai-coustics Quail VF 2.2 L); 1 of 11 engines beat raw here. ### Where nothing helped Reverberation is the clean failure case. None of these models are dereverberation models, and on the two reverberant conditions the best engine still lost to doing nothing. SI-SDR goes sharply negative there — down to around −17 dB — which is the metric telling you the processed signal has been moved a long way from the reference, not merely denoised. - Reverb: raw 5.8%, best engine 6.8% (+1.05 points, ai-coustics Quail L). Nothing beat doing nothing. - Reverb + noise 10 dB: raw 13.7%, best engine 14.7% (+1.05 points, GTCRN). Nothing beat doing nothing. - telephony_alaw: raw 7.4%, best engine 7.4% (±0.00 points, ai-coustics Quail L). Nothing beat doing nothing. - telephony_g722: raw 5.3%, best engine 5.3% (±0.00 points, FastEnhancer-L). Nothing beat doing nothing. - telephony_opus12k: raw 6.8%, best engine 6.8% (±0.00 points, FastEnhancer-S). Nothing beat doing nothing. - telephony_loss3: raw 11.1%, best engine 11.1% (±0.00 points, ai-coustics Quail VF 2.2 L). Nothing beat doing nothing. - telephony_loss10: raw 10.0%, best engine 10.5% (+0.53 points, ai-coustics Quail L). Nothing beat doing nothing. ### The resolution limit — what counts as a result Each condition carries roughly 190 reference words (3,420 across all 18 conditions), so one word is about 0.53 points of word error rate. Any per-condition delta smaller than that is a single word and is not a finding — including when it flatters us. This is why the telephony column is reported as "no meaningful improvement" rather than as a win: the best engine there lands −0.53 points from the control, which is inside one reference word, while the worst costs +10.00 points, which is not. The test set is deliberately small enough to say so plainly: 10 speakers is enough to separate large effects and not enough to resolve a point or two of WER. Treat rank orderings between adjacent engines as unresolved. ### The caveats that must travel with these numbers If you cite a figure from this benchmark, cite these with it. They are the harness's own statement of what it does not prove, published in the results file itself and reproduced here unedited. Every backend is scored at its own native rate, so a narrowband engine is resampled 16→8→16 to meet this matrix. To keep that handicap accountable rather than rhetorical, the table carries a resample-only control — the same round trip with no model in it. On the Telephony 8 kHz rows the control scores 5.8%, which makes it an easy condition, and easy conditions are exactly where enhancement has the least to win and the most to lose. Anecho's own engine is now in this table, and it loses. Pooled it scores well behind doing nothing, and the control is what makes that readable: on the wideband half the band limit costs about six points of word error rate and our model adds roughly thirty more, while on the narrowband half the band limit costs under a point and our model still adds six. It loses at its own operating point, to raw audio and to a piece of wire — the one exception being plain Telephony 8 kHz, where it lands on the control and ties doing nothing. We publish it because a benchmark that exempts its author measures nothing. - DeepFilterNet3 is measured through a block-online adapter: every public ONNX export of it lacks recurrent state tensors, so per-frame streaming is impossible with the published weights. Its RTF here is an upper bound and its latency (100 ms) is far above the ~40 ms a native (Rust/tract) runtime achieves. Read the DFN3 rows as a lower bound on quality and an upper bound on cost. - ai-coustics rows are a proprietary baseline run through the licensed aic-sdk. They are a reference point, never a shipping code path. - WER is measured with faster-whisper:base.en (CTranslate2 int8 CPU, greedy, model=base.en). A different recogniser will give different absolute numbers; the harness supports Deepgram and AssemblyAI so the *ordering* can be checked on another engine. - Clearline is an 8 kHz model and these conditions are stored at 16 kHz, so its row is measured through the harness's resample path — 16 k -> 8 k -> model -> 8 k -> 16 k (soxr), the same path the native-8 kHz ai-coustics baselines take. On `clean`, `noise_*`, `babble_snr5`, `competing_*` and `reverb*` that is not its operating point: the 4-8 kHz band is discarded before the model ever sees it. The `resample8k` control row is that round trip with no model in it, so the band limit and the model can be told apart; and the telephony sub-aggregate below reports the regime the model was actually built for. - The Clearline checkpoint benchmarked here is not shippable, and this row is not a claim that it is. Two gates it does not pass, both measured on its own eval (`runs/t4_300h_200ep_solo`, 1,800-clip test_v2) and neither of them visible in this table: (1) a solo regression — on clips with no competing speaker the model makes the audio worse, -0.874 dB SI-SDR with 37.1% of solo clips degraded (-0.595 dB / 33.0% on val_v2); (2) speaker conditioning contributes approximately nothing — the matched ablation scores FiLM +2.815 dB against no-conditioning +2.690 dB, a 0.125 dB difference against a 0.3 dB noise floor, so what is being measured here is a plain denoiser and not the speaker isolation the architecture is for. The row exists so Clearline can be *placed* next to the competition on identical audio, not to argue that it is ready. - Clearline's own eval (`clearline/evaluate.py`, 1,800 clips, difficulty buckets) and this leaderboard are not comparable and were never meant to be: different clips, different conditions, different pooling. Only the recogniser is shared. This row exists precisely so that a comparable number exists. - DeepFilterNet3 runs at 48 kHz on 16 kHz source material upsampled to 48 kHz — there is no real content above 8 kHz for it to work with. That is the honest telephony/VoIP situation, but it is not the condition DFN3 was designed for. - The test set is 10 speakers / ~190 reference words per condition. That is enough to separate large effects and not enough to resolve differences of a point or two of WER. - Run-to-run repeatability of the recogniser was measured, not assumed — and the standing warning did not reproduce here. Elsewhere in this project faster-whisper on CPU has been seen to return different transcripts for byte-identical input (three runs of one aggregate gave 121.65 / 122.15 / 121.40%), and the standing guidance from that is to treat differences below roughly 6 pp on a small aggregate as decoder noise. On this host and this harness it did not happen: 25/25 scopes were bit-identical across repeated decodes (raw on 18 condition(s), 180 clips x 3 decodes; clearline on 5 condition(s), 50 clips x 3 decodes), including the arms where the recogniser is already hallucinating past 100% WER. Independently, re-measuring `raw` and `gtcrn` in a fresh process two days after the published run reproduced all 54 rows exactly (max ΔWER 0.0000 pp). Three repeats cannot prove determinism and a different host, venv or thread count may well behave differently, so the conservative 6 pp guidance still governs anything quoted from another machine — but on this table the resolution limit is set by the size of the test set, not by the decoder: one reference word is 0.53 pp in a single condition (190 words), and about 0.03 pp pooled across all 18. Differences of a point or two between adjacent rows are still not results; differences of tens of points are. - onnxruntime is pinned to 1 intra-op thread with sequential execution: the measured cost is one stream on one core, not a fan-out across the whole machine. - The `telephony*` conditions in this table are stored at 16 kHz after a full 8 kHz carrier leg (band-limit, codec, packet loss), so every backend sees them at its own native rate. The genuinely narrowband experiment — where nothing is resampled — is a separate artefact (`bench/out/telephony.json`). - Packet loss is modelled as bursty (2-state Gilbert-Elliott) erasure of 20 ms frames with attenuated-repeat concealment, not i.i.d. loss with silence fill. i.i.d. loss flatters every decoder and silence fill overstates the damage. ### Licensing, reuse and citation The measured results — everything in the results JSON — are published under CC-BY-4.0. Quote them, republish them, argue with them; attribution to Null Test by Anecho Audio, Inc. is requested, and a link to the methodology is appreciated because it carries the caveats. The harness is Apache-2.0. The *derived audio* is a separate question and a stricter one: LibriSpeech is CC-BY-4.0. ESC-50 is CC-BY-NC-3.0 — the noise conditions are research/evaluation only. The competing-speaker and babble conditions are built from LibriSpeech alone and carry no NC restriction. The ai-coustics rows are produced through their licensed SDK as a proprietary reference point. They are a measuring stick and never a shipping code path in anything we sell. If you think we have measured your engine unfairly, the harness takes a new backend in one file — implement the four-method Enhancer protocol and it is picked up. Send a pull request and we will publish the corrected row. ### Questions this benchmark answers **Does noise suppression hurt speech-to-text accuracy?** In our measurement it depends entirely on the condition, and the honest answer is not a yes or a no. Pooled across 18 conditions, no engine beat unprocessed audio: raw scored 14.8% word error rate against 15.6% for the best enhancer. But per condition, at least one engine beat raw in 11 of 18 cases — the exceptions were both reverberant. Enhancement pays where the audio is genuinely bad (low SNR, competing speakers) and costs you accuracy where it is not. Measured with faster-whisper:base.en; a different recogniser will give different absolute numbers. **Which speech enhancement engine had the lowest word error rate?** ai-coustics Quail L at 15.6% pooled, but it is a proprietary baseline we do not ship. The best engine under a licence you can actually use in a product was GTCRN at 20.3%. Both are above the unprocessed control's 14.8%. Ranking between adjacent engines is inside the resolution of a 10-speaker test set and should not be treated as settled. **Does enhancement reduce hallucinated words in transcription?** No — it increased them here. Insertions rose from 101 on raw audio to 410 on Clearline (8 kHz, resampled). That is the direct opposite of the common vendor claim, and it is why this benchmark reports insertions, substitutions and deletions separately rather than only a WER total. **Should I run enhanced audio into my speech-to-text?** Not by default, and not with the same stream you feed your turn detector. Our data says the two consumers want different audio: enhancement cost word error rate on pooled average while cutting the voice-activity false-alarm rate by up to 29 points in the conditions that actually break barge-in (babble and competing speakers). The split pipeline — enhanced audio to VAD and turn-taking, original audio to the transcriber — is the routing that follows from the measurement. **Is DeepFilterNet3 really that bad?** No, and the rows should be read as a lower bound on its quality and an upper bound on its cost. Every public ONNX export of DeepFilterNet3 lacks recurrent state tensors, so per-frame streaming is impossible with the published weights and it has to run through a block-online adapter at 100 ms latency. Its native Rust runtime would be roughly 40 ms. It is also a 48 kHz model being handed 16 kHz source material upsampled to 48 kHz, which is the honest VoIP situation but not what it was designed for. **Can I reuse these results?** Yes. The results file is CC-BY-4.0 and the harness is Apache-2.0. Attribution to Null Test by Anecho Audio, Inc. is requested. Please carry the caveats with the numbers — particularly the single-recogniser caveat and the ~190-reference-word resolution limit, because most of the per-condition deltas that look interesting are one or two words wide. --- ## Writing ### Amazon Connect, Twilio, NICE CXone and Avaya: four media paths, two places to stand Canonical: https://anecho.ai/blog/where-you-can-insert-audio-processing · Anecho Engineering · published 2026-08-16 · 3715 words · tags: twilio, amazon-connect, nice-cxone, avaya, integration We read the published specifications for the real-time audio surfaces of four contact-centre platforms and asked one question of each: can a third party receive the audio, process it, and have the processed audio be what the platform's own downstream consumers hear? Twilio and Avaya document yes. Amazon Connect and NICE CXone document no. Along the way: Amazon publishes 8 kHz and 'raw PCM' and nothing else, Avaya is the only one of the five documenting a wideband codec, and Twilio never says the thing we have been saying they say. *Every specification claim below is quoted from the vendor's own published documentation and linked. Where we could not confirm something it is marked unconfirmed rather than filled in, and there is a consolidated list near the end. This post also sharpens something we have previously stated too strongly about Twilio; that correction is in its own section. If any of this is wrong or changes, tell us and we will correct it with a date.* The companion post reads [Genesys Cloud AudioHook](/blog/genesys-audiohook-8khz) the same way and finds a protocol that pins its sample rate in the type system and a media path with no insertion point in it. This one covers the other four surfaces our buyers actually run on, and the answer is less uniform than we expected. One question, asked of each platform: > Can a third party receive the call audio, process it, and have the **processed** audio be what the platform's own bot, recogniser, recording or far end consumes? | Platform | Surface | Audio out to you | Processed audio back in | Documented rate | |---|---|---|---|---| | Amazon Connect | Kinesis Video Streams | Yes | **No** | 8 kHz | | Twilio | `` | Yes | **Yes** | 8000 Hz µ-law | | Twilio | `` | No — JSON only | No | not published | | NICE CXone | Custom Agent Assist | Yes | **No** | 8 kHz µ-law | | NICE CXone | Custom Virtual Agent | Yes | Turn-based playback only | G.711, rate not stated | | Avaya Infinity | Real-time Contextual Media Streaming | Yes | **Yes** | 8 kHz, or **16 kHz on G.722** | Two of these are genuine man-in-the-middle positions. The rest are taps, forks or bot turns wearing the same vocabulary. #### Amazon Connect: a one-way fork, and two things everyone gets wrong **It is not started by an API.** `StartContactStreaming` is the obvious candidate and it is the wrong one — [its own reference](https://docs.aws.amazon.com/connect/latest/APIReference/API_StartContactStreaming.html) says it *"initiates real-time message streaming for a new chat contact"*, and its only configuration object is a chat streaming config pointing at an SNS endpoint. It is a chat API. Voice media streaming is enabled at the instance and then turned on inside a flow: > "After you enable live media streaming, add **Start media streaming** and **Stop media streaming** blocks to your flow." — [enable live media streams](https://docs.aws.amazon.com/connect/latest/adminguide/enable-live-media-streams.html) The block offers two options — stream from the customer, or to the customer — and applies to voice only; every other channel takes the error branch. As far as we can find, **there is no public API to start voice media streaming**, only the flow block. **The format is 8 kHz, and almost nothing else is published.** This is the sentence to quote: > "Media streaming uses Kinesis Video Streams multi-track support so that what the customer says is on a separate track from what the customer hears. **Audio sent to Kinesis uses a sampling rate of 8 kHz.**" — [plan live media streams](https://docs.aws.amazon.com/connect/latest/adminguide/plan-live-media-streams.html) The tracks are `AUDIO_FROM_CUSTOMER` and `AUDIO_TO_CUSTOMER`, and the consumer guide says a reader *"stores this data as a raw PCM file"*. That is the whole documented format. We grepped the four live-media-streaming pages for `L16`, `endian`, `bit` and `codec` and got nothing. **The commonly repeated "audio/L16, 16-bit, little-endian, mono" is not in the Amazon Connect documentation.** The only bit-depth signal anywhere in AWS's own material is an Audacity import instruction in a demo repository — *set encoding to signed 16-bit PCM and sample rate to 8000 Hz* — which tells you what the bytes turned out to be, not what the contract is. Sixteen-bit mono is strongly implied. It is not documented, and we are not going to write it down as though it were. Note also that Connect is the odd one out on encoding. Twilio and NICE both hand you companded G.711 µ-law; Connect hands you linear PCM at the same 8 kHz. Same band limit, different quantisation noise, and if you are benchmarking a model against one of them you are not benchmarking it against the other. **It is one-way, and AWS says so obliquely but unmistakably.** Every documented verb is capture or consume, there is no write path, and the planning page contains this: > "We recommend that you refrain from modifying the streams. Doing so can cause unexpected behavior." **It is a live stream you read fragments from, with a short default retention.** *"If you select No data retention, data is not retained and is available to be consumed for only 5 minutes."* Connect hands your consumer a fragment cursor through contact attributes — `StartFragmentNumber`, `StreamARN`, `StartTimestamp`, `StopTimestamp` — so your reader seeks into the stream rather than subscribing to a socket. That is an architectural difference worth planning for: it is a pull, not a push, and your latency budget includes whatever your consumer's polling loop costs. **What about the other Connect surfaces?** We checked them because the naming invites confusion: - **Contact Lens real-time** is transcripts, not audio. It delivers `Utterance` segments over Kinesis Data Streams — *"partial transcripts... to meet ultra-low latency requirements to assist agents on live calls"*. Text, downstream of a recogniser you do not control. - **External voice systems** ingest into Contact Lens over **SIPREC** — the configuration page names the host that *"will receive the SIPREC audio"*. This is audio flowing *into* Connect from a third-party telephony platform, not out of it for processing. The codec specification is deferred to Amazon Chime SDK guides that we did not read, so the format there is unconfirmed. - **`StartWebRTCContact`** returns Chime SDK meeting and attendee credentials — `AudioHostUrl`, `SignalingUrl`, `JoinToken`. That is how a client SDK joins the call as a participant. It is not a server-side raw-audio pipe. - **Third-party speech providers** is the one place Connect documents somebody else's model in the path: *"Connect routes audio to the chosen third-party speech-to-text provider"*, configured per bot locale with the provider's API key in Secrets Manager. Companion pages exist for third-party TTS. This is **vendor selection from a supported list**, not an insertion point for your own processing. One flag for anyone writing about Connect at the moment: the documentation table of contents now carries an end-of-support page for Amazon Connect Voice ID. Check its status before you build on it. **Verdict: no.** Kinesis Video Streams is a fork off to the side. You can listen to everything and change nothing. #### Twilio: the strongest general-purpose insertion point of the five Twilio publishes the tightest format contract of anyone here, and we verified these by reading the strings rather than trusting a summary. From the [WebSocket messages reference](https://www.twilio.com/docs/voice/media-streams/websocket-messages): > "`start.mediaFormat.encoding` — The encoding of the data in the upcoming payload. **Value is always `audio/x-mulaw`**. `start.mediaFormat.sampleRate` — **Value is always `8000`**. `start.mediaFormat.channels` — **Value is always `1`**." Three "always"es. Audio arrives as base64 inside JSON, with a `chunk` counter from 1 and a `timestamp` in milliseconds from stream start. The mechanism is a fork: *"Twilio forks the raw audio stream of the Call and streams it to your WebSocket server in near real-time"*, over `wss` only. Track selection — `inbound_track`, `outbound_track`, `both_tracks`, defaulting to `inbound_track` — belongs to the **unidirectional** form. Inbound is what Twilio receives from the other party; outbound is what Twilio generates toward the call. **The two forms are different products with one noun.** This is the distinction that decides your architecture: | | `` | `` | |---|---|---| | Direction | Fork out only | Bidirectional | | TwiML flow | *"immediately continues with the next TwiML instruction"* | Blocks: subsequent TwiML runs only *"after your server closes the WebSocket connection"* | | Created via REST | Yes | No | | Track selection | Yes | Scoped to unidirectional in the noun reference | For the bidirectional form, audio you send back **is played on the call**: > "Bidirectional Media Streams are those in which your WebSocket application both receives audio from Twilio and can send audio to Twilio, which is then played on the Call." — [Media Streams](https://www.twilio.com/docs/voice/media-streams) With a strict payload rule and a trap in it: *"The payload must be encoded `audio/x-mulaw` with a sample rate of `8000` and must be base64 encoded"*, and *"the `media.payload` should not contain audio file type header bytes. Providing header bytes causes the media to be streamed incorrectly."* If you generate µ-law with a library that writes a WAV header, you will hear it. You also get `mark` for playback-completion tracking and `clear` to flush buffered audio — which is the barge-in primitive, and the reason this surface can support a real interruption model rather than just talking over the caller. **Verdict: yes.** `` is a genuine man-in-the-middle: caller audio in, your audio out, played to the call, with playback and interrupt control. The cost is that `` is terminal — it owns the call for the duration — so you are not decorating a flow, you are becoming it. **One thing Twilio does not publish: the frame size.** We looked specifically, because everyone quotes 160 bytes and 20 ms. Twilio documents the `chunk` counter and the `timestamp` and gives no bytes-per-message or milliseconds-per-frame figure on either Media Streams page. If you need that number, measure it on your own traffic and label it as observed. Nor does Twilio publish a sample rate for real-time transcription. The `` noun documents track selection and labels and names Google or Deepgram as engines; we found no rate on the noun reference or the REST resource. #### The Twilio claim we have been making, stated correctly We have written, on this site, that Twilio's `` structurally cannot run a `` media fork alongside it. We went looking for the sentence that says so, and **there isn't one.** The correction matters more to us than the conclusion does, so here is exactly what the documentation does and does not contain. **What ConversationRelay is, confirmed.** It *"routes a call to the Conversation Relay service, providing advanced AI-powered voice interactions"*. Twilio performs the speech-to-text and text-to-speech itself, and your WebSocket exchanges **JSON only**: inbound `setup`, `prompt` (carrying `voicePrompt`, already-transcribed text), `dtmf`, `interrupt`, `error`; outbound text tokens, play-media by **URL**, send-digits, switch-language, end-session. No raw audio crosses that socket in either direction, and no codec or sample rate is published for it anywhere we could find. **What the documentation does not say, and we checked carefully:** - There is **no cardinality rule** on ``. The noun reference documents four nouns and two attributes and never states that `` may contain exactly one. - There is **no statement** that `` is incompatible with ``. We grepped the ConversationRelay overview, the noun reference, the onboarding page and the `` reference for "cannot", "not supported", "limitation", "incompatible", "only one", "simultaneously" and "alongside". Every hit was navigation chrome. No body text. There is no limitations section on any ConversationRelay page. - Whether a `` fork begun **before** a `` survives is **neither documented as working nor documented as failing.** Twilio simply never addresses the combination. **So the accurate statement is an inference, and we should have labelled it as one.** ConversationRelay's protocol carries no audio, and `` is documented as terminal, which together mean Twilio must be terminating and re-originating the media itself — leaving no documented place for a third-party processor between the caller and the recogniser. That reasoning is sound and it is not a citation. We have not run the empirical test of a preceding ``, and until we do, the honest form of the claim is: *Twilio does not document the combination; the architecture implies there is nowhere to insert a processor; we have not verified it.* If you have run that experiment either way, we would genuinely like the result, and we will publish it with attribution. #### NICE CXone: the documentation is not where you would look for it The public developer portal lists eighteen API families and none of them is live audio. "Real-Time Data" is metrics. "Media Playback" is recordings. Searching there leads you to conclude that CXone has no real-time audio surface, and that conclusion is wrong — the specification lives in the **administrator help centre**, under agent-assist integrations. Before the specifics, one negative finding worth having: we found no product called **CXone Real-Time Audio**, **RTMS**, **CXone Voice Streaming API** or **CXone Open Agent Assist**. Those names appear in third-party writing. They did not appear in a six-thousand-URL sweep of NICE's own documentation. Do not build a search on them. **Custom Agent Assist Integrations** is the real-time audio surface, and its format sentence is unambiguous: > "Audio packets are encoded as G711 μlaw 8-bit 8000 kHz raw audio. This is the same format as all NiCE CXone telephony audio." — [agent assist resources](https://help.nicecxone.com/content/aiassistantsandbots/customaiintegrations/aah/resources.htm) The unit is NICE's own typo; 8000 Hz is plainly meant. The second half of that sentence is the more useful half: **this is the same format as all CXone telephony audio.** The platform is telling you its internal media representation is narrowband G.711, everywhere. The channel model is per-connection rather than per-track, which is unusual and easy to get wrong: > "`streamPerspective: RX` — The audio being transmitted by the agent's phone: the agent talking. `TX` — The audio that the agent hears: the contact talking... `MIX` — Contains both agent and contact audio streams. **An individual websocket connection only contains audio from one perspective.**" So a stereo view of a call is **two sockets**, correlated by you, with no shared clock in the payload — because the payload has nothing in it but audio: > "**Only binary data flows through the webhook. For voice interactions, the only data sent from the call are audio bytes. No call control or other metadata is included.**" Operationally: at least 2,000 concurrent requests supported, no connection expiry or maximum duration, and a hold closes the socket while a resume opens a fresh one with an identical handshake. Note also that NICE permits an **unsecured** endpoint — *"The audio relay endpoint must be a websocket. It can be secured (WSS) or unsecured (WS)"* — which is a choice you should make deliberately rather than by copying a sample. **Direction: out only.** What you return is text and resources rendered to the agent's screen. Audio never re-enters the call. The Agent Assist Hub, which brokers these integrations, therefore streams **raw G.711 µ-law audio** to the third party and expects the third party to do its own speech-to-text. **Custom Virtual Agent** is the closest CXone comes to returning audio, and it is turn-based rather than streaming. Utterances arrive *"either as audio in the format of the G-711 codec or as transcribed text"*, and the response schema carries a field named — with NICE's own misspelling intact — `base64EndcodedG711ulawWithWavHeader`, described as the encoded WAV *"to be played at the next turn"*. Transport is REST over HTTPS. That is a bot reply, not an inline filter. **Verdict: no for agent assist, and only turn-based playback for virtual agent.** You can hear everything, at 8 kHz µ-law, on separate sockets per perspective, and you cannot change what anyone else hears. #### Avaya: the surprise, and the one wideband codec in the set We expected to write that Avaya's real-time audio surface is not publicly documented. It is — in detail, with no login, and it is the most explicitly bidirectional design of the five. **Avaya Infinity** publishes [Real-time Contextual Media Streaming](https://developers.avayacloud.com/avaya-infinity/docs/real-time-contextual-media-streaming), last updated 23 July 2026: > "Real-time Contextual Media Streaming is an **open WebSocket-based protocol** that lets you connect your own AI services directly to Avaya Infinity... **No proprietary SDKs. No middleware.** Just a WebSocket connection carrying **bidirectional audio** and structured JSON messages." The media contract, verbatim: > "**Codecs: PCMU (8 kHz), PCMA (8 kHz), G.722 (16 kHz)** / Frame size: **20ms default, configurable per session** / Channels: Customer audio, agent audio, or both — negotiated at session setup" That is the only wideband codec in this entire survey. Every other platform here, and Genesys, documents narrowband and only narrowband. Avaya documenting G.722 at 16 kHz does not mean your calls will be wideband — the carrier leg still decides — but it means the platform is not the thing forcing the band limit, which is a different situation from the other four. The direction model is explicit and per-service: > "**Egress (out)** (Avaya Infinity → Your Server) — live audio from the endpoint, streamed to your server for processing... **Ingress (in)** (Your Server → Avaya Infinity) — audio generated by your server, played back to the endpoint." / "A Virtual Agent bot typically uses both flows... A recording service uses egress only. A TTS service uses ingress only. **Your server controls which flows are active per endpoint.**" Two transport modes are documented — `avaya-wss`, binary frames over a TLS WebSocket, and `avaya-wss-rtp`, SRTP over UDP *"for ultra-low latency environments"* — both using JSON for control and both supporting dynamic codec negotiation. Avaya is the WebSocket client and you are the server, so *"no inbound firewall rules required on your side"*; authentication is a signed JWT that your server validates; TLS 1.2 or better with public CA certificates. One socket multiplexes multiple endpoints with frames tagged by endpoint ID. **The caveat that keeps this honest: only Virtual Agent is live.** The same page carries a "Coming Soon" table listing Agent Assist, Recording, Text-to-Speech, Speech Recognition, Transcription and Translation as future services on the same protocol. So today, the bidirectional path exists and the way to use it is to be the virtual agent. Do not read this section as "Avaya lets you insert third-party ASR or recording today", because the vendor does not say that. Two further scoping notes. **Avaya Experience Platform has no equivalent** — its media handling is signed upload and download of files, not live audio; real-time streaming is on Avaya Infinity, the newer platform. And **we did not investigate Avaya Aura, AES or Aura Media Server at all.** That is an open question in this post, not a negative finding, and we would rather say so than let silence read as absence. #### What this adds up to Two independent conclusions, and they point in different directions. **On sample rate, the wedge holds and is slightly narrower than we have been saying.** Four of the five surfaces document 8 kHz and nothing else, and NICE says outright that narrowband G.711 is *"the same format as all NiCE CXone telephony audio"*. That is the channel the money arrives on, and [essentially every speech enhancement model on the market is trained wideband](/blog/8khz-is-where-voice-ai-breaks) — the DNS Challenge, the benchmark series most published denoisers are tuned against, never had a narrowband track. But we should stop writing "it is 8 kHz µ-law everywhere": Amazon Connect is linear PCM, not companded, and Avaya documents G.722 at 16 kHz. Two exceptions out of six surfaces is not a rounding error. **On integration surface, the field splits cleanly, and it is not the split the marketing suggests.** Every one of these platforms will let you *listen*. Two of them — Twilio's `` and Avaya Infinity's RCMS — document a way to put audio back into the live media path. Genesys, Amazon Connect and NICE CXone do not, and on the surfaces that do return audio, you are being asked to be the **bot**, not the filter: the audio you send goes to the human, not to somebody else's recogniser. That distinction is the one to carry into a build decision. "Can I process the caller's audio before the platform's speech recogniser hears it?" is a different question from "can I send audio into this call?", and only one of them has a yes anywhere in this survey — on the platforms where you also own the recogniser. Which is where our own numbers become the constraint rather than the platform's. On our seven telephony conditions, [no enhancer we tested produced a meaningful improvement](/blog/nulltest-open-benchmark) over the unprocessed audio, several degraded it by five to ten points of word error rate, and [our own model did worse than every one of them](/blog/we-put-our-model-in-our-benchmark-and-it-lost). The insertion points are the easy part. Having something worth inserting is not solved yet, by us or by anyone whose numbers we can check. #### What we could not confirm - **Amazon Connect: bit depth, endianness and channel count.** 8 kHz and "raw PCM" is the entire published format. Sixteen-bit mono is implied by AWS's own demo material and is not documented. - **Amazon Connect: any latency figure.** None published on any page we read. - **Amazon Connect: the SIPREC ingest codec profile.** Deferred by AWS to Amazon Chime SDK configuration guides, which we did not read. - **Twilio: bytes or milliseconds per media message.** Not published. The widely quoted 160 bytes / 20 ms is not in Twilio's documentation. - **Twilio: ConversationRelay audio format and sample rate.** Not published anywhere we could find, including the voice-configuration page, which covers TTS providers and prosody settings and no codecs. - **Twilio: whether a `` fork survives into a subsequent ``.** Undocumented in both directions, and untested by us. - **Twilio: real-time transcription sample rate.** Not published. - **NICE CXone: the Custom Virtual Agent audio sample rate.** The codec is stated as G.711 µ-law; the rate is not stated on the schema page. - **NICE CXone: per-provider formats behind Agent Assist Hub.** The integrations page does not restate them; the only format specification is on the custom-integration pages. - **Avaya: the wire-level protocol specification.** Avaya publishes a protocol specification PDF, version 1.1, which we did not fetch — so session lifecycle and message definitions are unverified here. - **Avaya Aura, AES and Avaya Aura Media Server.** Not investigated. Open question, not a negative finding. A note on method, because it changes how much weight these findings carry: this survey was done by walking vendor-owned documentation indexes and reading the pages directly, not by searching. That is more reliable for confirming what a document says and less reliable for proving a document does not exist. Every "not documented" above means "we did not find it by crawling the vendor's own index", which is a weaker statement than "the vendor does not document it" — and weaker still than the confirmed quotes, which are verbatim. ### We put our own model in our benchmark and it lost Canonical: https://anecho.ai/blog/we-put-our-model-in-our-benchmark-and-it-lost · Anecho Engineering · published 2026-08-16 · 2849 words · tags: nulltest, benchmark, clearline, negative-result, wer Rule 2 of Null Test is that we score ourselves in the same tables under the same rules. So we did. Clearline posts 39.6% pooled WER against raw's 14.8% — last place in a matrix of twelve, behind every competitor and behind a resample-only control with no model in it. Here is the number, the error decomposition, the control that separates the band limit from the model, and why our own corpus said the opposite. When we published [Null Test v0.1](/blog/nulltest-open-benchmark) we wrote down six rules. Rule 1 was that the raw control is always in the matrix and is allowed to win. Rule 2 was that we score ourselves under the same rules, in the same tables, rather than on a separate marketing page. Clearline is now a row. It came last. Pooled across all eighteen conditions, on the same 180 clips, through the same recogniser as every other backend: **39.6% word error rate against raw's 14.8%.** Thirteen backends in the matrix and ours is the worst of them — behind DeepFilterNet3 measured through a block-online adapter, behind ai-coustics' Quail VF at 18.1%, and behind a **resample-only control with no model in it at all** at 18.6%. That is the whole post in three sentences, and none of the rest of it is an excuse. What the rest of it is, is the decomposition: which part of that number is the band limit, which part is the model, where the model is merely bad versus catastrophic, and why the same checkpoint looked like the best thing this company had produced when we scored it on our own corpus a week earlier. #### Validate the instrument before you trust the result A benchmark that only produces bad news about competitors is a marketing asset. A benchmark that produces bad news about the vendor who runs it is only worth something if the vendor cannot quietly blame the harness — so the first thing we did with a result this bad was try to make it the harness's fault. Two checks, both before we looked at Clearline's row. **The published rows still reproduce.** Re-running `raw` and `gtcrn` through the new code path in a fresh process reproduced **all 54 published rows exactly**, max ΔWER 0.0000 pp. Whatever changed to admit a resampled 8 kHz backend into the matrix did not move anything already in it. **The decoder is not adding noise on this host.** Elsewhere in this project faster-whisper on CPU has returned different transcripts for byte-identical input — three runs of one aggregate gave 121.65 / 122.15 / 121.40% — and the standing guidance from that is to treat differences under roughly 6 pp on a small aggregate as decoder noise. On this host it did not happen. **25 of 25 scopes were bit-identical across repeated decodes** (raw on 18 conditions, 180 clips × 3 decodes; Clearline on 5 conditions, 50 clips × 3 decodes), including the arms where the recogniser is already hallucinating past 100% WER. So the resolution limit on this table is set by the size of the test set, not by the decoder: one reference word is **0.53 pp** in a single condition (190 words), about 0.03 pp pooled. Three repeats cannot prove determinism, and a different host, venv or thread count may well behave differently — the conservative 6 pp guidance still governs anything quoted from another machine. But on this table, a 24.9 pp gap is not a decoder artefact. It is the model. #### The number, and the shape of it | Backend | Pooled WER | S | I | D | VAD F1 | RTF | Latency | |---|---|---|---|---|---|---|---| | Raw (no processing) | 14.77% | 344 | 101 | 60 | 0.950 | — | 0 ms | | ai-coustics Quail L | 15.56% | 349 | 126 | 57 | 0.945 | 0.124 | 30 ms | | ai-coustics Quail VF 2.2 L | 18.10% | 383 | 161 | 75 | 0.945 | 0.076 | 30 ms | | **Resample-only 16→8→16 (control)** | **18.63%** | 413 | 173 | 51 | 0.948 | 0.00005 | 0 ms | | GTCRN (MIT) | 20.26% | 443 | 210 | 40 | 0.947 | 0.050 | 16 ms | | DeepFilterNet3 | 35.18% | 626 | 380 | 197 | 0.901 | 0.133 | 100 ms | | **Clearline (8 kHz, resampled)** | **39.65%** | **776** | **410** | **170** | **0.835** | 0.048 | 16 ms | 3420 reference words per backend. The rows omitted here — the passthrough control, the five FastEnhancer variants and the DFN3 warm-up variant — are unchanged from the published run and are in [the full matrix](/benchmark). **It hallucinates rather than deletes.** Insertions go from raw's 101 to **410**. Deletions go from 60 to 170. An insertion is not a garbled word; it is the recogniser emitting a word that nobody said. This is the same failure mode we documented for every other suppressor in this matrix — [insertions went up, not down](/blog/does-noise-suppression-help-stt) — except that ours does four times as much of it as the unprocessed control, and roughly twice as much as GTCRN. **It is not a cost problem.** RTF 0.048 on one pinned core, p99 block time 4.08 ms against its own 16 ms block period — 3.9x real-time headroom — and 16 ms of algorithmic latency. It is fast enough to ship and not good enough to ship, which is the least useful combination available. **The signal metrics disagree with each other, which is the point of this company.** Against the same pooled audio: on PESQ-WB, Clearline scores **1.685** and the resample-only control scores **1.683** — indistinguishable, a model that reads as doing nothing. On STOI it is 0.744 against the control's 0.840. On SI-SDR it is +0.69 dB absolute where raw is +10.92 and the control is +2.89, so the model costs **2.20 dB beyond the band limit**. On WER it is 21 points worse than the control. Four measurements, four different magnitudes of bad, and only one of them is the one a voice agent actually pays for. That is the [objective mismatch](/blog/does-noise-suppression-help-stt) argument, demonstrated on our own model rather than someone else's. **The VAD row is worse than it looks.** F1 falls from 0.950 to 0.835 — by far the largest drop in the matrix — while the false-alarm rate *improves*, 17.2% to 15.1%. Fewer false alarms and much lower F1 means the losses moved into misses: it is gating out speech. The one branch where enhancement reliably earns its place in our own architecture, [the turn detector](/blog/the-split-pipeline), is a branch this checkpoint would also damage. #### The control that separates the band limit from the model Clearline is an 8 kHz model and these eighteen conditions are stored at 16 kHz, so its row is measured through the harness's resample path: 16 kHz → 8 kHz → model → 8 kHz → 16 kHz, `soxr`, the same path the native 8 kHz ai-coustics baselines take. On `clean`, the noise conditions, `babble`, `competing` and `reverb` that is **not its operating point** — the 4 to 8 kHz octave is discarded before the model ever sees it. The obvious objection is therefore that we handicapped our own model, and the obvious answer is a control. `resample8k` is that exact round trip **with no model in it**: same decimation, same interpolation, same library, nothing in between. Whatever the band limit costs, it costs the control too. Whatever is left over is ours. | Regime | Raw | Resample-only control | Clearline | Band costs | Model adds | |---|---|---|---|---|---| | Wideband (11 conditions, 2090 ref words) | 18.99% | 24.88% | 55.45% | **+5.9 pp** | **+30.6 pp** | | Telephony (7 conditions, 1330 ref words) | 8.12% | 8.80% | 14.81% | **+0.7 pp** | **+6.0 pp** | Read the wideband row first, because it is the one the objection was about. Throwing away the top octave of a 16 kHz recording costs **5.9 points**. Running our model on what is left costs a further **30.6**. The handicap is real and it is one sixth of the damage. Five sixths of it is the model. Then read the telephony row, because that is the regime the model was built for. There the band limit is nearly free — 0.7 points, which is a bit over one reference word per condition — and the model still adds **6.0 points**. It is not being punished by an unfair sample rate on the conditions that matter. It is losing at its own operating point, and it loses to everybody there. On the same seven-condition telephony sub-aggregate: raw 8.12%, the resample-only control 8.80%, ai-coustics Quail L 9.02%, Quail VF 9.70%, Clearline **14.81%**. Beaten by the competition, beaten by doing nothing, and beaten by a piece of wire. #### Where it loses, condition by condition The seven telephony conditions, which are the only ones we would ever quote for a narrowband model: | Condition | Raw | Resample-only control | Clearline | Δ vs raw | |---|---|---|---|---| | `telephony` (G.711 µ-law round trip) | 5.79% | 4.21% | **5.79%** | 0.00 | | `telephony_snr10` | 10.53% | 7.89% | 14.74% | +4.21 | | `telephony_alaw` | 7.37% | 8.95% | 13.68% | +6.32 | | `telephony_g722` | 5.26% | 7.37% | 9.47% | +4.21 | | `telephony_opus12k` | 6.84% | 11.05% | 27.37% | +20.53 | | `telephony_loss3` | 11.05% | 11.58% | 17.89% | +6.84 | | `telephony_loss10` | 10.00% | 10.53% | 14.74% | +4.74 | About 190 reference words per condition, so one word is 0.53 pp — read the `telephony` row as a genuine tie and everything below it as a genuine loss. There is exactly one cell in this table that is not a loss. On plain `telephony` — a clean G.711 µ-law round trip with no added noise — Clearline scores **5.79%, identical to raw to the word**. It ties doing nothing. It is still three words behind the resample-only control's 4.21%, which is to say a piece of wire beat it on the one condition where it drew. `telephony_opus12k` is the cell to stare at: 27.37% against a raw baseline of 6.84%, with insertions at 8 where raw has 2. Opus at 12 kbps is a parametric codec doing aggressive things to the residual, and the model responds to that the way FastEnhancer-M responds to A-law companding — it hallucinates. Whatever this checkpoint learned about what speech looks like, low-bitrate Opus artefacts are outside it. The wideband conditions are worse and less interesting, because they are out of regime. Two of them are worth naming anyway, because they show the failure mode at full size: `babble_snr5` at **138.4%** WER with **169 insertions** where raw has 10, and `reverb_snr10` at **116.3%** with **117 insertions** where raw has zero. A WER above 100% means the recogniser emitted more wrong words than there were words to get right. That is a model producing structured, speech-shaped output out of material it does not understand, and a recogniser dutifully decoding it. #### Why it looked good a week earlier Here is the part that is actually useful to anyone else, because the mistake is not specific to us. Scored on **our own evaluation corpus** — 1,800 competing-speaker clips, same recogniser, `faster-whisper base.en` — the identical checkpoint beat raw audio: | Metric | Raw | Clearline | Δ | |---|---|---|---| | Word error rate | 79.6% | 70.9% | **−8.70 pp** | | Insertions | 5,737 | 3,269 | −43% | | VAD false alarms | 62.8% | 14.6% | −48.2 pp | | SI-SDRi | — | +3.19 dB | — | That was, on its own terms, the best result this company had produced: the first engine anywhere in our own testing to beat raw audio on word error rate. It is also, in isolation, an advertisement — and the reason we are showing it here rather than on the landing page is that the two tables are not comparable and were never meant to be. Different clips, different conditions, different pooling, different difficulty distribution. Only the recogniser is shared. Null Test exists precisely so that a comparable number exists. The comparable number is 39.6% and last place. **That contrast is the lesson, and it generalises.** A model evaluated on the distribution it was trained on will report the performance of that distribution. Our corpus is competing-speaker mixtures rendered under a specific set of assumptions — a specific level, a specific onset order, a specific interference-to-target ratio range — and on material drawn from those assumptions the model does what it was taught. Null Test is eighteen conditions it never saw: reverberation, low-bitrate codecs, packet-loss concealment, broadband noise at five SNRs, clean speech with nothing to remove. It is overfit to its training distribution, and the only way to find that out was to score it against a corpus we did not build for it. Every vendor claim in this category, including the two we quote most often, is a number from the vendor's own corpus. We now have a measured example of how far apart those two numbers can be for one checkpoint: **8.7 points better than raw on ours, 24.9 points worse on a neutral one.** We do not know that anyone else's gap is that wide. We do know nobody has published theirs. #### The two gates this checkpoint does not pass, neither of which is in the table Both are measured on Clearline's own eval, `runs/t4_300h_200ep_solo`, and neither is visible anywhere in Null Test: - **A solo regression.** On clips with no competing speaker at all, the model makes the audio worse: **−0.874 dB SI-SDR with 37.1% of solo clips degraded** (−0.595 dB / 33.0% on the val split). Most phone calls have exactly one person on the line. A lone caller attenuated as if they were interference is worse than doing nothing. - **The speaker conditioning contributes approximately nothing.** The matched ablation scores FiLM conditioning at **+2.815 dB** against no conditioning at all at **+2.690 dB** — a 0.125 dB difference against a 0.3 dB noise floor. What is being measured in this table is therefore a plain denoiser, not the primary-speaker isolation the architecture is for, which is the entire product claim. There is also no ONNX export, so it cannot run in a browser or through the Node SDK. **Clearline is not shippable, this row is not a claim that it is, and nothing on this site sells it.** The row exists so the model can be *placed* next to the competition on identical audio. #### Caveats, in full These attach to every number above. - **Single recogniser.** faster-whisper `base.en`, CTranslate2 int8 CPU, greedy. A different recogniser gives different absolute numbers. The harness supports Deepgram and AssemblyAI so the ordering can be re-checked. - **About 190 reference words per condition**, 3420 pooled. One reference word is 0.53 pp in a condition, about 0.03 pp pooled. Differences of a point or two between adjacent rows are not results; differences of tens of points are. - **Clearline's row is measured through a resample path**, 16 kHz → 8 kHz → model → 8 kHz → 16 kHz, because these conditions are stored at 16 kHz after their carrier legs. The `resample8k` control is that round trip with no model in it, which is how the band limit and the model are told apart above. - **The ESC-50 noise conditions are CC-BY-NC-3.0** and therefore evaluation-only. LibriSpeech is CC-BY-4.0, and the babble and competing-speaker conditions are built from LibriSpeech alone. - **DeepFilterNet3 is measured through a block-online adapter** because no public ONNX export carries recurrent state tensors. Its rows are a lower bound on quality and an upper bound on cost. Never quote its number without this sentence. - **The ai-coustics rows are a proprietary baseline** run through the licensed SDK — a reference point, never a shipping code path for us. - **onnxruntime pinned to one intra-op thread**, sequential execution. The RTF column is one stream on one core. - **Clearline's own eval and this leaderboard are not comparable.** Different clips, different conditions, different pooling. That is the whole reason both appear above. - **Still missing from the condition set: acoustic echo with non-linear speaker distortion.** It is the impairment behind the half of this product that cannot be closed with a patch, and until it is a condition, none of these tables measure the thing we most want measured. #### What happens to the row It stays. Rule 3 is that every cell publishes, including the ones we lose, and no condition gets dropped after we see the result. The next checkpoint gets scored under the same rules against the same control, and if it also loses, that publishes too. The diagnosis we are working from is in [the warm-up contract post](/blog/the-warm-up-contract): the corpus this checkpoint learned from always had the target speaker talk alone first, which is not what a phone call does, and the measured cost of that assumption is large enough to account for a lot of what is above. Fixing a corpus is slower than fixing a loss function and it is the actual work. The full matrix, including every cell where we lose, is at [/benchmark](/benchmark). The raw results file is at [/data/nulltest/v1/results.json](/data/nulltest/v1/results.json) and the harness is at [github.com/anecho](https://github.com/anecho). If you think we configured our own model badly, the fix is a pull request against a YAML file, and we will rerun and republish. ### We trained a speaker-isolation model on an assumption that does not hold in real calls Canonical: https://anecho.ai/blog/the-warm-up-contract · Sofia Marchetti · published 2026-08-16 · 2330 words · tags: clearline, speaker-isolation, training-data, negative-result, telephony Every clip in our corpus had the target speaker talk alone first. Real calls do not — the television is already on when you dial. With a background running at t=0 the model's speaker embedding locks onto the background in 96% of clips: −14.63 dB SI-SDR, 86% of clips made worse, and 131.9% WER against 121.9% for not running the model at all. Splicing 0.8 s of target-alone speech onto the front recovers it to +4.91 dB. A 19.5 dB swing from timing alone, on the same audio. Target-speaker extraction has to answer a question that generic denoising never asks: of the two people talking, which one am I supposed to keep? There are a few ways to tell a model the answer. You can enrol the speaker in advance and hand the network an embedding — accurate, and useless for an inbound call from a stranger. You can pick the loudest, or the nearest, which is not speaker isolation, it is a level meter. Or you can use the structure of the call itself: **the primary speaker is the one talking alone at the start.** Somebody says "hello" before the noise begins. Pool a speaker embedding over that opening window, condition every subsequent frame on it, and you have an enrolment-free system with a definition of "primary" that requires no configuration at all. That is the warm-up contract, and it is what we built. Our model commits its embedding at **736 ms** — 46 frames at a 128-sample hop — and every clip in every corpus we generated before this month rendered the target from t=0 with the interferers entering afterwards, because that is what makes the contract well defined. Phone calls do not do that. The television was already on. The open-plan office was already loud. The other person on the speakerphone was already mid-sentence when the line connected. In a large fraction of real calls, **the interference is running at t=0 and the caller starts talking into it.** We measured what that costs. It is the largest single number this project has produced, and it is not in our favour. #### The experiment Four arms, rendered from **one room and one pair of sources**. Same room, same simulated impulse responses, same utterances, same target-to-interference ratio, same codec leg. The arms differ in **timing and nothing else**, which is what makes this an ablation rather than four unrelated experiments. ```text A trained contract target ──────────────────────▶ bed ┌────────────────▶ (target alone for 0.8 s) B reality target ┌────────────────▶ bed ──────────────────────▶ (bed running at t=0) C bed only target bed ──────────────────────▶ (no target — control) D B + lead-in target ────┐ ┌───────────────▶ bed └─────────────────▶ (0.8 s of B's own target, spliced on) ``` Arm D is the one that decides what the result means. It is arm B's audio with a genuine target-alone lead-in spliced onto the front — same speaker, same room, same recording. If D recovers arm A's performance, the defect is the warm-up contract. If D stays broken, the problem is simply that this material is out of distribution and timing has nothing to do with it. 48 items × three TIRs (0, 6 and 12 dB) = **144 matched pairs per arm**, scored offline against the `t4_300h_200ep_solo` checkpoint at epoch 199 — the same checkpoint that appears, and [comes last](/blog/we-put-our-model-in-our-benchmark-and-it-lost), in our public benchmark. #### The result | Arm | In SI-SDR | Out SI-SDR | Δ | Clips degraded | cos → target | cos → background | Embedding picked the background | |---|---|---|---|---|---|---|---| | **A** trained contract | +6.75 dB | +11.76 dB | **+5.01 dB** | 6.9% | 0.948 | 0.862 | 9% | | **B** reality (bed at t=0) | +4.71 dB | −9.92 dB | **−14.63 dB** | **86.1%** | 0.866 | 0.971 | **96.5%** | | **D** B with 0.8 s lead-in | +5.67 dB | +10.58 dB | **+4.91 dB** | 6.9% | 0.948 | 0.862 | 9% | | B with an oracle embedding | +4.71 dB | +6.52 dB | +1.81 dB | 20.1% | — | — | — | Arm B is not "the model helps less". It is **−14.63 dB**: the output is fifteen decibels further from the target than the input was, and **86% of clips come out worse than they went in**. A model that reliably damages six clips in seven is not underperforming, it is doing the wrong job confidently. And arm D is the answer to the objection. Splice 0.8 seconds of the target's own voice onto the front of the identical mixture and the model recovers to **+4.91 dB, with 6.9% degraded — arm A's numbers to two decimal places on the cosine columns.** The audio in D and B is the same audio. The room is the same room. **A 19.5 dB swing, from timing alone.** The material is not out of distribution. The contract is. The downstream number is the one a voice agent actually pays: | Arm | WER | |---|---| | raw 8 kHz, **no model in the path** | 121.9% | | A trained contract | 64.5% | | **B reality** | **131.9%** | | D B with lead-in | 76.5% | 39 items, TIR +6 dB, `faster-whisper base.en`. Word error rates above 100% mean the recogniser emitted more wrong words than there were words to get right, which is what happens when the surviving signal is a second person talking fluently. On its trained contract the model takes 121.9% down to 64.5% — a large, real win, and the result that made us optimistic. On the case a phone call actually presents it lands at **131.9%, worse than not running the model at all.** #### The mechanism is visible in one number The two cosine columns are the whole diagnosis. In arm A the pooled embedding sits at cosine **0.948** to the target speaker and 0.862 to the background. In arm B those swap: **0.866** to the target, **0.971** to the background. The embedding is not confused, degraded or noisy. It is a clean, confident embedding **of the wrong voice**, and every downstream frame is conditioned on it. The model then does exactly what it was trained to do — keep the speaker matching the embedding and suppress everything else — which on arm B means keeping the television and suppressing the caller. It does that in **96.5% of clips**. This is not a tail failure to be fixed with more data of the same kind. It is deterministic behaviour following from a warm-up window that pools whatever is making noise when the window opens. The oracle row bounds how much of the damage the embedding is responsible for. Hand the model the *correct* embedding and run it on arm B's audio anyway, and it recovers to **+1.81 dB with 20.1% degraded** — from a 14.6 dB hole to a modest positive. So the embedding is most of the defect and not all of it: even correctly conditioned, arm B material only reaches +1.81 dB against arm A's +5.01, because a mixture with no target-alone region anywhere is genuinely harder. Fixing the conditioning is necessary. It is not sufficient. #### What we could and could not fix at inference The corpus takes days to regenerate and a training run takes longer, so we asked what a front end could do in the meantime. The honest answer is: half of it. There are two distinct versions of the warm-up defect, and they are not equally reachable. **The quiet start** — nobody is talking yet, and the window pools room tone. A front-end gate can fix this, because there is a real acoustic onset to find. Holding the embedding until a sustained rise over a primed background estimate takes that arm from **+1.83 to +4.03 dB**, cuts degraded clips from 21% to 10%, and drops the rate at which the embedding lands on the wrong source from **69% to 4%**. Cosine to target goes 0.882 → 0.966. That is a genuine fix and it shipped. **The bed running at t=0** — the arm B case — a front-end gate cannot reach, and the reason is structural rather than a tuning failure. The bed *is* speech. Any input-side speech detector fires at t=0, because something really is talking. The gate opens at 112 ms and the embedding still lands on the interferer. Measured: **−14.11 → −13.98 dB, +0.13 dB against a 0.3 dB noise floor**, with the wrong-source rate going *up*, 94% to 97%. We also quantified the aggressive setting — never take the "something is already talking, open immediately" escape — because it looks like the obvious answer. It buys arm B **+1.98 dB against a 14 dB hole** and costs **−1.03 dB on arm D**, whose degraded rate goes 10% to 18%. With a background running continuously there is no clean onset to find, so the gate simply fires about 370 ms later on the background's own fluctuations. **No front-end gate setting meaningfully fixes this.** It is a training-data problem and it has to be paid for as one. One more finding from that work, because it is the sort of thing that makes a metric lie. The same front end also normalises level, which is a large real win on the trained-contract arm — **+9.65 dB** at the quiet input level a real user reported, and a no-op at the level the corpus was normalised to. On arm B, level normalisation makes the number *worse*: −2.34 → −14.10 dB. It does not create the defect. It **un-masks** it. Arm B at the trained level was already −14.11 with the front end off, so being too quiet had been accidentally protecting users from a model acting confidently on an embedding of the wrong person. Restoring the correct operating point restored the damage. We shipped that as a stated regression, in the UI and in `/health`, rather than as a quiet improvement to an average. #### The corpus fix, and the check that it did not invalidate the past The corpus now renders all three onset orders instead of one: | Onset order | Share of eligible clips | |---|---| | Target first (the old contract) | 45% | | Interferer first | 40% | | Simultaneous | 15% | Verified on 6,000 drawn specifications before any rendering: 45.5 / 40.1 / 14.4% over eligible clips, 30.6% interferer-first corpus-wide once solo clips and the same-distance control are counted — those are forced to target-first because their label is otherwise undefined. One parameter matters more than the ratio. The interferer's head start is drawn from **200 to 2500 ms**, deliberately longer than the 300–800 ms warm-up window. If the lead were always shorter than the window, "the background is already playing" would still resolve inside the pooling window every time, and the model could keep using the same shortcut while appearing to have learned the harder case. The point of the range is that it *cannot*. Target-first keeps a plurality, because a lead-in is still the cleanest available definition of who to keep. It is no longer the only thing the model has ever seen. And because changing a data generator silently invalidates every run that came before it: re-rendering 60 existing validation specifications through the revised generator reproduces the shards on disk **bit for bit** — worst absolute int16 delta 0, zero mismatched clips. All 60 old specifications deserialise under the new field defaults as target-first at the old level, so the revision cannot have changed any run trained before it, and old-versus-new comparisons stay valid. A data pipeline change that cannot be shown to be inert on old inputs is a change that quietly rewrites your history. #### Caveats - **144 matched pairs per arm, 48 items × three TIRs**, one checkpoint, offline forward pass. Enough to establish a 19.5 dB effect; not a shipping validation. - **The WER figures come from a smaller sample and a noisier instrument.** 39 items at TIR +6 dB. Running the identical items three times gave 121.65 / 122.15 / 121.40% on the unprocessed control, so the standing guidance from that measurement is to treat differences below roughly 6 pp on an aggregate this size as decoder noise. The 131.9% versus 121.9% gap is 10 pp and clears that bar; the 64.5% versus 76.5% gap between arms A and D is larger still. Do not read one-point differences in this table. - **These are simulated rooms**, convolution with simulated impulse responses, one interferer, a codec leg. Not recorded calls. The direction of the effect is what we would defend; the exact decibel figure is corpus-specific. - **This is one architecture's warm-up window.** A system that enrols speakers in advance does not have this failure mode, and has a different one. - **The checkpoint measured here is the same one that lost our public benchmark** and it remains [not shippable](/blog/we-put-our-model-in-our-benchmark-and-it-lost) for two further reasons that this post does not fix: it degrades single-speaker clips, and its speaker conditioning contributes about 0.125 dB against a 0.3 dB noise floor, which means it is currently behaving as a plain denoiser rather than as speaker isolation. #### The general version The failure is not that the model is weak. On the distribution it was trained on it does the job well — +5.01 dB, 121.9% WER down to 64.5%. The failure is that the distribution encoded an assumption nobody wrote down as an assumption. "The target speaks first" entered the corpus as a *convenience*, because it makes the label well defined, and it left the corpus as a *requirement*, silently, with no test asserting it and no metric reporting when it was violated. That is the shape to look for in your own data. Somewhere in every generated corpus there is a decision that was made to keep the labels clean, and if the world does not honour it, your evaluation will never tell you — because your evaluation was generated by the same code, under the same convenience. The cheap detection is the one used above: build a matched pair that differs in the suspect variable and nothing else, and see how far apart they land. Ours were 19.5 dB apart. We would rather have known that before the training run than after it, and the only reason we know it now is that somebody asked what happens when the television is already on. The public benchmark row for this checkpoint, including every condition where it loses, is at [/benchmark](/benchmark). ### The resampler that ate 92% of a microphone Canonical: https://anecho.ai/blog/the-resampler-that-ate-92-percent-of-a-microphone · Daniel Reiss · published 2026-08-16 · 1920 words · tags: dsp, resampling, voice-agents, debugging, streaming soxr.ResampleStream is a burst emitter. Fed the browser's 128-sample AudioWorklet quantum it returns an empty array on 92% of calls and then hands back 95 ms at once. We read that emptiness as 'still priming' and skipped the block — which also skipped the untouched microphone buffer sitting next to it, so 8% of the caller ever reached the model. Our own telemetry reported it as 0.08 for days. This is a control-flow bug, not the resample-tax claim we retracted, and here is the four-line probe that finds it. An agent we were building would connect, show a live session, stream audio for the whole call, and never once answer. The socket stayed open. Audio kept leaving the browser. Nothing came back. The cause was four tokens of Python: ```python x8 = down(xin) if not x8.size: continue ``` `down` is a stateful `soxr.ResampleStream`. The guard reads as "the polyphase filter has not primed yet, skip this block and come back next time." That is a reasonable thing to believe about a resampler, and it is wrong about this one. `soxr.ResampleStream` does not prime once and then emit steadily. **It buffers internally and emits in bursts, for the entire session**, and if you feed it small blocks it returns an empty array on the overwhelming majority of calls. Before the diagnosis, one thing this post is not. We have previously [retracted a claim](/blog/8khz-is-where-voice-ai-breaks) that resampling per se costs you a fixed fraction of caller audio; we built the control and our measurement did not support it, and we are not quietly reintroducing it here. Nothing below is a property of resampling. It is a property of one library's emission schedule meeting one `continue` statement, and the audio was thrown away by our control flow, not by any filter. #### The measurement Four lines, reproducible on any machine with `soxr` installed. Feed a stream resampler the browser's own render quantum and count how often it gives you anything back. ```python import numpy as np, soxr rs = soxr.ResampleStream(16000, 8000, 1, dtype="float32", quality="VHQ") sizes = [rs.resample_chunk(np.zeros(128, np.float32)).size for _ in range(400)] print(sum(s > 0 for s in sizes), "non-empty of", len(sizes)) ``` On this host — `soxr` 1.1.0, macOS on Apple silicon — that prints **33 non-empty of 400**. So 91.8% of calls return nothing at all, and the 33 that do return a *fixed* 758 samples each: 94.8 ms of 8 kHz audio, arriving in one lump, over and over, for as long as the session lasts. The block size is a property of the resampler, not of your input. Vary the chunk you feed it and only the *frequency* of the bursts changes: | Input chunk | Chunk duration at 16 kHz | Calls returning nothing | Burst size out | |---|---|---|---| | 128 samples (AudioWorklet quantum) | 8 ms | **91.8%** | 758 samples (94.8 ms) | | 256 samples | 16 ms | 83.2% | 758 samples | | 512 samples | 32 ms | 66.5% | 758 samples | | 1024 samples | 64 ms | 32.5% | 758 samples | | 2048 samples | 128 ms | 0% | 758 or 1516 samples | 400 calls per row, one fresh `ResampleStream` each, 16 kHz to 8 kHz at `VHQ`. Two things fall out. First, there is a threshold, and it is exactly the input needed to fill one output burst: at 2:1 decimation, 758 output samples take 1516 input samples, and feeding 1516 per call takes the empty returns down to the single priming call (0.5% of 200). At 2048 even that one comes back. Which is why nobody hits this with 20 ms server-side frames from a SIP leg and everybody hits it with a browser worklet. Second, the burst size tracks the filter, not the plumbing: at `HQ` it is 830 samples, at `LQ` 470, at 48 kHz → 16 kHz it is 1100. Those numbers are this version on this host and you should measure your own rather than trust ours; the *shape* is what transfers. None of this is a defect in `soxr`. A polyphase resampler with a long anti-alias filter has to accumulate input before it can produce correctly-filtered output, and emitting in whole internal blocks is the efficient way to do that. The library is behaving exactly as a stream resampler should. The bug is entirely in what we concluded from an empty return. #### Why an empty array meant something different in two places We had the same shape of guard in two services, and it was harmless in one and severe in the other. The difference is worth stating precisely, because "we had this bug twice" is not the interesting part. On the enhancement endpoint the code is: ```python x8 = down(pcm) if not x8.size: continue # harmless here y8 = model.process(x8) await ws.send_bytes(up(y8)) ``` Everything downstream is *derived from* `x8`. An empty `x8` genuinely means there is nothing to process yet, nothing has been dropped, and the audio is sitting inside the resampler's buffer waiting for its burst. The `continue` is correct. On the agent endpoint the same three lines sat at the top of a function that had **two** parallel jobs: run the caller's audio through the model at 8 kHz, and separately keep an untouched 16 kHz copy for the control arm of an A/B. The `continue` returned before the second job ran. So with the filter switched **off** — the arm with no model in it at all, the arm that is supposed to be the honest baseline — the raw microphone buffer was discarded along with the empty `x8` that had nothing to do with it. The result, measured end to end against the real endpoint: **3.7 seconds of uplink across a 46 second call.** About 8% of the caller. Zero turns, ever. The rule we wrote into the code afterwards is narrower than "never skip a block", because sometimes skipping is right: > Nothing may condition one arm on the other arm's resampler having produced output. #### It was in the telemetry the whole time This is the part that should be uncomfortable, and it is the reason the post exists rather than a commit message. The service already emitted a `stats` message roughly every sixteen blocks, and that message already carried a field called `uplink_realtime` — uplink seconds divided by wall-clock seconds since the session went live. On a healthy call it sits near 1.0. On the broken calls it sat at **0.08**, in a JSON blob, in a browser console, for days. Nobody read it, because the user-visible symptom ("the agent does not answer") pointed at the model, the prompt, the credentials and the API before it pointed at arithmetic. A number that says *8% of the audio you think you are sending is being sent* was one scroll away from the person debugging, and it lost to a more interesting hypothesis every time. Two changes came out of that, and they are cheap enough that we would recommend both to anyone running a live audio path: - **Publish a ratio, not a count.** Bytes sent is unfalsifiable — it goes up on a broken call too. Seconds of audio delivered per second of wall clock has a known correct value, so a wrong value is legible without context. It now ships in the `session` and `stats` messages with a stated threshold: below roughly 0.9 means audio is being lost or arriving late. - **Say what the number implies, in the payload.** The `stats` message now carries the sentence "Gemini's VAD cannot end a turn on audio it never received" next to the number, because the number alone did not connect to the symptom for anyone who had not already found the bug. #### The second-order bug: the arms were not comparable Fixing the discard exposed a subtler problem in the same code, and it is the one that would have quietly invalidated the experiment the endpoint exists for. That endpoint is a single switch: same speaker, same background noise, filter on versus filter off, mid-conversation. For that comparison to mean anything the two arms must differ in the audio and in *nothing else*. They did not. The filtered arm inherited the resampler's ~95 ms burst cadence; the unfiltered arm passed the browser's raw 8 ms quantum straight through. Measured on the same call: **10.6 frames per second against 125.0** — an 11.8x difference in how one conversation was packetised. Any behavioural difference a listener attributed to the filter was confounded with frame cadence, and 125 frames per second is far more than the API expects. Both arms are now coalesced to a fixed 100 ms frame before they are sent, so cadence is identical by construction, and the per-arm rate is *reported* rather than assumed — the health payload carries `uplink_frame_hz` for both arms, and they are supposed to read the same. There is an interlock worth naming, because it is the kind of thing that turns one fix into a different bug. A part-built frame is audio that has not been sent. Coalescing is only safe because something else guarantees the buffer always has a producer — otherwise the tail of the caller's last sentence sits in a half-full frame at exactly the moment the server needs it in order to detect that the caller has stopped. That guarantee is the keep-alive described in [the Gemini Live post](/blog/gemini-live-audio-stream-end), and the two changes are not separable. #### Why use a stateful resampler at all The obvious escape is to call the stateless `soxr.resample()` once per chunk and never see an empty return. Do not. A stateless call restarts the polyphase filter at every chunk boundary, which stamps a discontinuity into the signal about thirty times a second on a 32 ms chunk. It is audible as a buzz at the chunk rate, and it is the kind of artefact a listener will quite reasonably blame on your model. The stateful resampler carries its filter state across calls precisely so that does not happen, and the bursty emission is the visible cost of that correctness. That cost is honest and bounded: it is a fixed start-of-session delay, not jitter. We measure it rather than assume it — `roundtrip_delay_ms()` runs a correlation between input and output through a down/up resampler pair and reports the lag, and the result is published in `/health` as transport delay, kept separate from the model's own 16 ms algorithmic latency. Two different delays with two different owners, reported as two numbers. #### The checklist If you run a browser-to-server audio path with a sample-rate conversion in it: 1. **Count what you actually send.** Seconds of audio delivered over wall-clock seconds, per session, exposed to whoever is debugging. Not bytes. 2. **Never let a `continue` cover more than the thing it is guarding.** If a block does two independent jobs, an early return from one is a silent failure of the other. Split them. 3. **Probe your resampler's emission schedule at your real block size**, with the four lines above. If you feed it 128-sample quanta, expect nothing back most of the time; if you feed it 20 ms server frames, you will never see this at all. 4. **Feeding it a full output block's worth of input per call makes the empty return almost disappear** — but do not build on that number, because it is version- and ratio-dependent, and the priming call is still empty. Handle the empty case correctly instead. 5. **Do not compare two arms with different packetisation.** Coalesce to one frame size before the transport, and report the per-arm frame rate so a divergence is visible rather than inferred. And the general one, which is the only part of this that is not about audio: the instrumentation that would have caught this already existed and already had the right value in it. Building the metric is the easy half. Making the metric say what it implies, next to the symptom the person is actually looking at, is the half we got wrong. ### Genesys AudioHook is 8 kHz µ-law, and there is nowhere to stand Canonical: https://anecho.ai/blog/genesys-audiohook-8khz · Sofia Marchetti · published 2026-08-16 · 2668 words · tags: genesys, telephony, contact-centre, narrowband, integration The Genesys Cloud AudioHook protocol pins its sample rate in the type system: MediaRate is the literal 8000, and 16000 appears nowhere in the specification. AudioHook Monitor is a one-way tap whose server output is explicitly discarded; Audio Connector can play audio back, but it is a bot fork inside the IVR that pauses the flow and never reaches an agent. A spec-by-spec reading of what a third party can and cannot do in that media path, with the four things we could not confirm named as unconfirmed. *Every specification claim below is quoted from Genesys' own published documentation and linked. Where we could not confirm something, it is listed as unconfirmed rather than filled in — there is a section for that near the end. If any of this is wrong or changes, tell us and we will correct it with a date.* Contact centres are where voice AI meets actual revenue, and Genesys Cloud is one of the largest of them. So the question that matters for anyone selling audio processing is narrow and answerable: **what does the media path look like, and is there a place in it where a third party can stand?** The Genesys answer is unusually legible, because they publish the protocol rather than describing it. Here is the whole sample-rate question, as it appears in their [type definitions](https://developer.genesys.cloud/devapps/audiohook/type-definitions): ```typescript type MediaFormat = 'PCMU'; type MediaRate = 8000; ``` Not a default. Not a recommendation. A literal type. The [protocol reference](https://developer.genesys.cloud/devapps/audiohook/protocol-reference) restates it in prose — *"Sample rate of the media format in Hertz. The Genesys Cloud client currently only supports 8000Hz"* — and the string `16000` does not appear anywhere in the AudioHook specification. That is the wedge this whole site is about, stated by the platform itself. Every third party integrating audio into Genesys Cloud through the documented protocol receives narrowband G.711 µ-law, and [essentially every speech enhancement model on the market is trained wideband](/blog/8khz-is-where-voice-ai-breaks). #### The transport, briefly AudioHook inverts the roles you might expect. **Genesys is the WebSocket client; your service is the server.** > "The AudioHook protocol uses WebSockets over TLS as the transport and is designed to make it easy to implement servers that accept it." — [introduction](https://developer.genesys.cloud/devapps/audiohook/introduction) Concretely, from the [security page](https://developer.genesys.cloud/devapps/audiohook/security) and the [session walkthrough](https://developer.genesys.cloud/devapps/audiohook/session-walkthrough): | Property | Value | |---|---| | Port | 443, and only 443 | | TLS | 1.2 and 1.3; certificates must be signed by a public CA (no self-signed) | | Framing | JSON in text frames, raw audio in binary frames | | Maximum message size | 64,000 bytes, text and binary alike | | Open / ping / close timeouts | 5,000 ms / 5,000 ms / 10,000 ms | | Application ping interval | every 5 seconds per connection | | Auth | `X-API-KEY` header, plus optional HMAC-SHA256 request signing with a mandatory nonce | | Retries | up to 5, exponential backoff | Two details worth pulling out because they bite implementers. **The ping is not the WebSocket ping.** These are application-level `ping`/`pong` JSON messages — *"a protocol feature distinct from the WebSocket ping/pong messages (which are not used)"* — and an unsolicited `pong` is a protocol error. You have five seconds to answer, every five seconds, or Genesys may treat the connection as lost and re-establish it. **There is no WebSocket subprotocol.** We looked for one specifically. The documented handshake carries `Audiohook-Organization-Id`, `Audiohook-Correlation-Id`, `Audiohook-Session-Id` and `X-API-KEY` headers, and no `Sec-WebSocket-Protocol` at all; the string does not appear on any AudioHook specification page. If you were planning to route by subprotocol, route by path instead. #### The media negotiation, and what the two channels actually are The `open` message carries an offer and your `opened` response picks from it, SDP-style. The rule is strict: > "The server **must** choose exactly one of the entries offered and **must not** modify an offered media format (including the channels and their order). This is similar to the offer-answer exchange of an SIP/SDP media negotiation, just simpler." — [session walkthrough](https://developer.genesys.cloud/devapps/audiohook/session-walkthrough) A typical offer is three entries, all PCMU at 8000 Hz, differing only in channels: `["external", "internal"]`, `["external"]`, `["internal"]`. The channel names are the part people get wrong, and they are not "caller" and "agent". AudioHook follows a **participant**, not a leg: > "Currently, two values are supported: `external`, which represents what the party represented by they participant speaks and `internal`, which represents what they hear." — [protocol reference](https://developer.genesys.cloud/devapps/audiohook/protocol-reference) So `internal` is *everything that participant hears* — which can be IVR prompts, ACD hold music, or, in a conference, a mix. Genesys says so explicitly: *"the audio of the 'internal' channel in our AudioHook represents a mix of the audio from both agents."* If you are building anything that assumes one voice per channel, that assumption fails on a warm transfer. In a stereo answer, left is index 0 and right is index 1. #### Frame sizes: there aren't any This is the single most likely source of a bug in a first implementation, and the specification is blunt about it. > "The number of samples per frame is variable and is up to the client... **The server must not make any assumptions about audio frame sizes** and maintain a timeline of the audio stream by counting the samples." — [session walkthrough](https://developer.genesys.cloud/devapps/audiohook/session-walkthrough) The only concrete number given is illustrative: a 100 ms frame of two-channel PCMU at 8000 Hz is 1600 bytes, headerless and interleaved. Your framing budget is the 64,000-byte message ceiling and the rate limits — [10 binary messages per second on average with a burst of 25](https://developer.genesys.cloud/organization/organization/limits), the same for text. And the stream is not paced like a phone call. Genesys buffers: > "The client maintains a history buffer of at least 20 Seconds of audio... it will send the buffered audio to the server faster than real-time and then continue with the real-time stream. The rate at which the client 'catches up' is undefined." So a session opens with a burst of history at an unspecified rate, then settles into real time. Anything you build with a fixed block size — which is every streaming enhancement model, ours included — needs its own re-framing buffer in front of it, and anything that measures real-time factor from wall-clock arrival will read nonsense for the first few seconds. When the client cannot keep up it tells you, with a `discarded` message carrying `start` and `discarded` durations and the invariant `position = start + discarded`. That is a gap announcement, not a control command: *"whenever there is a discontinuity in the audio stream due to unexpected loss of audio."* Explicit pause and resume do not produce one. If you are counting samples to maintain a timeline — and the spec tells you to — `discarded` is how you learn your timeline just moved. #### The question that decides everything: can you send audio back? It depends entirely on which of the four AudioHook features you are implementing, and the answers are not similar to each other. | Feature | Audio to your server | Audio back into the call | What you may return | |---|---|---|---| | **AudioHook Monitor** | Yes, `external` and/or `internal` | **No** | Nothing in band | | **Audio Connector** | Yes, `external` only | **Yes** — played to the caller | Audio, as bot prompts | | **Transcription Connector** | Yes | No | `transcript` events | | **Bot Transcription Connector** | Yes (PCMU or L16) | No | `transcript` and speech events | **AudioHook Monitor is a tap and says so.** The [features page](https://developer.genesys.cloud/devapps/audiohook/features) does not leave room for interpretation: > "here the client will stream the conversation audio to the server but **no in band data can be sent back to the client. Data returned by the server will be silently ignored and discarded.**" Note *silently*. Not an error, not a warning — your processed audio disappears and the session looks healthy. The [introduction](https://developer.genesys.cloud/devapps/audiohook/introduction) frames the whole protocol the same way: *"Think of this as 'taps' on the audio streams to and from participants' parties."* **Audio Connector genuinely is bidirectional**, and this is the one that looks, at first, like the opening: > "This feature supports bi-directional audio streaming... **Any audio data sent to the client will be played to the caller.** Only the `external` audio channel is sent from the client to the server." — [features](https://developer.genesys.cloud/devapps/audiohook/features) It has the machinery you would expect for that: `playback-started` and `playback-completed` client messages, a `barge_in` server event, a `dtmf` message. Server-to-client audio is capped at *"no more than 64,000 bytes per message"* and paced — *"the server should not send audio more often than every 200ms"*. There is one documented contradiction worth knowing about before you spend a day on it. The error table in the same protocol reference still says the opposite: > "`415` Unsupported Media Type — The server sent a binary message to the client... **sending audio from the server to the client is not supported**." Our reading is that 415 predates Audio Connector and applies to Monitor and Transcription Connector sessions, where binary from the server is indeed illegal; the [changelog](https://developer.genesys.cloud/devapps/audiohook/what-changed) shows bidirectional audio, DTMF, playback and barge-in arriving in a later update. But the specification contains both statements today, and if you are building against it you should expect to discover which one applies to your session type empirically. #### Why Audio Connector is still not an insertion point Here is where the architecture, rather than the protocol, closes the door. Audio Connector is a **bot turn**, not an inline filter, and three properties of it are individually disqualifying for anyone trying to clean audio before someone else's bot: > "The Call Audio Connector action in Architect **forks the voice stream**, sends it to the configured URL, and then **pauses the flow execution** at this point, until the bi-directional stream ends... **The bi-directional streaming session is active only in the IVR channel. It does not transfer to an agent.**" — [Audio Connector overview](https://help.genesys.cloud/articles/audio-connector-overview/) 1. **The audio you send goes to the caller, not to the bot.** You are being asked to be the voice on the other end. There is no "return the cleaned caller audio and let the platform's speech recogniser hear it instead" direction, because the direction that exists points at the human. 2. **It is a fork with a blocking flow.** Execution stops at the action and resumes when your session ends. That is a dialog turn, not a pipeline stage. 3. **It never reaches an agent.** IVR only. Plus the smaller constraints, all documented on the same page: one bidirectional stream, `external` channel only, [300 concurrent Audio Connector calls and a 900-second per-call ceiling](https://developer.genesys.cloud/organization/organization/limits), and no support under BYOC Premises. So the honest conclusion, checked surface by surface: **Genesys Cloud documents audio egress, audio dialog in the IVR leg, and metadata return. It documents no way for a third party to receive audio, process it, and have the processed audio be what the platform's own downstream consumers hear.** We checked Monitor, Audio Connector, Transcription Connector, Bot Transcription Connector, Genesys Agent Assist (a knowledge-surfacing feature with no audio path at all), Genesys Bot Connector (text and intents, message flows only), and Bring Your Own Interactions (post-hoc ingestion, not live media). #### Where you can stand, and it is exactly one place There is a real insertion point, and it is not the one people look for. **If you are the speech recogniser, you can clean the audio before your own recogniser.** The Bot Transcription Connector exists precisely to let a third party be the STT engine — *"allows audio to be streamed to the AudioHook-based server that will serve as an STT engine"* — and Genesys' own help centre describes it as the way to *"integrate third-party ASR engines using the Genesys AudioHook protocol"*. That path receives PCMU, and uniquely also offers L16, and returns `transcript` and speech-start events rather than audio. Which relocates the question rather than answering it. Inside that boundary you own the whole chain: 8 kHz µ-law in, your processing, your recogniser, transcript out. Nothing in the platform stops you from putting a filter there. What stops you is that on our own measurements, **that filter is currently a bad idea**: across our seven telephony conditions no engine we tested produced a meaningful improvement over the unprocessed audio, several degraded it by five to ten points of word error rate, and [our own model did worse than all of them](/blog/we-put-our-model-in-our-benchmark-and-it-lost). The place to stand exists. The thing worth standing there with does not exist yet, and that is the entire reason this company is building one. The second-order consequence is worth stating for anyone doing platform selection. **On AudioHook, the audio-quality decision for your highest-volume channel is made upstream of you and cannot be changed by you.** You get 8 kHz µ-law and whatever the carrier leg did to it, and the only lever you have is what you do after it arrives. #### What we could not confirm Named explicitly, because a developer-facing post that quietly fills gaps is worth nothing. - **Genesys' own transcription sample rate.** We could not find any Genesys statement of what rate their own speech-to-text runs at, and no 8 kHz versus 16 kHz comparison anywhere. Their engines are vendor-backed (Google, Microsoft Azure, AWS Transcribe, a Genesys native engine, Deepgram) and no rate is published for any of them. Trunk codecs [can include G.722 and Opus](https://help.genesys.cloud/articles/configure-the-preferred-codec-list/), so the carrier leg is not necessarily narrowband — but what reaches the transcriber is undocumented. **The 8000 Hz figure in this post is the AudioHook third-party path only.** Do not extend it to Genesys' own pipeline on our say-so. - **Any supported way to insert an SBC or media server into a BYOC SIP path to modify audio.** BYOC Cloud lets you define SIP trunks to third-party carriers, and trunk codecs are configurable. We found no Genesys document that endorses, describes or forbids putting a processing element in that path. Absence of documentation is not permission and it is not prohibition. - **A maximum number of concurrent AudioHook Monitor connections.** Only integration and monitor counts are published — five AudioHook Monitor integrations, up to 20 integration installations, up to 500 configured monitors, ten Transcription Connector integrations. Audio Connector's 300 concurrent calls is the only concurrency figure we found. - **The wire format of server-to-client audio in Audio Connector.** The specification gives the cadence (no more often than 200 ms) and the size ceiling (64,000 bytes) but never restates the format. It is presumably the negotiated one. It does not say so. - **"AudioHook Media Service"** does not appear to be a Genesys product name. We searched the developer centre's sitemap and the full help-centre article index and found nothing. If you have seen the term, it did not come from Genesys documentation. We also note that Genesys publishes a [reference server implementation](https://github.com/purecloudlabs/audiohook-reference-implementation) in TypeScript with a client tool and a test suite, and a separate [Audio Connector reference server](https://github.com/GenesysCloudBlueprints/audioconnector-server-reference-implementation) whose source pins `MediaFormat` to `'PCMU' | 'L16'`, `MediaRate` to `8000` and the maximum binary message to 64000 — which is a second, independent confirmation of the numbers above, in code rather than prose. #### If you are implementing this 1. **Count samples, do not count frames.** The spec requires it and the catch-up burst punishes anything that doesn't. 2. **Buffer into your model's block size**, and never take arrival timing as a proxy for audio timing during the first twenty seconds. 3. **Handle `discarded` as a timeline event**, not as an error to log and move on from. Your sample counter is wrong afterwards if you don't. 4. **Answer `ping` within five seconds, always**, from a path that cannot be blocked by inference. If your model and your control plane share an event loop, a slow block eats your connection. 5. **Decide which feature you are before you design anything.** Monitor and Audio Connector look like the same protocol and are opposite architectures, and the error table will not stop you from building the wrong one. 6. **Note the PCI boundary.** AudioHook [adheres to PCI DSS compliance when secure pause runs during secure flows](https://help.genesys.cloud/articles/about-audiohook-monitor/), and you cannot use it to stream audio during a secure flow. Whatever you build has a hole in it by design, and that hole is deliberate. The companion post covering [Amazon Connect, Twilio, NICE CXone and Avaya](/blog/where-you-can-insert-audio-processing) reads the same question against the other four platforms — and one of them turns out to document something Genesys does not. ### audio_stream_end=True does not end a turn in the Gemini Live API Canonical: https://anecho.ai/blog/gemini-live-audio-stream-end · Daniel Reiss · published 2026-08-16 · 2218 words · tags: gemini-live, voice-agents, vad, turn-taking, streaming Measured on Vertex against gemini-live-2.5-flash-native-audio with one 5-second clip: audio alone returns 0 bytes, audio plus audio_stream_end=True returns 0 bytes, and audio plus 1.5 seconds of trailing silence returns 229,994 bytes. The server ends a turn when its own VAD hears silence, so your uplink has to carry a pause as real samples and never as absent chunks. Google's own documentation says two different things about this field, and the Vertex docs do not mention it at all. If you are building a voice agent on the Gemini Live API and it connects, streams audio, shows a healthy session and never answers, this post is the two hours you are about to spend. The measurement first. One 5-second LibriSpeech clip, sent three ways to the same Vertex endpoint, same model, same session configuration. The only thing that varies is how the client signals that the caller has stopped talking: | What the client sent | Bytes returned | |---|---| | Audio only, no end signal | **0** | | Audio, then `send_realtime_input(audio_stream_end=True)` | **0** | | Audio, then **1.5 s of trailing silence** as actual samples | **229,994** | 229,994 bytes of 24 kHz 16-bit mono is about 4.8 seconds of reply. The middle row is the one worth staring at. The field is called `audio_stream_end`, you set it to `True`, the socket stays open, no error comes back — and the model does not respond. Ever. We waited. **The server ends a turn when its own voice activity detector hears silence.** That is the only mechanism. So the silence has to reach it, as samples, at real-time rate. A client that simply stops sending is not a client that has finished speaking; it is a client that is indistinguishable from a dead one. #### What the documentation says, which is two different things We went looking for this in the docs after we measured it, and the docs are genuinely divided against themselves. All three quotes below are from Google's own published reference and guide. The normative field reference, in [`BidiGenerateContentRealtimeInput`](https://ai.google.dev/api/live), describes a **stream state**, not a turn: > "Indicates that the audio stream has ended, e.g. because the microphone was turned off. This should only be sent when automatic activity detection is enabled (which is the default). The client can reopen the stream by sending an audio message." The [Live API guide](https://ai.google.dev/gemini-api/docs/live-guide), in its automatic-VAD section, says the same thing in operational terms — flush, not finish: > "When the audio stream is paused for more than a second (for example, because the user switched off the microphone), an `audioStreamEnd` event should be sent to flush any cached audio. The client can resume sending audio data at any time." And the same guide, in its hybrid-VAD section, says something quite different: > "The server treats the `audio_stream_end` signal as an **immediate finalization prompt, bypassing the default server-side silence detection delay** and returning the transcript and model response with minimal latency." Those cannot both be the general behaviour of one field, and the third is the only sentence anywhere that says it produces a response. Meanwhile the SDK's own type for the containing message states the model's actual contract: > "End of turn is not explicitly specified, but is rather derived from user activity (for example, end of speech)." There is one more thing we think is the real explanation for why this is so easy to get wrong. **The Vertex documentation does not mention `audioStreamEnd` at all.** We could not find the field on a single Vertex Live API page — and Vertex is where `gemini-live-2.5-flash-native-audio` lives, generally available, as the recommended model for low-latency voice agents. So on the platform we measured, the field is not documented as doing anything, and the paragraph that says it finalises a turn is on the other product's guide. We are not claiming Google's docs are wrong. We are claiming that on Vertex, with automatic VAD at its defaults, on this model, on this date, **the signal produced nothing and 1.5 seconds of silence produced a full response** — and that if you read only the sentence about immediate finalization, you will build something that hangs. We are not the first to hit it. [python-genai issue #1328](https://github.com/googleapis/python-genai/issues/1328) reports exactly this, step for step: send `audioStreamEnd` when the microphone goes off, then *"observe that the model does not end the user turn or react, even after waiting over a minute."* That issue was closed by a stale bot in January with no fix and no confirmation, so treat it as corroboration from another developer rather than as an official statement — which is precisely why we are publishing a measurement instead of a link. #### The invariant that fixes it We wrote this into the service as a one-sentence contract, because everything else in the file is downstream of it: > Once a call is live, a continuous 16 kHz PCM16 stream reaches the API at real-time rate, and a pause is carried as **actual silent samples** — never as absent chunks. That is not a style preference. It is the protocol, restated as something you can test. Every place in your code that could decline to forward a chunk is a place that can hang the call forever with the socket still open, and both of the places we had were bugs. The second consequence is the one that catches browser clients, and it is not obvious: **your client will stop sending audio during a pause whether you want it to or not.** - Chrome hands an `AudioWorkletProcessor` an **empty input array** whenever the upstream bus is flagged silent. Not a buffer of zeros — an empty array. Naive code forwards nothing. - A muted track, or a device switch, stops the callbacks outright. So "the user went quiet" and "the microphone stopped producing callbacks" arrive at your server as the same event, and the one thing the server-side VAD needs in order to answer is exactly the thing the browser has decided not to give you. Our fix is a keep-alive that **synthesises** the silence the client should have sent and pushes it through the *identical* path — the same resamplers, the same model — so the pipeline stays coherent and the stream carries true silent samples rather than a hole. Two details matter more than the idea: - **Key it on client arrival, never on your own uplink.** Our uplink is legitimately bursty, because [the stream resampler emits in roughly 95 ms bunches](/blog/the-resampler-that-ate-92-percent-of-a-microphone). A keep-alive that watched its own output would read every burst gap as a stalled client and inject silence into the middle of live speech. Ours triggers after 150 ms of nothing from the browser, which is about nineteen missed 128-sample quanta: a stall, not jitter. - **Cap the fill per tick.** A long stall gets repaired steadily rather than as one enormous late blob that arrives out of proportion to the conversation. And report it. The amount of silence synthesised per call is a number in our health payload, not a hidden repair, because a session that is 40% synthetic silence is telling you something about the client that you want to know. #### The interlock nobody expects Coalescing came next, for a different reason — two arms of an A/B were being packetised differently, [10.6 against 125.0 frames per second](/blog/the-resampler-that-ate-92-percent-of-a-microphone) — and both arms now assemble into fixed 100 ms frames before they are sent. Google's own guidance is *"send small chunks (between 20 ms and 40 ms) to minimize latency"*, and 125 frames per second is far outside it in the other direction. But a part-built frame is audio that has **not** been sent. So coalescing is only safe *because* of the keep-alive: without a producer behind the buffer, the end of the caller's last sentence sits in a half-full frame at exactly the moment the server needs it in order to detect that the caller has stopped. The keep-alive guarantees the buffer always has a producer, so a partial frame is always flushed by the silence that follows it. Two independent-looking changes that are not separable. If you take the frame coalescing without the keep-alive you have built a subtler version of the same hang. #### Make the turn-detection settings explicit The last thing we changed was not code so much as visibility. This service used to leave every VAD parameter implicit, which is how "the agent never answers" was able to look like a mystery instead of a setting. The documented surface, for reference: | Field | What the docs say | |---|---| | `automaticActivityDetection.disabled` | Defaults to `false` — server-side VAD is on | | `silenceDurationMs` | *"The server's internal default is approximately 800ms"*; the guide recommends 500–800 ms | | `prefixPaddingMs` | No default documented that we could find | | `startOfSpeechSensitivity` / `endOfSpeechSensitivity` | The SDK states different defaults per platform — `LOW` on the enterprise branch, `HIGH` on the Gemini API branch | | `activityStart` / `activityEnd` | *"can only be sent if automatic (i.e. server-side) activity detection is disabled"* | Two traps in that table. First, the guide's own configuration example uses `prefix_padding_ms: 20` and `silence_duration_ms: 100` — illustrative values, and 100 ms sits inside the range the same page calls *"Too low"*, where *"the system ends speech turns during natural pauses, splitting a single utterance into multiple small audio fragments."* Do not copy the sample into production. Second, manual activity detection is a different world, not a supplement. With `disabled: true` you send `activityStart` and `activityEnd` yourself, and — per the guide — *"an `audioStreamEnd` isn't sent in this configuration."* If you own the VAD, you own the turn boundary and the silence problem goes away; you have simply bought the harder half of the problem, which is [detecting turns on a noisy line](/blog/the-split-pipeline). Our own choice is to expose all five as environment variables and per-call query parameters, and to leave them **unset by default**, so the status endpoint reports "vertex default" rather than a number the service invented. The two configurations we have measured working were measured at Vertex's defaults, and quietly moving them would change a result that is now proven. #### The rest of the audio contract, since you will need it anyway From the documentation, and it is worth getting right in one pass: - **Input:** raw, little-endian, 16-bit PCM. *"Input audio is natively 16kHz, but the Live API will resample if needed so any sample rate can be sent."* Declare the rate in the MIME type of every blob, e.g. `audio/pcm;rate=16000`. The Vertex troubleshooting page adds the channel count: a single mono channel. - **Output:** raw, little-endian, 16-bit PCM, **always 24 kHz**. Note the asymmetry — you send 16 kHz, you receive 24 kHz. Any downstream mixing or recording has to resample one of them. - **Chunking:** 20 to 40 ms; do not buffer around a second before sending. - **Session length:** without context-window compression, audio-only sessions are documented at 15 minutes and audio-video at 2. Separately, *"the lifetime of a connection is limited to around 10 minutes due to WebSocket connection constraints"*, with a `goAway` notice sent 60 seconds before the end. Those are two different limits and the shorter one is the connection, so **reconnection is a normal part of a long call, not an error path.** We treat it as one: bounded, drop-oldest mic queue across the gap, so a reconnect does not "recover" by replaying a growing backlog and answering questions from ten seconds ago. One caveat on where these docs live: the Vertex Live API pages have been rebranded and moved, and the old `cloud.google.com/vertex-ai/generative-ai/docs/live-api` path now redirects. If a link in your notes is dead, that is why. #### Caveats - **This is n = 1 on the stimulus.** One 5-second clip, three conditions, one model (`gemini-live-2.5-flash-native-audio`), one region, one date. It is a clean, reproducible demonstration of a behaviour, not a survey. We would not report the byte count as if it were a benchmark. - **The 0 / 0 / 229,994 result is a demonstration that the silence works and the flag did not**, on that configuration. It does not establish that `audio_stream_end` never does anything, on any model, at any setting. If you can show it finalising a turn on Vertex, we want to see the configuration. - **The behaviour is measured on Vertex.** The Gemini API branch documents different sensitivity defaults, and the hybrid-VAD paragraph that describes finalization is on that branch's guide. - **The community issue is corroboration, not confirmation.** It was closed by automation with no resolution. - **We are not measuring quality here.** Nothing in this post says anything about how well the model heard the caller. That is a separate experiment and [the audio path in front of it](/blog/8khz-is-where-voice-ai-breaks) is what we actually work on. #### The short version 1. **The server's own VAD is the only thing that ends a turn.** Design for that and nothing else. 2. **Carry pauses as samples.** If your client goes quiet, synthesise the silence and send it through the same path as real audio. 3. **Never key a keep-alive on your own uplink** — a bursty transmit path will look like a stalled client and shred live speech. 4. **Coalesce to a fixed frame, but only once something guarantees the buffer has a producer.** The two changes are one change. 5. **Publish a delivered-audio ratio per session.** Uplink seconds over wall-clock seconds has a known correct value near 1.0, which makes a wrong value legible. Ours read 0.08 for days while the UI cheerfully reported a live session, and that is [its own post](/blog/the-resampler-that-ate-92-percent-of-a-microphone). 6. **Do not copy the VAD example values.** Leave them unset, report what is in force, and change them deliberately. ### Null Test v0.1: ten speech enhancers, eighteen conditions, and nothing beat raw on the pooled average Canonical: https://anecho.ai/blog/nulltest-open-benchmark · Anecho Engineering · published 2026-08-13 · 3125 words · tags: nulltest, benchmark, methodology, open-source, release The harness, the dataset, the configs and now the results. Ten enhancement engines scored against a raw control on WER, error decomposition, VAD and cost. Raw won on pooled WER — but at least one engine beat raw in 11 of the 18 individual conditions, and insertions went up rather than down. *Updated 14 August 2026: the matrix was regenerated with five additional telephony carrier conditions, taking it from thirteen to eighteen. Absolute numbers moved; the ordering did not. Every figure below is quoted from the published `benchmark.json`.* Today we are publishing Null Test v0.1 — the harness, the dataset manifest, the configs, the rules, and the first full set of results. It measures what happens to speech-to-text accuracy when you put a speech enhancement engine in front of a recognizer, across eighteen acoustic conditions, against a null-hypothesis control that is allowed to win. On the pooled numbers, it won — and the per-condition breakdown underneath that is where the actual engineering lives. A null test is the canonical audio-engineering measurement: subtract the processed signal from the reference and whatever remains is exactly what the processor did — including the damage. That is the whole ambition of this project, so it is the name. #### Why this exists The public claims are irreconcilable. Krisp markets roughly a 46% WER reduction. ai-coustics markets roughly 43%. An independent study ([arXiv:2512.17562](https://arxiv.org/abs/2512.17562)) found enhanced audio scored worse than raw in all 40 configurations it tested — though only with MetricGAN+, on medical speech, in semantic WER. AssemblyAI, citing it, separately reported that Krisp's output roughly doubled WER when fed to its STT while cutting false VAD triggers about 3.5x. Deepgram has published the position that enhancement hurts accuracy, and we have not found accompanying data. Nobody shares a corpus, a condition set, a configuration, or a control. The long version of that disagreement is in [Does noise suppression actually help speech-to-text?](/blog/does-noise-suppression-help-stt). #### What was run | | | |---|---| | Recogniser | faster-whisper `base.en`, CTranslate2 int8 CPU, greedy decoding | | VAD | Silero VAD, threshold 0.5, 512-sample blocks | | Test set | 10 speakers, 180 clips, roughly 1406 seconds of audio, seed 20260101 | | Reference words | about 190 per condition, 3420 total per backend | | Conditions | 18 | | Backends | 12 (raw, passthrough control, and ten enhancers) | | Host | Apple M1 Max, onnxruntime 1.28.0, pinned to 1 intra-op thread | The eighteen conditions: `clean`; broadband noise at -5, 0, 5, 10 and 20 dB SNR; `babble_snr5`; competing speaker at 0 and 5 dB; `reverb` and `reverb_snr10`; and seven telephony conditions — `telephony` (a G.711 µ-law round trip), `telephony_snr10`, plus A-law, G.722, Opus at 12 kbps, and bursty packet loss at 3% and 10%. The ten enhancers: GTCRN (MIT), five FastEnhancer variants T/S/B/M/L (MIT), DeepFilterNet3 in two configurations (MIT or Apache-2.0), and two ai-coustics models, Quail L and Quail VF 2.2 L, run through the licensed SDK as a proprietary reference baseline. **Our own models are not in this matrix.** Chamber and Clearline will be rows in it, under the same rules, against the same control that just beat everybody. Publishing the referee before the contestant was the point of the ordering. #### The results Pooled across all eighteen conditions: | Backend | WER | Insertions | Deletions | RTF | Latency | |---|---|---|---|---|---| | Raw (no processing) | 14.77% | 101 | 60 | — | 0 ms | | ai-coustics Quail L | 15.56% | 126 | 57 | 0.124 | 30 ms | | ai-coustics Quail VF 2.2 L | 18.10% | 161 | 75 | 0.076 | 30 ms | | GTCRN (MIT) | 20.26% | 210 | 40 | 0.050 | 16 ms | | FastEnhancer-L (MIT) | 20.44% | 159 | 107 | 0.376 | 25.8 ms | | FastEnhancer-M (MIT) | 22.40% | 217 | 83 | 0.110 | 22 ms | | FastEnhancer-T (MIT) | 22.46% | 145 | 109 | 0.013 | 16 ms | | DeepFilterNet3 | 35.18% | 380 | 197 | 0.133 | 100 ms | **No enhancer beat raw on pooled WER.** The best was ai-coustics Quail L at 15.56% against raw's 14.77% — about eight tenths of a point worse, which on this sample size we would call a tie rather than a loss. Everything below it lost by five points or more, and that is not a tie. Pooled answers "should enhancement be on by default" and almost nothing else; the per-condition sections below are the ones to act on, and they say something different. **Insertions went up, not down.** Raw produced 101. GTCRN produced 210. DeepFilterNet3 produced 380. This is the direct opposite of the vendor claim that enhancement removes hallucinated words, and it is worth dwelling on: an insertion is not a garbled word, it is the recogniser emitting a word that nobody said. Our reading is that what a suppressor leaves behind — musical noise, gated silence, smeared transients — is more decodable-as-speech than the noise it removed. GTCRN's decomposition makes the point: 210 insertions against only 40 deletions, where raw had 101 and 60. It is not eating your function words, it is adding words nobody spoke. Opposite failure modes, and one WER number hides which one you bought. That is why every cell publishes S, D and I separately. #### Where the vendor claim does hold Stopping at the pooled number would be doing exactly what we criticise the vendors for. Enhancement wins real margins in specific conditions — **in 11 of our 18, at least one engine beat raw** — and they are the conditions you would predict: where noise or a competing voice is genuinely destroying phonetic cues. The clearest is speaker isolation. ai-coustics Quail VF is a Voice Focus model: its job is to isolate the primary speaker and reject other voices. On competing speaker at 5 dB it cut insertions from **19 to 10** and WER from **26.3% to 20.5%**. On babble at 5 dB, **24.7% to 20.5%** with insertions down 10 to 7. Those are the strongest results any enhancer posted here. Low-SNR broadband noise shows the same shape — at 0 dB, eight of the ten engines beat raw, the best by 4.7 points. The seven conditions where nothing beat raw are just as informative: the two reverberant ones, and the five telephony carrier conditions. In four of those five the best an engine manages is an exact tie with doing nothing, which is its own kind of answer. What none of that supports is the general claim. Pooled, Quail VF is more than three points worse than raw, because it also runs on clean, reverberant and narrowband material where it can only remove information. **Enhancement pays where the audio is genuinely bad and costs you where it is not**, and a single scalar — ours included — hides which side of that line your traffic sits on. #### VAD: the metric everyone quotes barely moves, the one that matters moves a lot The same run scored Silero VAD on every stream. Pooled VAD F1: **0.950 raw against 0.947** for the best enhanced stream. Grade enhancement on VAD F1 and you conclude it does nothing. The false-alarm rate is a different story, and a false alarm is what a spurious barge-in actually is: | Condition | False-alarm rate, raw | Enhanced | Engine | |---|---|---|---| | Babble at 5 dB | 97.7% | 68.2% | FastEnhancer-L | | Competing speaker at 5 dB | 58.1% | 37.2% | ai-coustics Quail VF | Under babble at 5 dB the raw stream false-alarms on essentially every non-speech frame — a VAD that has stopped working as a gate. Enhancement takes that to 68.2%, still bad and thirty points better. So the same run says enhancement costs you five points of pooled WER and buys you a third off your false-alarm rate. That is not a contradiction, it is an architecture: enhanced audio to the turn detector, raw audio to the recogniser. We shipped that as the SDK default and the reasoning is in [The split pipeline](/blog/the-split-pipeline). #### Telephony: no enhancer helped, and several hurt badly On the G.711 µ-law round trip, raw is 5.8% WER. Against that: ai-coustics Quail L +1.6 points, Quail VF +2.6, GTCRN +3.2, FastEnhancer-T +6.8, DeepFilterNet3 +9.5. On the same round trip with noise at 10 dB, raw is 10.5% and the spread runs from +0.5 to +9.5. **No enhancer produced a meaningful improvement on telephony audio, and several caused large degradations.** Two rows come out fractionally under raw and neither is a result: FastEnhancer-S is -0.5 on the clean row — at 190 reference words, inside a single word of doing nothing — and +3.7 on the noisy one; Quail VF is -1.1 on the noisy row and +2.6 on the clean one. Ties, in both directions. The five carrier conditions added in this update say the same thing more bluntly. On A-law, G.722, Opus at 12 kbps and 3% bursty loss, the best any engine manages is an exact tie with raw; on 10% loss nothing reaches it. The worst cells are severe — FastEnhancer-M posts **49.0%** on A-law against raw's 7.4%, with 61 insertions where raw has 2. Companded quantisation noise is outside what that model expects and it responds by hallucinating. That matters because telephony is where the call volume is, and the shape of it is not an accident: the DNS Challenge, the benchmark series most published denoisers are tuned against, never had a narrowband track. The field optimised for 16 and 48 kHz because that is what was scored. It matters with a large asterisk. Raw WER of 5.8% on the base telephony condition means it is an easy one, and not the thing that breaks production phone deployments — see the caveats. The full narrowband story, including a separate 8 kHz end-to-end experiment with native narrowband engines and a resample round-trip control, is in [8 kHz is where voice AI actually breaks](/blog/8khz-is-where-voice-ai-breaks). Its short version: native 8 kHz beats resample-and-hope by 0.7 points for the same vendor, and nothing at all beats the untouched caller audio. #### Caveats, in full A result without its limits is an advertisement. These attach to every number above. - **Single recogniser.** faster-whisper `base.en`. A different recogniser gives different absolute numbers. The harness supports Deepgram and AssemblyAI so the *ordering* can be re-checked, and we would like someone who does not trust us to do it. - **About 190 reference words per condition**, 3420 pooled. Enough to separate large effects, not enough to resolve one or two points of WER: one word is worth roughly 0.5 points per condition and 0.03 pooled. Read the raw-versus-Quail-L gap as a tie; the five-to-twenty point gaps below it are not. - **DeepFilterNet3 is measured through a block-online adapter.** No public ONNX export carries recurrent state tensors, so per-frame streaming is impossible with the published weights. Its 35.18% is a **lower bound on quality and an upper bound on cost**, and its 100 ms latency is far above what a native Rust runtime achieves. Never quote the number without this sentence. It also runs at 48 kHz over 16 kHz material upsampled to 48 kHz — the honest VoIP situation, not the condition it was designed for. - **ai-coustics rows are a proprietary baseline** run through the licensed SDK. A reference point, never a shipping code path for us. - **Licensing constrains reuse.** ESC-50 noise is CC-BY-NC-3.0, so those conditions are evaluation only. LibriSpeech is CC-BY-4.0, and the babble and competing-speaker conditions are built from LibriSpeech alone, so they carry no NC restriction. - **onnxruntime pinned to 1 intra-op thread**, sequential execution. The RTF column is one stream on one core, not a fan-out. - **English only, batch scoring, no listening panel, no turn-level metrics.** Enhancement is driven block by block with state carried across blocks, but the recogniser is scored on complete utterances, and VAD is frame-level. "Fewer false barge-ins" as a product metric needs labelled turn boundaries and an agreed definition. - **The telephony conditions are more realistic than they were, and still not a phone call.** They now include A-law and G.722 legs, Opus at 12 kbps, and bursty two-state packet loss at 3% and 10% with attenuated-repeat concealment — rather than independent random loss with silence fill, which flatters decoders and overstates damage respectively. **Still missing: acoustic echo with non-linear speaker distortion**, handset-side suppression, and AMR-WB or EVS legs. Echo is the gap we care most about, because it is the impairment behind the part of our own roadmap that cannot be closed with a patch. - **These eighteen conditions are stored at 16 kHz** after an 8 kHz carrier leg, so every backend sees them at its own native rate. The genuinely narrowband experiment — 60 clips, 8 kHz throughout, nothing resampled, native 8 kHz engines included — is a separate artefact with its own resolution limit of 1,140 reference words, where one word is 0.09 points. #### Methodology, in enough detail to attack **Streaming is simulated honestly.** Every enhancer is driven block by block at its declared block size, carrying state across blocks, with no lookahead beyond its declared algorithmic latency. Offline whole-file processing flatters an enhancer and is not what a live agent gets. Latency is reported as algorithmic latency — frame plus lookahead — a property of the model, not of our laptop. **RTF is measured, not quoted.** Single pinned core, warmup passes discarded, hardware string recorded in every row. A model that wins on WER at 4x real time is not a shipping option, so the accuracy number and the cost number travel together. **Noise is mixed, not sourced pre-mixed.** Clean speech plus noise at a target SNR computed over active speech regions only, so leading and trailing silence does not distort the ratio. Reverberation is convolution with measured room impulse responses. Telephony is a real codec round trip — resample to 8 kHz, G.711 µ-law encode and decode, resample back — rather than a low-pass filter standing in for a channel. A real codec leg, and still not a real phone call; the caveat above lists what it leaves out. **Text normalization is fixed and published.** WER is `(S + D + I) / N` after one normalizer applied identically to every hypothesis and reference: case folding, punctuation removal, number expansion, contraction handling. Normalization can move WER by multiple points, so it is versioned in the repo and recorded in every row. **Determinism, and per-utterance outputs.** Fixed seeds for noise selection, SNR draw and RIR selection; the dataset manifest is content-addressed with SHA-256 per file, verified on fetch, so the same config on the same commit produces the same rows. Every run also emits per-utterance hypotheses and references alongside the aggregates, so anyone can recompute, slice differently, or find the ten files carrying a result. Aggregate-only benchmarks cannot be checked, which is most of how the field got here. #### The rules that make it a referee 1. **The raw control is always in the matrix and is allowed to win.** In v0.1 it won. 2. **We score ourselves under the same rules, in the same tables.** Chamber and Clearline will be rows, not a separate marketing page. 3. **Every cell publishes, including the ones we lose.** No condition gets dropped after we see the result. 4. **Configs are in the repo.** If a vendor believes we configured their engine badly, the fix is a pull request against a YAML file, and we rerun and republish. 5. **Right of reply.** Any maintainer or vendor can request a rerun at a different setting. We publish both. 6. **Licensing is stated per row.** Where a vendor's terms prohibit publishing benchmark results, we say the cell is blocked and name the reason rather than quietly omitting it. #### The output format The full document is at [/benchmark](/benchmark). One backend's summary: ```json { "id": "gtcrn", "label": "GTCRN", "license": "MIT", "wer": 0.2026, "werSubstitutions": 443, "werInsertions": 210, "werDeletions": 40, "werRefWords": 3420, "vadF1": 0.9472, "vadFalseAlarm": 0.1638, "rtf": 0.0498, "latencyMs": 16.0 } ``` Each backend also carries a `conditions` object with the same decomposition per condition — where the interesting disagreements live. #### Running it yourself ```bash git clone https://github.com/anecho/nulltest cd nulltest uv sync # fetch and verify the dataset against the content-addressed manifest python -m nulltest.data fetch --manifest data/manifest.json # one condition, two backends, one recogniser — a few minutes on a laptop python -m nulltest.run \ --condition babble_snr5 \ --backend gtcrn --backend raw \ --stt faster-whisper:base.en \ --out out/ python -m nulltest.report out/ --format md ``` Adding an engine is one file. The benchmark only ever talks to this interface: ```python class Enhancer(Protocol): id: str sample_rate: int block_size: int # samples per process() call algorithmic_latency_ms: float def reset(self) -> None: ... def process(self, block: np.ndarray) -> np.ndarray: """block: float32 mono, shape (block_size,), range -1 to 1. Returns same shape.""" ``` If your engine implements `reset` and `process`, it can be in the matrix. We would particularly like the vendors whose published claims we quoted to submit their own adapters and configurations. #### What happens next, and the commercial part v0.2 adds turn-level endpointing metrics via Onset, streaming STT with latency-to-final, a second and third recogniser so the ordering can be checked rather than trusted, multi-language conditions, and an echo condition with non-linear speaker distortion. Chamber and Clearline enter the matrix when they have something worth scoring. On that last point, plainly, because a benchmark run by a vendor is worth nothing if the vendor is coy about its own progress: **Clearline's first proof of concept has trained and converged, and it is weak.** It reaches **+1.41 dB SI-SDR** over the unprocessed mixture on 1,200 held-out validation mixtures, where competent target-speaker-extraction systems reach +8 to +12 dB. It has no WER number because it is not yet worth scoring through this harness, and it will not appear in any table above until it is. Architecture experiments are running. The commercial part, stated once and without decoration: this benchmark is not the product. A benchmark has no buyer and no retention — nobody renews a leaderboard. What we intend to sell is [Clearline](/blog/8khz-is-where-voice-ai-breaks), a telephony audio path: native 8 kHz primary-speaker isolation, residual echo suppression, per minute, with an instant API key and no sales call. Note exactly what v0.1 licenses us to say about that. It says nothing on the shelf helps on a narrowband channel, and that nobody ships speaker isolation at 8 kHz at all. It does not say ours will be better — that number does not exist until Clearline is a row here, on telephony conditions worth the name, against a control allowed to beat it. Null Test is how we found out what to build, and how you check whether we are lying about it. Harness and issues: [github.com/anecho](https://github.com/anecho). Full matrix: [/benchmark](/benchmark). Quickstart: [/docs](/docs). If you want a condition, an engine or a recogniser added, open an issue — including if you expect it to make us look bad. ### 8 kHz is where voice AI actually breaks Canonical: https://anecho.ai/blog/8khz-is-where-voice-ai-breaks · Sofia Marchetti · published 2026-01-20 · 3978 words · tags: telephony, pstn, narrowband, signal-processing, stt Every voice AI demo is 16 kHz from a laptop mic. Most revenue is 8 kHz from a phone line. We ran a full narrowband benchmark — nothing resampled, three native 8 kHz engines, a round-trip control — and the best path in it ties doing nothing by one word. We also tested the 'resample tax' we had been repeating, and it is not there. *Updated 14 August 2026 with the regenerated [Null Test](/blog/nulltest-open-benchmark) run (18 conditions) and with a new end-to-end narrowband experiment. **This update retracts a claim.** An earlier version of this post repeated a community report that a major framework's default resample path discards ~16% of caller audio. We built the control for it, and our measurement does not support it. The section is rewritten below rather than deleted.* Every voice AI demo you have seen runs at 16 kHz from a laptop microphone over WebRTC. Almost every dollar of production call volume arrives at 8 kHz from a phone line that has already thrown away the top half of the spectrum. Those are not the same problem, and the second one is where deployments quietly fail. The gap is not a configuration detail. It is an octave of missing signal, a codec chain nobody controls, a platform layer where — as of writing, and this is our reading of the vendors' docs — you often cannot insert your own processing at all, and a model shelf on which nothing is actually built for the channel. #### What a phone line actually removes PSTN and SIP telephony are 8 kHz narrowband, carried as G.711 µ-law or A-law, with a passband of roughly 300 to 3400 Hz. Sampling at 8 kHz puts the Nyquist limit at 4 kHz, and the channel filter takes another 600 Hz off the top. Everything above 3.4 kHz is simply not present — and that band is where a specific, unusually costly set of phoneme cues lives. | Cue | Where its energy is | Survives a 300–3400 Hz channel? | |---|---|---| | `/s/` versus `/f/` versus `/θ/` | sibilant energy peaks roughly 4–8 kHz | No. This is the classic narrowband confusion. | | Vowel identity (F1, F2) | 250–2500 Hz | Yes | | `/r/` and F3 cues | roughly 2000–3500 Hz | Marginal, right at the cutoff | | Plosive burst transients | broadband, much of it above 4 kHz | Partly; bursts are flattened | | Nasal murmur | below 1 kHz | Yes | | Speaker fundamental F0 | 85–255 Hz for most adults | Usually below the 300 Hz cutoff; reconstructed perceptually from harmonics | So narrowband errors are not uniformly distributed across your transcript. They cluster exactly where voice agents carry the most business risk: digits ("six" and "fix" lose their most reliable discriminator), spelled-out names and email addresses, confirmation codes, and plurals that hinge on a word-final `/s/` or `/z/`. A 3% aggregate WER lift can be a 30% error rate on the one field the call existed to capture. The other half of the problem is the recognizer. Whisper and most modern STT models are trained on 16 kHz audio, so narrowband speech is out of distribution before any noise is added. #### We tested the resample tax, and it is not there An earlier version of this post led with a LiveKit community PSA of 7 August 2026 reporting that the default `FrameProcessor` resample pattern continuously discards audio — roughly **16% of the caller's speech**, permanently, on every call — with one operator in the thread recovering nearly eighteen points of STT accuracy by fixing it. We repeated that. It was the most quotable thing in the post. Then we built the control, and our own measurement does not support it. The control is simple. Take the narrowband caller audio, send it 8 kHz → 16 kHz → 8 kHz, and measure the SI-SDR of what comes back against what went in. If a resample round trip destroys a fixed fraction of the signal, that number is bad, and it is worse for the sloppier resampler. | 8 kHz → 16 kHz → 8 kHz | SI-SDR vs its own input | |---|---| | `naive` — sample-repeat up, sample-drop down, no filtering at all | **150 dB** | | `linear` — linear interpolation, no anti-alias filter | **150 dB** | | `soxr_hq` — polyphase/sinc, library default | 38.45 dB | | `soxr_vhq` — polyphase/sinc, the correct implementation | **38.28 dB** | Read that twice. The deliberately broken resampler scores a **perfect null** — 150 dB is the ceiling our harness reports when the difference signal is numerically zero — and the correct one scores **112 dB worse**. Both results are right, and the reason is arithmetic rather than measurement error. At an exact 2:1 ratio, sample-repeat followed by sample-drop is *algebraically the identity function* — every original sample comes back bit-exact, so there is nothing left to measure. The polyphase resampler scores lower precisely because it does its job: it applies an anti-alias filter, and a filter changes the signal. **Round-trip SI-SDR, used naively, rewards the broken implementation.** That is the trap, and it is why this needed a control rather than an intuition. (At 44.1 kHz, where the ratio is not an integer, `linear` drops to 32.49 dB — the identity property is specific to the 2:1 case that telephony actually hits.) Downstream, on transcription, it is the same story. With no model in the path at all, resampler choice is worth **0.18 points of WER**: `soxr_vhq` scores 10.96% against `linear`'s 11.14%, versus a raw 8 kHz control of 11.14%. On 1,140 reference words that is **two words**. It is a tie. And with a model in the path, the result inverts. Across all six model-and-resampler pairs we tested, the *naive* resampler transcribed **better** — mean −3.0 points, every single pair negative, from −0.4 (`aic-quail` with `naive`) to −5.6 (`fastenhancer-b` with `linear`). The likely mechanism is that a polyphase resampler is an extra filtering stage the model was never trained on, while 2:1 nearest-neighbour is an identity, so the naive path hands the model a *less* altered signal. So: we are not claiming resampling destroys caller audio, we are not repeating the 16% figure as if it were ours, and we will not sell resampling correctness as recovered accuracy. We do not know what the PSA author's instrumentation was measuring — frame accounting inside a specific pipeline is not the same experiment as ours, and their fix plainly helped their deployment. What we can say is that the general claim, in the form we repeated it, does not reproduce here. If you take one action from this post, it is still: record the audio your recognizer actually receives, not what your SIP trunk delivers, and compare their durations. Frame accounting is worth checking. It is just not worth a WER point. #### The upsampling that fixes nothing Even a correct resampler restores nothing. You resample 8 kHz to 16 kHz because the model demands 16 kHz input; that satisfies the interface and adds no information. The 4 to 8 kHz octave stays empty, because there was never any energy there to interpolate. What the acoustic model sees is a 16 kHz spectrogram with a dead top half and a hard shelf at 3.4 kHz — a pattern that appears in its training data mostly as a degradation, if at all. ```text 16 kHz mic capture 8 kHz PSTN, upsampled to 16 kHz 8k ┤▓▓▒▒░░ fricatives 8k ┤ (nothing) 6k ┤▓▓▓▒▒░ 6k ┤ (nothing) 4k ┤▓▓▓▓▒▒ 4k ┤──────────── hard shelf 2k ┤████▓▓ formants 2k ┤████▓▓ formants 0 ┤█████▓ 0 ┤ ███▓▓ (HPF at ~300 Hz) └──────── time └──────── time ``` So "we support telephony" in most stacks means "we resample and hope." The interesting failure is not the resampler — we just spent a section establishing that — it is that the model on the other side of it is being fed a distribution it never trained on. #### What we measured: nothing on the shelf helps on the phone line We ran the telephony path through [Null Test](/blog/nulltest-open-benchmark), our open benchmark. Seven of its eighteen conditions are telephony: a G.711 µ-law round trip — resample to 8 kHz, encode, decode, resample back — the same round trip with noise at 10 dB SNR, and five carrier variants (A-law, G.722, Opus at 12 kbps, and bursty packet loss at 3% and 10%). Provenance, because these numbers are worth exactly what their method is worth: single recogniser, **faster-whisper base.en** (CTranslate2 int8 CPU, greedy), Silero VAD, 10 speakers / 180 clips / roughly 1406 seconds of audio, about 190 reference words per condition, 18 conditions, 12 backends, Apple M1 Max, onnxruntime 1.28.0 pinned to one intra-op thread. At that sample size one word is worth roughly 0.5 points of WER; read the delta columns accordingly. | Backend | G.711 | Δ vs raw | G.711 + noise 10 dB | Δ vs raw | |---|---|---|---|---| | Raw (no processing) | 5.8% | — | 10.5% | — | | FastEnhancer-S (MIT) | 5.3% | -0.5 | 14.2% | +3.7 | | FastEnhancer-M (MIT) | 5.8% | 0.0 | 12.1% | +1.6 | | ai-coustics Quail L | 7.4% | +1.6 | 12.1% | +1.6 | | ai-coustics Quail VF 2.2 L | 8.4% | +2.6 | 9.5% | -1.1 | | GTCRN (MIT) | 8.9% | +3.2 | 11.1% | +0.5 | | FastEnhancer-T (MIT) | 12.6% | +6.8 | 17.4% | +6.8 | | DeepFilterNet3 | 15.3% | +9.5 | 20.0% | +9.5 | **No enhancer produced a meaningful improvement on telephony audio, and several caused large degradations** — FastEnhancer-T at +6.8 points on both conditions, DeepFilterNet3 at +9.5 on both. Be precise about the two apparent wins, because they are the kind of thing that gets quoted without its error bar. FastEnhancer-S is -0.5 points on the clean row — inside one reference word of doing nothing — and 3.7 worse once noise is added. Quail VF is -1.1 on the noisy row and +2.6 on the clean one. Neither is a win, neither is a loss. The useful reading is that the upside is zero-shaped and the downside runs to ten points. The five carrier conditions we added in this run tighten that reading rather than changing it. In none of them does any engine beat raw, and in four of the five the best an engine manages is an exact tie: | Condition | Raw | Best engine | Worst engine | |---|---|---|---| | A-law | 7.4% | 7.4% (Quail L, tie) | 49.0% (FastEnhancer-M) | | G.722 | 5.3% | 5.3% (FastEnhancer-B, tie) | 6.8% (FastEnhancer-M) | | Opus 12 kbps | 6.8% | 6.8% (FastEnhancer-S, tie) | 12.1% (GTCRN) | | 3% bursty loss | 11.1% | 11.1% (Quail VF, tie) | 24.7% (DeepFilterNet3) | | 10% bursty loss | 10.0% | 10.5% (Quail L) | 35.3% (FastEnhancer-T) | The FastEnhancer-M row on A-law is worth staring at: 49.0% against a raw baseline of 7.4%, driven by 61 insertions where raw has 2. Companded quantisation noise is not what that model expects, and it responds by hallucinating. This is the shape of the whole category on a phone line — the ceiling is a tie and the floor is a catastrophe. DeepFilterNet3's number needs its caveat every time: we measure it through a **block-online adapter**, because no public ONNX export carries recurrent state tensors, so per-frame streaming is impossible with the published weights. Its rows are a lower bound on quality and an upper bound on cost. The ai-coustics rows are a proprietary baseline run through the licensed SDK — a reference point, never a shipping code path for us. #### The experiment that table was missing: 8 kHz end to end Raw WER on our G.711 condition is 5.8%. That is an **easy** condition, and every backend in the matrix above is a wideband model pointed at a narrowband signal. It says something specific and limited — wideband enhancers do not help on a narrowband channel and can hurt a lot — and nothing about whether processing natively at 8 kHz beats resampling. So we built the experiment that answers that. Sixty clips, six narrowband conditions, **8 kHz from end to end with nothing resampled anywhere in the arm**, three genuinely native 8 kHz models, and the STT ingest resampler held constant at `soxr_vhq` in every arm including the controls. 1,140 reference words, so **one word is worth 0.09 points of WER** — that ratio travels with every number below. | Path | WER | vs raw 8 kHz | |---|---|---| | **raw 8 kHz (control)** | **11.14%** | — | | resample only, no model, `soxr_vhq` | 10.96% | −0.18 | | resample only, no model, `linear` | 11.14% | 0.00 | | **native 8 kHz** `quail-l-8khz` | **11.05%** | **−0.09** | | **native 8 kHz** `quail-s-8khz` | 11.32% | +0.18 | | **native 8 kHz** `rook-l-8khz` | 12.46% | +1.32 | | upsampled `quail` (16 kHz weights), `soxr_vhq` | 11.75% | +0.61 | | upsampled `quail-vf` (16 kHz weights), `soxr_vhq` | 12.98% | +1.84 | | upsampled `gtcrn`, `soxr_vhq` | 16.58% | +5.44 | | upsampled `fastenhancer-b`, `soxr_vhq` | 17.63% | +6.49 | | upsampled `fastenhancer-m`, `soxr_vhq` | 21.14% | +10.00 | | upsampled `fastenhancer-t`, `soxr_vhq` | 24.82% | +13.68 | Three things fall out of it, and only one of them is the one we went looking for. **Native 8 kHz does beat resample-and-hope — by 0.7 points.** Same vendor, same model family, the only difference being 8 kHz weights against 16 kHz weights on a resampled path: 11.05% against 11.75%. As far as we can find, that comparison has never been published by anyone. It is also **about eight words out of 1,140**, which makes it a real, directional, modest result and not a headline. It is a reason to build native. It is not a reason to buy anything. **Nothing beat the untouched caller audio.** The best path in the table, `quail-l-8khz` at 11.05%, sits 0.09 points below the raw control's 11.14% — **one word**. That is a tie, not a win. Every other path is worse than doing nothing. On a narrowband phone call, in August 2026, the null hypothesis is undefeated. **And there is barely anything to test.** Look at how short the native section is. Those three models are the *entire population* of genuinely narrowband enhancement weights we could obtain. GTCRN, FastEnhancer and DeepFilterNet3 have **no 8 kHz checkpoint at all** — not a worse one, none — so on a phone call they can only be run through a resample wrapper, which is why they appear in this table exclusively as upsampled rows. One vendor ships narrowband weights. Nobody ships narrowband *speaker isolation*: every ai-coustics Voice Focus model and their primary-speaker VAD is 16 kHz only, while the same vendor does ship 8 kHz for their general enhancer and their human-listening model. Krisp's telephony answer is a variant explicitly positioned as lighter than their main line. That is the actual finding of this post, and it took us a while to see it because we were busy looking for damage. **The problem on the phone line is not that something is breaking your audio. It is that the thing you would want to run on it was never built.** What the matrix still omits is best seen against the real signal chain below, of which our conditions model some stages and not others. #### Nobody in the chain is on your side Even before the band limit, the signal has been through several processing stages you do not control and cannot inspect. | Stage | Typical processing | Artifact class it introduces | Modelled now? | |---|---|---|---| | Handset | AMR-WB or EVS encode, device noise suppression, AGC | already-suppressed speech, musical noise, pumping gain | No | | Carrier interconnect | transcode to G.711 µ-law or A-law, G.722, low-bitrate Opus; downsample to 8 kHz | band limit, companding quantization noise | **Yes** | | SIP trunk | jitter buffer, packet loss concealment, DTX, comfort noise | time-domain splices, repeated frames, synthetic noise floor | **Yes** — bursty two-state loss at 3% and 10% | | Echo path | hybrid or speakerphone return of the agent's own TTS | fluent, confident insertions | **No — the biggest remaining gap** | | Your stack | resample to 16 kHz, often a second suppression pass | empty upper band, cascaded suppression | Partly | The carrier and packet-loss rows moved from "No" to "Yes" in this update, and they are worth a note each. **Packet loss concealment is not silence.** When a packet is lost, the jitter buffer synthesises a replacement from previous frames: a plausible waveform with no linguistic content, which a recognizer will happily decode words out of. These are insertions no enhancer can fix, because the signal is not degraded — it is fabricated. We model this as bursty two-state (Gilbert-Elliott) erasure of 20 ms frames with attenuated-repeat concealment, rather than independent random loss with silence fill: independent loss flatters every decoder and silence fill overstates the damage. **The handset already suppressed.** The audio is usually suppressed once on the device before you touch it, so you are cascading a second neural suppressor over the artifacts of a first. Nobody has characterised that publicly, us included — though our results do show that a suppression pass over material that does not need one is reliably expensive; see [The split pipeline](/blog/the-split-pipeline). **Echo is the row that is still missing, and it is the one we most want.** Acoustic echo with non-linear speaker distortion is the impairment behind the half of this product that cannot be fixed with a patch, and we have not built the condition for it yet. Until we do — and until Clearline is a row in these tables against a `passthrough` control allowed to beat it — we have measured the shelf, not our answer to it. #### The platform gap There is a structural reason 8 kHz is under-served: on the highest-volume path, the major platforms do not give you a place to stand. As of writing, and stated as our reading of the published docs: - **LiveKit's SIP trunk integration exposes standard Krisp noise cancellation only** — not the background-voice-cancellation variant available on other paths. The telephony leg, the one with the worst audio, gets the less capable of the two bundled options. - **On Twilio there is no documented place to tap the media alongside `` — but this is our inference, not their statement.** `` is documented as terminal and the ConversationRelay socket carries JSON only, so Twilio must terminate and re-originate the media. Twilio does not say `` is incompatible, and it does not say it works either; a grep of the `` reference, the ConversationRelay overview and the onboarding pages returns no body text either way. We stated this as fact before checking, which was wrong of us. It is settled by one phone call and we have not made it. - **ai-coustics' own representative steers telephony customers away from their 8 kHz models toward the 16 kHz one.** We read that as a candid engineering answer rather than a slip: the narrowband models are not where their quality is. - **Every speaker-isolation model in the market is 16 kHz.** ai-coustics ships 8 kHz weights for their general enhancer and for their human-listening model, but every Voice Focus variant and their primary-speaker VAD is 16 kHz only. Krisp's telephony entry is a variant positioned as lighter than their main line. Telephony is exactly where competing speakers are worst — contact-centre bullpens, speakerphones, a television in the background — and it is the one place the isolation models will not run. - **The DNS Challenge — the field's main denoising benchmark series, which most published denoisers are tuned against — never had a narrowband track.** If the benchmark that defines the state of the art never scored 8 kHz, the state of the art was never optimised for 8 kHz. Our telephony rows are what that omission looks like from the receiving end. If any of those readings is wrong, or changes, we will correct this post and note the date. We have already done that once, at the top. The consequence is that the audio-quality decision on your highest-revenue channel is made for you, by a bundled component you did not choose and cannot measure. Basic noise cancellation is free everywhere by now — that is not the scarce thing. The scarce thing is a model trained for the actual channel, and evidence about whether it helps downstream. #### What Clearline is, and what it has to prove One thing it deliberately does not lead with is bandwidth extension. Predicting the missing 4 to 8 kHz band is possible, but it is generative: it does not recover the caller's `/s/`, it synthesises a plausible one. That is a real gain for a listener and a large PESQ number, and for a recognizer it is a new way to produce a confident substitution — the error moves from "unrecoverable" to "confidently wrong". The [same objective mismatch as everywhere else](/blog/does-noise-suppression-help-stt), at its widest on narrowband. Clearline is a telephony product rather than a general denoiser. Three commitments, ordered by how much of the argument each one carries: - **Native 8 kHz primary-speaker isolation.** The gap the table above exposes. The model runs at the channel's sample rate, and it keeps the person on the phone rather than the room behind them. Nobody ships this at 8 kHz today; that is the whole reason to build it. - **Residual echo suppression.** A hybrid or a speakerphone leaks the agent's own TTS back into the recogniser as fluent, confident insertions. That is not ambient noise and a denoiser does not fix it. It also does not yet have a condition in our benchmark, which we consider a bigger hole than any number in this post. - **Correct resampling.** Sample counts in equal sample counts out, verified in CI. **We are demoting this one deliberately.** We used to call it a product feature; the control earlier in this post is why we now call it hygiene. It is correct because correctness is cheap, not because it buys you accuracy. **Where the model itself actually stands, since we would rather say it than be asked.** The first Clearline proof of concept has trained and converged, and it is **weak**: **+1.41 dB SI-SDR** over the unprocessed mixture (2.90 → 4.31) across 1,200 held-out validation mixtures, where competent target-speaker-extraction systems reach **+8 to +12 dB**. Architecture experiments are running now. What that result buys us is a working training pipeline on licence-clean data and a model that demonstrably learns; what it does not buy us is a competitive model. **Clearline has no WER number, and it appears in no table on this site until it earns one.** The number arrives when Clearline is a row in Null Test, on telephony conditions that include echo, against a `passthrough` control allowed to beat it — a control that, on the evidence above, currently beats everything. If it does not clear that control, that publishes too. Anything else makes the rest of the numbers on this site worthless, which is the actual asset here. Also unfilled: how much of narrowband WER is codec artifacts versus band limit, whether bandwidth extension is net-positive for WER, and whether streaming recognizers degrade differently from batch on a narrowband channel. #### Where to go next If you run phone traffic, the first step is not to buy anything. Measure your own 8 kHz path against a passthrough control and read the error decomposition rather than the aggregate. Insertions clustered around silence usually mean packet loss concealment or echo return. Substitutions on sibilants and digits mean the band limit. Deletions of short function words mean something upstream is already suppressing too hard. The harness is open at [github.com/anecho](https://github.com/anecho) and the full matrix — including every cell where we lose — is at [/benchmark](/benchmark). Trying Clearline on your own audio takes an API key, not a sales call: [/docs](/docs). ### The split pipeline: why your VAD and your STT want different audio Canonical: https://anecho.ai/blog/the-split-pipeline · Daniel Reiss · published 2025-12-09 · 2459 words · tags: architecture, vad, turn-taking, sdk, latency Turn detection wants a clean stream. Transcription wants the raw one. We now have our own numbers for both halves of that claim: enhancement cost us five points of pooled WER and cut VAD false alarms by a third — while beating raw outright in 11 of 18 individual conditions. *Updated 14 August 2026. When this post was first published, the evidence for the split was other people's — one narrow study and a stated vendor position. It is now ours and it is measured, and the numbers below have since been regenerated across eighteen acoustic conditions rather than thirteen. Absolute values moved, the ordering did not, and everything here is quoted from the published `benchmark.json`.* Your voice agent has one microphone stream and two consumers with opposite requirements. The turn detector wants the cleanest possible signal so it does not barge in on a television. The recognizer wants the audio it was actually trained on, artifacts and all. Almost every stack we have looked at hands both of them the same buffer, and then tunes one at the expense of the other. This post is the engineering argument for splitting that stream, the measurements that now back it, and the parts that are annoying to get right: delay alignment, pre-roll, and where the split belongs in a LiveKit or Pipecat pipeline. #### The two consumers want different things **Turn-taking wants clean.** VAD and endpointing operate on energy and spectral cues over short frames. They are threshold machines. Drop the SNR and every threshold you calibrated moves: onset detection fires late because the speech-to-noise ratio needs longer to cross, offset detection fires early or never because the noise floor never drops below the hangover threshold. The highest-value failure mode in a voice agent is a barge-in false trigger — a second speaker in the room, a TV, café babble with an intelligible fragment in it — and that is exactly the case where a suppressor that isolates the primary speaker turns an unusable signal into a usable one. **Transcription wants raw.** Modern STT models are trained on very large quantities of real, noisy audio. Noise is in-distribution. Neural suppression artifacts — spectral holes, musical noise, transient smearing, a noise floor that goes unnaturally dead between words — are not. An independent study ([arXiv:2512.17562](https://arxiv.org/abs/2512.17562)) found enhanced audio scored worse than raw in all 40 configurations tested, with the caveat that it used only MetricGAN+, on medical speech, in semantic WER. AssemblyAI, citing it, separately reported Krisp's noise cancellation roughly doubling WER when its output was fed to STT, while cutting false VAD triggers about 3.5x. That was public evidence, and thin. It is no longer the only evidence we have. #### What our own run says about both halves We ran ten enhancement engines across eighteen acoustic conditions against a raw control, scoring WER and VAD in the same pass. Provenance: single recogniser, **faster-whisper base.en** (CTranslate2 int8 CPU, greedy), Silero VAD at threshold 0.5, 10 speakers / 180 clips / roughly 1406 seconds, about 190 reference words per condition, Apple M1 Max, onnxruntime 1.28.0 pinned to one intra-op thread. Full method and caveats in [the benchmark launch post](/blog/nulltest-open-benchmark). **On the transcription branch, enhancement lost the pooled average.** Pooled WER was 14.77% raw against 15.56% for the best engine (ai-coustics Quail L) and 20.26% for GTCRN. Insertions — the recogniser inventing words nobody said — went from 101 raw to 210 for GTCRN and 380 for DeepFilterNet3. None of the ten beat doing nothing on that pooled average. **Per condition it is a different sentence, and the split pipeline depends on both of them being true.** At least one engine beat raw in **11 of the 18** conditions, and the wins are large where they happen: at 0 dB broadband noise, eight of the ten engines beat raw, the best by 4.7 points. Enhancement pays where the audio is genuinely bad and costs you where it is not. The full table and its caveats are in [Does noise suppression actually help speech-to-text?](/blog/does-noise-suppression-help-stt). **On the turn-taking branch, the picture inverts — but not in the metric people usually quote.** Pooled VAD F1 barely moved: 0.950 raw against 0.947 for the best enhanced stream. If you were grading enhancement on VAD F1 you would conclude it does nothing. The false-alarm rate tells a completely different story, and false alarms are what a barge-in actually is. | Condition | VAD false-alarm rate, raw | Enhanced | Engine | |---|---|---|---| | Babble at 5 dB | 97.7% | 68.2% | FastEnhancer-L | | Competing speaker at 5 dB | 58.1% | 37.2% | ai-coustics Quail VF | Under babble at 5 dB, the raw stream false-alarms on essentially every non-speech frame the detector sees — 97.7% is a VAD that has stopped functioning as a gate. Enhancement takes that to 68.2%. That is still bad, and it is a thirty-point improvement in the one number that maps directly onto your agent interrupting a caller who has not spoken. This is the measured basis for the split. Enhancement earns its place on the turn-taking branch unconditionally, and earns its place on the transcription branch only in specific conditions — in the same run, on the same audio, with the same recogniser. Those are not two opinions to be traded off. They are two branches that want two different signals, and one of them wants a different signal depending on the room. Two honest limits on that. F1 barely moving means enhancement is trading false alarms against misses to some degree, and 190 reference words per condition separates large effects rather than small ones. And these are frame-level VAD statistics, not labelled turn boundaries: "fewer false barge-ins" as an end-to-end product metric still needs an agreed definition of a false barge-in, which we do not have yet. #### The architecture Single-stream is a compromise nobody chose deliberately: ```text ┌──────────────┐ mic ──▶ enhance ──▶│ same buffer │──▶ VAD / endpointing (happy) └──────────────┘──▶ STT (five points worse, pooled) ``` The split pipeline is the same analysis, two outputs: ```text ┌─▶ enhanced ──▶ VAD / endpointing / barge-in mic ──▶ Chamber┤ └─▶ raw (delay-matched) ──▶ ring buffer ──▶ STT ▲ └── turn boundaries index into HERE ``` One inference pass, two taps. The enhancer already computes a speech-presence estimate internally; the extra cost of emitting both streams is a copy and a delay line, not a second model. #### The part everyone gets wrong: the raw stream is early An enhancer has algorithmic latency — the analysis frame plus any lookahead. Across the models in our manifest that is 12 to 30 ms for everything except DeepFilterNet3, which we measure at 100 ms through a block-online adapter. The enhanced output therefore lags the raw input by that amount. If you emit both streams naively, the turn detector produces boundaries in enhanced-stream time and you apply them to raw-stream indices, so every utterance you slice for STT is shifted by one frame. At 30 ms that is enough to clip a word-initial plosive, and a clipped `/p/` is a substitution or a deletion in your transcript. The bug is subtle because it does not fail loudly; it just costs you a fraction of a point of WER forever. The fix is a delay line on the raw path, matched to the enhancer's declared `algorithmicLatencyMs`, so both streams share a sample clock. Our SDK does this by default: ```ts import { Anecho } from '@anecho/sdk'; const anecho = new Anecho({ apiKey: process.env.ANECHO_API_KEY }); const session = await anecho.createSession({ model: 'chamber', sampleRate: 16000, splitPipeline: true, // default alignRaw: true, // delay raw by algorithmicLatencyMs, default prerollMs: 300, // raw audio retained before onset, default }); // 20 ms of float32 mono at 16 kHz = 320 samples for await (const block of micBlocks) { const { enhanced, raw, speech, score } = session.process(block); turnDetector.push(enhanced, speech); // clean stream drives boundaries rawRing.write(raw); // raw stream is what STT will read } ``` `chamber` is our wideband enhancement model; on a phone leg you would pass `clearline` instead, which processes natively at 8 kHz. `speech` is Onset's frame-level decision computed on the enhanced signal. `score` is Nyquist's running call-quality estimate, which is the thing you alert on when a caller's line degrades mid-conversation. #### Pre-roll, or why your first word disappears A VAD declares speech after it has seen speech. Every detector has an onset delay — the frames it needed in order to be confident. If you start filling the STT buffer at the moment the VAD fires, you have already thrown away the attack of the first word. This is why the ring buffer matters more than the split does. You want a continuously-written raw ring buffer with a few hundred milliseconds of history, and turn boundaries that index into it: ```ts session.on('turn', async (turn) => { // turn.startSample / turn.endSample are in raw-stream sample indices const utterance = rawRing.slice( turn.startSample - session.prerollSamples, turn.endSample + session.hangoverSamples, ); await stt.transcribe(utterance); }); ``` Two defaults worth stating explicitly, because they are the ones people tune first: - **Pre-roll 300 ms.** Cheap insurance. At 16 kHz mono float32 that is 19.2 KB of memory per stream. - **Hangover past the offset.** Trailing fricatives and unreleased final stops are low-energy and get cut by an eager endpointer. If your transcripts systematically lose plural `/s/`, this is why. In Python the shape is the same: ```python import os from anecho_sdk import Anecho anecho = Anecho(api_key=os.environ["ANECHO_API_KEY"]) session = anecho.create_session( model="chamber", sample_rate=16000, split_pipeline=True, ) for block in mic_blocks(320): # 20 ms at 16 kHz out = session.process(block) turn_detector.push(out.enhanced, out.speech) raw_ring.write(out.raw) ``` #### Where the split lives in a real stack In an agent framework, the split has to happen before the framework's own routing, because the framework assumes one audio track. Our plugins insert at that point and hand the two streams to the two consumers. The `route` parameter is the whole idea in one field: ```python from anecho_livekit import AnechoPlugin session = AgentSession( stt=deepgram.STT(), vad=AnechoPlugin.vad(model="onset"), audio=AnechoPlugin.enhance( model="chamber", route="vad-only", # STT receives the unprocessed, delay-matched stream ), ) ``` `route` takes `vad-only` (default), `both`, or `stt-only`. We ship `vad-only` as the default because that is where our own measurements point, on our corpus, with one recogniser — not because we have proven it for your traffic. `anecho-pipecat` exposes the same three values as a frame processor. The routing matrix, stated plainly: | route | Turn-taking receives | STT receives | When to use it | |---|---|---|---| | `vad-only` | enhanced | raw, delay-matched | default; noisy environments, barge-in problems | | `both` | enhanced | enhanced | competing-speaker traffic with a Voice Focus model, where our data says STT also wins | | `stt-only` | raw | enhanced | rare; only if your VAD is already noise-robust and your STT is not | | disabled | raw | raw | control, and what you should measure against | The `both` row is not hypothetical, and it is not rare either. Enhancement beat raw on transcription in **11 of our 18 conditions** — most cleanly where a competing voice or low-SNR broadband noise is genuinely destroying phonetic cues. The largest single result is ai-coustics Quail VF on competing speaker at 5 dB, cutting insertions from 19 to 10 and WER from 26.3% to 20.5%. If your callers sit in rooms with another human talking, `both` is defensible on evidence. Pooled across all eighteen conditions the same model is more than three points worse than raw, because it also runs on clean, reverberant and carrier-degraded audio where it can only remove information — which is why it is not the default, and why this is a per-deployment measurement rather than a setting we can pick for you. #### Cost, and what the split does not cost The objection we hear is that this doubles the audio you are moving. It does not, if you keep the raw path local. Enhancement runs where the audio already is — in the browser worklet, in your media server, or in your Node process — and only the stream a given consumer needs crosses a network boundary. The enhanced stream feeds a VAD that is usually in-process. The raw stream goes to your STT vendor, which it was going to do anyway. What it does cost: - **Memory.** One ring buffer per session. Roughly 19 KB per 300 ms at 16 kHz mono float32; call it 64 KB per session with headroom. - **A delay line.** One frame of enhancer latency added to the raw path so the clocks match. This does not add end-to-end response latency, because the raw path was ahead, not behind — you are aligning to the slower stream you already had. - **One more thing to reason about.** Two streams means two places a bug can hide. The mitigation is that both are derived from one `process()` call with one sample clock, rather than two independent pipelines that can drift. #### What is still unknown - **Some recognizers may be trained on enhanced audio.** If a vendor pre-processes in their own ingest, feeding them raw is right for a different reason, or wrong for a subtle one. None of us can see those training sets. - **The SNR crossover.** Enhancement clearly wins at low SNR and clearly loses on clean audio, but we cannot yet put a number on where it turns over. The density of winners peaks at 0 dB — eight of ten engines — and thins in both directions, with one engine still beating raw even on the `clean` condition. That is a slope, not a threshold, and it will differ per engine. Locating it properly is what would make conditional enhancement shippable. - **Turn-taking gains are frame-level, not turn-level.** We report VAD F1 and false-alarm rate. "Fewer false barge-ins" needs labelled turn boundaries and an agreed metric. That is the next thing we owe the benchmark. - **Telephony is its own case.** On an 8 kHz narrowband path the recogniser is out of distribution before you touch anything, no enhancer in our matrix produced a meaningful improvement there and several degraded it badly, and the reasoning changes; see [8 kHz is where voice AI actually breaks](/blog/8khz-is-where-voice-ai-breaks). We have since run a separate experiment at 8 kHz end to end, with three genuinely native narrowband models in it: native beats resample-and-hope by 0.7 points for the same vendor, and nothing at all beats the untouched caller audio. The reason to encode the split in the SDK rather than in a blog post is that it makes the question testable per deployment. Flip `route` between `vad-only`, `both`, and disabled, run your own audio through the harness, and read the ΔWER against the passthrough control. If `both` wins on your traffic, use `both` — the SDK does not care which answer you get, and neither do we as long as the measurement is real. Quickstart and the full `createSession` reference are at [/docs](/docs). The harness that produces the ΔWER numbers is at [github.com/anecho](https://github.com/anecho), and the published matrix is at [/benchmark](/benchmark). ### Does noise suppression actually help speech-to-text? We measured it. Canonical: https://anecho.ai/blog/does-noise-suppression-help-stt · Anecho Engineering · published 2025-11-18 · 2432 words · tags: speech-to-text, noise-suppression, benchmarks, wer Krisp markets a 46% WER reduction. We ran ten enhancement engines across eighteen conditions against a raw control. None beat doing nothing on pooled word error rate, though at least one engine did beat raw in 11 of the 18 individual conditions — and insertions went up rather than down. *Updated 14 August 2026. The original version of this post argued that the field was under-measured and described the harness we were building. The harness has since run twice: the illustrative example that used to sit in the middle was replaced by real numbers, and those numbers have now been regenerated across eighteen conditions rather than thirteen. Absolute values moved, the ordering did not, and every figure here is quoted from the published `benchmark.json`.* Two vendors market speech enhancement as a way to cut speech-to-text error rates roughly in half. A major STT provider tells its customers to switch it off. If you are shipping a voice agent, one of those camps is costing you accuracy or money, and when we first wrote this, nobody had published the evidence that settles it. So we built the harness and ran it. The short answer, on our data: **no enhancement engine we tested beat raw audio on pooled WER**, the errors it added were mostly the recogniser inventing words that were never said, and **there is a substantial band of conditions where the vendor claim holds up cleanly — 11 of our 18 had an engine that beat raw.** All three of those deserve more than a headline, and the third is the one that gets dropped when this post is quoted. #### The claims that cannot all be true | Source | Claim as published | What we read it as measuring | |---|---|---| | Krisp | roughly 46% WER reduction from its noise cancellation | vendor-run evaluation; conditions, corpus and STT engine not independently reproducible | | ai-coustics | roughly 43% WER reduction | same shape of claim, different model family and corpus | | Chondhekar et al., [arXiv:2512.17562](https://arxiv.org/abs/2512.17562) | *"Original noisy audio achieves lower semWER than enhanced audio in all 40 tested configurations"*, with degradations of 1.1%–46.6% absolute | independent academic work — but note the scope: **MetricGAN+ only**, a non-streaming 2021 research model, on **medical speech**, scored in **semantic** WER. Not a production real-time enhancer. | | AssemblyAI | cited the study above, and separately reported that Krisp's noise cancellation cut false VAD triggers ~3.5x while roughly doubling WER when its output was fed to STT | an STT vendor evaluating enhancement as a preprocessing step in front of its own models | | Deepgram | published the position that enhancement hurts STT accuracy | a stated position; we have not found accompanying data | Read those rows carefully and the disagreement is smaller than it looks. Nobody is running the same corpus, the same acoustic conditions, the same enhancement configuration, or the same recognizer. There is no shared control. Every number in that table is defensible inside its own experiment and useless for comparing across them. That is not a scandal. It is what a field looks like before it has a benchmark. #### Five hidden variables that flip the sign Each of these is enough on its own to turn a 40% improvement into a 40% regression. **1. The enhancement objective is not the STT objective.** Suppression models are trained against signal-level or perceptual targets: SI-SDR, PESQ, STOI. Those reward a waveform close to a clean reference and pleasant to a human ear. None of them reward preserving the specific cues an acoustic model uses to discriminate phonemes. A model can improve PESQ by half a point while smearing the exact fricative energy that separates "six" from "fix". The loss function was never asked about that. **2. The recognizer is a hidden variable.** Modern STT models are trained on enormous quantities of real-world audio, much of it noisy. Noise is in-distribution for them. The artifacts of a neural suppressor — spectral holes, musical noise, transient smearing, an unnaturally silent noise floor between words — are not. A recognizer trained on messy real audio can be more robust to the original noise than to the cleanup. **3. The condition is a hidden variable.** On already-clean audio, enhancement can only remove information; the best possible outcome is a tie. The case for enhancement lives at low SNR, where noise is genuinely destroying phonetic cues. A benchmark weighted toward clean recordings will conclude enhancement is harmful; one weighted toward 0 dB café babble will conclude it is essential. Both are "measuring WER". **4. Double processing.** By the time audio reaches your agent it has often been through a handset noise suppressor, a codec, and a platform-level canceller. Adding another suppressor on top is not one enhancement pass, it is a cascade nobody characterised. This is especially true on telephony paths, covered in [8 kHz is where voice AI actually breaks](/blog/8khz-is-where-voice-ai-breaks). **5. Aggregation.** WER over a corpus is dominated by its hardest files. A single scalar averaged over conditions can be flat while hiding a large win in one condition and an equally large loss in another. #### Aggregate WER hides the mechanism WER is `(S + D + I) / N` — substitutions plus deletions plus insertions, over reference tokens. It is one number carrying three different failure mechanisms, and for a voice agent those mechanisms are not equally bad. Over-aggressive suppression has a characteristic signature. It tends to delete rather than substitute, and what it deletes first are short unstressed function words — "a", "the", "of", "to", "not" — because they carry low energy and short duration, and because gating decisions built around a speech-presence probability treat them as noise. Untreated noise has a different signature: substitutions where a phoneme is masked, and insertions where background speech or transient noise gets decoded as words. Attention-based models are particularly willing to hallucinate fluent text out of babble. So two systems can post identical WER and be in completely different trouble. A pipeline whose errors are insertions of background chatter is annoying but often survivable, because a downstream LLM discards the incoherent fragment. A pipeline whose errors are deletions of function words is dangerous, because dropping "not" from "I do not want to renew" produces a fluent, confident, inverted sentence the LLM will act on. This is why we record insertions, deletions and substitutions separately in every cell, and why a WER-only claim — from anyone, including us — is not enough information to act on. #### What we measured Provenance first, because these numbers are worth exactly what the method is worth. Single recogniser, **faster-whisper base.en** (CTranslate2 int8 CPU, greedy), Silero VAD, 10 speakers / 180 clips / roughly 1406 seconds of audio, about 190 reference words per condition and 3420 in total, 18 acoustic conditions, 12 backends, Apple M1 Max, onnxruntime 1.28.0 pinned to one intra-op thread. Pooled across all 18 conditions: | Backend | WER | Insertions | Deletions | RTF | Latency | |---|---|---|---|---|---| | Raw (no processing) | 14.77% | 101 | 60 | — | 0 ms | | ai-coustics Quail L | 15.56% | 126 | 57 | 0.124 | 30 ms | | ai-coustics Quail VF 2.2 L | 18.10% | 161 | 75 | 0.076 | 30 ms | | GTCRN (MIT) | 20.26% | 210 | 40 | 0.050 | 16 ms | | FastEnhancer-L (MIT) | 20.44% | 159 | 107 | 0.376 | 25.8 ms | | FastEnhancer-M (MIT) | 22.40% | 217 | 83 | 0.110 | 22 ms | | FastEnhancer-T (MIT) | 22.46% | 145 | 109 | 0.013 | 16 ms | | DeepFilterNet3 | 35.18% | 380 | 197 | 0.133 | 100 ms | **No enhancer beat raw on pooled WER.** The best result any engine posted was ai-coustics Quail L at 15.6% against raw's 14.8% — still worse, by eight tenths of a point, which at this sample size is a tie. Everything else was worse by five points or more, and those are not ties. **Insertions went up, not down.** Raw produced 101. GTCRN produced 210. DeepFilterNet3 produced 380. This is worth dwelling on, because it is the precise opposite of the marketing claim. An insertion is not a mangled word; it is the recogniser emitting a word that nobody said. The industry story is that enhancement removes the noise that causes hallucinated text. On our data enhancement *produced* hallucinated text — roughly twice as much of it for GTCRN, close to four times for DFN3, against an unprocessed control. Our reading is that the suppressor's residue, the musical noise and the abrupt gated silences, is more decodable-as-speech than the noise it removed. Note GTCRN's row in particular: 210 insertions but only 40 deletions, against raw's 101 and 60. It is not deleting your function words. It is adding words. Those are opposite failure modes and a single WER figure conceals which one you bought. #### Where the vendor claim does hold We would be doing exactly what we accuse the vendors of doing if we stopped at the pooled number. **In 11 of our 18 conditions, at least one engine beat raw.** The margins are real and sizeable, and they land where you would predict: where noise or a competing voice is genuinely destroying phonetic cues, rather than where the recogniser was coping fine on its own. The clearest is speaker isolation. ai-coustics Quail VF is a Voice Focus model — its job is to isolate the primary speaker and reject other voices. On competing speaker at 5 dB it cut insertions from **19 to 10** and WER from **26.3% to 20.5%**. On babble at 5 dB it cut WER from **24.7% to 20.5%** with insertions down 10 to 7. Those are the strongest results any enhancer posted in our matrix, and they vindicate a narrow, well-targeted use. If your failure mode is a second human voice in the room, a Voice Focus model is the right tool and our data says so. Low-SNR broadband noise shows the same shape: at 0 dB, eight of the ten engines beat raw, the best by 4.7 points. **The general pattern, and the one sentence to take from this post: enhancement pays where the audio is genuinely bad and costs you where it is not.** The seven conditions with no winner are the mirror image and just as legible: the two reverberant ones, where every engine made things worse, and the five telephony carrier conditions, where in four of five the best an engine achieves is an exact tie with doing nothing. What none of that vindicates is the general claim. Quail VF pooled across eighteen conditions lands at 18.1%, more than three points worse than doing nothing, because it also runs on clean, reverberant and narrowband audio where it can only remove information. A model that wins its design conditions by five or six points and loses the rest is not a preprocessing default. It is a routing decision, which is the entire argument of [The split pipeline](/blog/the-split-pipeline). #### Caveats, in full These belong attached to every number above. - **Single recogniser.** faster-whisper base.en. A different recogniser gives different absolute numbers. The harness supports Deepgram and AssemblyAI, so the *ordering* can be re-checked on another engine, and we would like someone to do that. - **About 190 reference words per condition**, 3420 pooled. Enough to separate large effects, not enough to resolve one or two points of WER: one word is worth roughly 0.5 points per condition, 0.03 pooled. Treat the raw-versus-Quail-L gap as a tie; the five-to-twenty point gaps below it are not ties. - **DeepFilterNet3 is measured through a block-online adapter**, because no public ONNX export carries recurrent state tensors, making per-frame streaming impossible with the published weights. Its 35.2% is a lower bound on quality and an upper bound on cost, not a fair reading of the model. Do not quote that number without this sentence. - **ai-coustics rows are a proprietary baseline** run through the licensed SDK. A reference point, never a shipping code path for us. - **Licensing.** ESC-50 noise is CC-BY-NC-3.0, so those conditions are evaluation-only. LibriSpeech is CC-BY-4.0, and the babble and competing-speaker conditions are built from LibriSpeech alone, so they carry no NC restriction. - **onnxruntime pinned to one intra-op thread.** The RTF column is one stream on one core, not a fan-out across the machine. #### What we still do not know - **Whether any STT vendor trains on enhanced audio.** If a recognizer has seen suppressor artifacts in training, its sensitivity to them is completely different, and none of us can see inside those training sets. - **How the enhancement models in the published vendor claims were configured.** Aggressiveness is usually tunable, and the same model at two settings can land on opposite sides of zero. - **How much of this transfers to streaming STT.** Our enhancement is driven block-by-block with state carried across blocks, but the recogniser is scored on complete utterances. Partial stability and latency-to-final are a separate measurement. - **Whether this holds outside English, or on a larger corpus.** It will not necessarily. The manifest and the harness are public so that this is a question someone can answer rather than debate. #### What we do with the result Two things follow, and neither of them is "enhancement is useless." The first is architectural. The same run measured VAD, and there the picture inverts: enhancement barely moves VAD F1 but cuts the false-alarm rate substantially in exactly the conditions that cause spurious barge-ins. That is a real, measured reason to send enhanced audio to your turn detector and raw audio to your recogniser, which is what our SDK does by default — the numbers and the reasoning are in [The split pipeline](/blog/the-split-pipeline). The second is where our product went. On the narrowband conditions, no enhancer produced a meaningful improvement and several degraded badly — FastEnhancer-T by 6.8 points and DeepFilterNet3 by 9.5 — and that is where the call volume is. That is [Clearline](/blog/8khz-is-where-voice-ai-breaks), and it is the thing we sell. Two honesty notes attached: our telephony conditions are a G.711 round trip over clean speech, with no echo, packet loss or handset acoustics, and no natively-8 kHz engine has been benchmarked yet, including ours. So this is a negative result about the existing field, not a positive result about us. The benchmark is not the product — a benchmark has no buyer and no retention. It is how we know what to build and how you check whether we are lying. Every cell, including the ones we lose, is at [/benchmark](/benchmark). The harness, dataset manifest and configs are at [github.com/anecho](https://github.com/anecho), and the method is described in [What we shipped: an open benchmark for speech enhancement](/blog/nulltest-open-benchmark). If you want a condition, an engine or a recogniser added, open an issue. --- ## Documentation ### Quickstart Canonical: https://anecho.ai/docs Install the SDK, get an API key without talking to anyone, and process your first stream in under five minutes. ### The split pipeline Canonical: https://anecho.ai/docs/split-pipeline Turn-taking wants clean audio. Transcription may want the original. How to route both from one capture without breaking sample alignment. ### LiveKit integration Canonical: https://anecho.ai/docs/livekit Install the LiveKit Agents plugin, keep SIP audio at 8 kHz end to end, and route clean audio to turn detection and raw audio to your STT. ### Pipecat integration Canonical: https://anecho.ai/docs/pipecat Add the Pipecat frame processor, tag audio frames by channel, and let the STT and turn-taking services consume different audio. ### API reference Canonical: https://anecho.ai/docs/api Constructor options, session options, the processed frame, model ids and error codes for the Node, browser and Python SDKs. ### Install **npm** ```bash npm install @anecho/sdk # browser build npm install @anecho/sdk-web ``` **pnpm** ```bash pnpm add @anecho/sdk pnpm add @anecho/sdk-web ``` **pip** ```bash pip install anecho-sdk # orchestrator plugins pip install anecho-livekit-plugin anecho-pipecat ``` ### Integration **LiveKit SIP** ```python from livekit.agents import AgentSession from livekit.plugins import deepgram, openai, cartesia from anecho_livekit_plugin import AnechoAudio session = AgentSession( # A SIP participant arrives at 8 kHz. Keep it there. # No resample on the way in, no resample on the way out. audio_processor=AnechoAudio( model="clearline-s", sample_rate=8000, echo_suppression=True, split_pipeline=True, # clean -> turn-taking, raw -> STT ), stt=deepgram.STT(), llm=openai.LLM(), tts=cartesia.TTS(), ) ``` LiveKit's SIP trunk exposes standard Krisp noise cancellation and nothing else, so anything different has to run inside your agent process — which is where this plugin runs. Clearline processes the narrowband signal natively instead of upsampling it first. **Pipecat** ```python from anecho_pipecat import AnechoProcessor pipeline = Pipeline([ transport.input(), AnechoProcessor(model="clearline-s", sample_rate=8000), stt, # consumes frames tagged raw context_agg.user(), llm, tts, transport.output(), ]) ``` The processor tags each audio frame with its channel; the STT service consumes the raw tag and the turn-taking service consumes the enhanced tag. **Node** ```ts import { Anecho } from "@anecho/sdk"; const anecho = new Anecho({ apiKey: process.env.ANECHO_API_KEY! }); const session = await anecho.createSession({ model: "clearline-s", // Clearline: native narrowband sampleRate: 8000, // no implicit resampling, ever splitPipeline: true, }); for await (const block of pcmStream) { const { enhanced, raw, speech, score } = session.process(block); if (speech) turnDetector.push(enhanced); stt.push(raw); metrics.record(score); } ``` enhanced and raw are the same block aligned sample-for-sample — no delay compensation on your side. **Browser** ```ts import { Anecho } from "@anecho/sdk-web"; const anecho = new Anecho({ apiKey }); const session = await anecho.createSession({ model: "chamber-xs" }); // Runs in an AudioWorklet; the model is ONNX and stays on device. const node = await session.createWorkletNode(audioContext); micSource.connect(node); node.connect(audioContext.destination); ``` Browser capture is wideband, so this is Chamber. The extra-small variant is the one small enough to ship to a tab. **Python** ```python from anecho_sdk import Anecho anecho = Anecho(api_key=os.environ["ANECHO_API_KEY"]) session = anecho.create_session( model="clearline-s", sample_rate=8000, split_pipeline=True, ) for block in blocks: # float32, shape (128,) out = session.process(block) turn_detector.push(out.enhanced) stt.push(out.raw) ``` Install with pip install anecho-sdk. --- ## Models | Model | Role | Notes | |---|---|---| | Clearline | Flagship. primary-speaker isolation at 8 kHz, enrollment-free, streaming | THE GAP: ai-coustics ships 8 kHz weights (Quail and Rook — see $rookQuail below, they are the same network at two operating points), but EVERY Voice Focus / speaker-isolation model is 16 kHz only — quail-vf-2.1/2.2 L+S and vad-vf-2.0 are all 16 kHz. Krisp's telephony answer is 'VIVA Tel Lite', a downgraded variant. Verified 2026-08-14 against docs.ai-coustics.com/reference/sdk/models. Nobody ships speaker isolation at 8 kHz — and telephony is exactly where competing speakers are worst (contact-centre bullpens, speakerphone, TV/family in background). | | Chamber | Wideband (16/48 kHz) speech enhancement + primary-speaker isolation | As in anechoic chamber. The wideband sibling to Clearline. | | Onset | Voice activity detection and turn-taking | The onset is the instant a sound begins. This is the branch our own data says enhancement actually helps. | | Nyquist | Call-audio quality scoring and failure prediction | Observability layer — the 'will this call break my pipeline' model. ai-coustics' Tyto is the only productized competitor and it is bundled, not sold. | The flagship is Clearline: primary-speaker isolation at 8 kHz, enrollment-free and streaming. None of these appear on the Null Test leaderboard yet, because they have not been run through it. When they are, they will be scored by the same code as everyone else's and published the same way, including if they lose. --- ## Pricing Canonical page: https://anecho.ai/pricing Usage-based, self-serve, published end to end. 10,000 free processed minutes every month, forever, no card. No monthly floor, no annual commitment, and no sales call for any tier — including the volume rate. ### The metered ladder (graduated, applied automatically) | Minutes per month | Rate per processed minute | |---|---| | First 10,000 | $0.00 | | 10,000 – 100,000 | $0.00100 | | 100,000 – 1,000,000 | $0.00080 | | 1,000,000 – 10,000,000 | $0.00065 | | 10,000,000+ | $0.00050 | ### Plans **Free — $0 forever.** For building the thing and finding out whether it works. - 10,000 processed minutes every month, not a trial - All four models, no feature gating - Instant key — no card, no call - Full Null Test dataset and harness - Hard concurrency cap instead of surprise overage billing **Usage — $0.0010 per minute.** For production traffic. The bill is the meter reading, nothing else. - Everything in Free, plus the free 10,000 minutes each month - Volume rates apply automatically: $0.0008 over 100k, $0.00065 over 1M, $0.0005 over 10M - No monthly minimum, no annual commitment, no sales call - Split-pipeline routing and shadow scoring included - Minimum billable invoice $5/month — smaller balances accrue **Edge — $199 per month, flat.** On-device inference costs us nothing per minute, so we do not meter it. - Unlimited on-device minutes, unlimited devices, commercial use - Browser AudioWorklet and native runtimes, audio never leaves the device - Edge Free tier at $0 for non-commercial and pre-revenue use, with attribution - Edge + Source at $1,500/month adds vendorable ONNX weights and a redistribution grant - Self-hosted container from $25,000/year — published, no call ### Add-ons | Line | Price | Rationale | |---|---|---| | Clearline — native 8 kHz + residual echo suppression | Standard ladder, no premium | It is the cheapest configuration we run and the highest-value to the buyer. Putting a surcharge on the wedge is how you fail to drive it in. | | Chamber — wideband 16/48 kHz enhancement | Standard ladder | 48 kHz is edge-only: hosted, a 48 kHz model costs more per minute than the list price. | | Onset — VAD and turn-taking | $0, bundled | Silero VAD is MIT and excellent. It is also the branch our own benchmark says enhancement actually helps, so it has to be bundled for the split pipeline to be the default. | | Nyquist — call-audio scoring | $0.0004 per scored minute | Free under 100,000 scored minutes a month. | | Null Test Private — the benchmark run on your own audio | $2,500 one-time | It converts precisely because the report sometimes says turn enhancement off. | ### Pricing questions **What counts as a processed minute?** Wall-clock audio pushed through a session, summed across streams and rounded up per session to the nearest second. Silence counts, because the model runs on it. Sessions you open and never feed cost nothing. **What happens when the free minutes run out?** On a test key, sessions keep running and return the unprocessed signal, so your agent degrades to the audio it would have had without us rather than failing. On a metered key it just bills. We never auto-bill overage on a free account — a runaway loop hits a hard concurrency cap instead of generating an invoice. **Why is there no monthly floor?** Because the nearest comparable vendor's floor is $135 a month committed annually — $1,620 before you have processed a second of audio — and that number is what stops a developer trying the product on a Saturday. We would rather have the Saturday. **Do I have to talk to someone to get the volume rate?** No. The whole ladder is published and applied automatically, including the $0.0005 rate above 10 million minutes a month. That is the number our competitors make you call for. **Do you train on my audio?** No. Metered traffic is processed and discarded; nothing is retained for training. If you need that in writing it is in the DPA, and the on-device and self-hosted options remove the question entirely. **Is the benchmark part of the paid product?** No. Null Test — the dataset, the harness and every published result — is open and free, including for people who use it to argue that a competitor beats us. We sell the private run on your own audio, never the public one. **How does on-device pricing work?** It is flat, at $199 a month for unlimited commercial on-device minutes across unlimited devices. On-device inference costs us exactly zero per minute, per-device telemetry is unenforceable, and metering something that costs us nothing is just a reason for you to say no. --- ## About Anecho Audio, Inc. — Voice AI breaks on the phone line. We fixed the audio path. Native 8 kHz enhancement, residual echo suppression and provably correct resampling for telephony voice agents — with every claim measured in the open. Website: https://anecho.ai · GitHub: https://github.com/anecho-official · Contact: hello@anecho.ai Content licence: the Null Test results are CC-BY-4.0 and the harness is Apache-2.0. Site prose may be quoted with attribution. See https://anecho.ai/ai.txt.