Your VAD and your transcriber want different audio.
Most voice stacks apply one filter to one track and hand the result to everything downstream. That is a routing decision disguised as a default, and it is usually the wrong one. Here is the alternative, and how to wire it.
This is not a hunch. Every engine below raised word error rate on our test set and several of them cut the voice-activity false-alarm rate by twenty to thirty points in the conditions that break barge-in. Same audio, two consumers, opposite verdicts.
| Engine | Transcription | Turn-taking (Onset) | ||||
|---|---|---|---|---|---|---|
| WER raw → enh | ΔWER | Ins raw → enh | F1 raw → enh | FA Babble 5 dB | FA Competing 5 dB | |
| ai-coustics Quail L | 14.3→15.2 | +0.86 | 103→128 | 0.948→0.943 | 98→96 | 58→58 |
| anecho.ai_focus_model_16khz_v8_2 | 14.3→16.6 | +2.27 | 103→119 | 0.948→0.941 | 98→92 | 58→53 |
| ai-coustics Quail VF 2.2 L | 14.3→17.6 | +3.27 | 103→163 | 0.948→0.945 | 98→75 | 58→37 |
| Resample-only 16→8→16 (control) | 14.8→18.6 | +3.86 | 101→173 | 0.950→0.948 | 98→97 | 58→55 |
| GTCRN | 14.3→20.1 | +5.82 | 103→223 | 0.948→0.945 | 98→92 | 58→57 |
| FastEnhancer-S | 14.3→20.3 | +6.01 | 103→173 | 0.948→0.943 | 98→85 | 58→56 |
| FastEnhancer-L | 14.3→20.4 | +6.04 | 103→171 | 0.948→0.941 | 98→68 | 58→39 |
| FastEnhancer-B | 14.3→20.6 | +6.32 | 103→163 | 0.948→0.943 | 98→87 | 58→56 |
| FastEnhancer-T | 14.3→21.7 | +7.34 | 103→146 | 0.948→0.940 | 98→89 | 58→57 |
| FastEnhancer-M | 14.3→22.1 | +7.81 | 103→234 | 0.948→0.940 | 98→75 | 58→48 |
| DeepFilterNet3 (96-frame warm-up) | 14.3→30.7 | +16.37 | 103→375 | 0.948→0.914 | 98→93 | 58→57 |
| DeepFilterNet3 | 14.3→34.0 | +19.72 | 103→381 | 0.948→0.901 | 98→92 | 58→57 |
Every engine we tested raised word error rate. Several of them cut the VAD false-alarm rate by twenty to thirty points in the conditions that actually break turn-taking. That is not a contradiction — it is two different consumers with two different requirements, and it is why the SDK emits two channels instead of picking a winner on your behalf.
Raw → your transcriberEnhanced → Onset
| Consumer | Default channel | Reasoning |
|---|---|---|
| VAD and endpointing | Enhanced | Endpointers trigger on energy and spectral change. Background speech and broadband noise produce false starts, clipped turns and barge-in that fires at the wrong person. |
| Barge-in / interrupt detection | Enhanced | The failure you are avoiding is a second talker in the room stopping the agent mid-sentence. Suppressing that talker is exactly the job. |
| Speech-to-text | Raw, by default | Recognisers are trained on noisy audio and have their own robustness. A suppressor removes acoustic detail the acoustic model was relying on, and the errors show up as deletions of short unstressed words. |
| Recording / compliance | Raw | Processed audio is not evidence of what was said in the room. Archive the capture, not the interpretation. |
| Human listening / QA review | Enhanced | The one case where sounding better is the requirement. |
The split is two taps on one capture, in your own process. The enhanced tap is the processor’s output. The raw tap is your own copy of the input, pushed through a delay line exactly model.audioDelay samples deep so both taps share one sample clock.
import { Model, Processor, UsageReporter } from "@anecho-official/sdk";
const reporter = new UsageReporter(process.env.ANECHO_API_KEY!);
const model = Model.fromFile("anecho.ai_focus_model_16khz_v4_1.anecho");
const proc = new Processor(model, await reporter.ensureLicense());
// Your own delay line, model.audioDelay samples deep — a ring buffer,
// a dozen lines you write once.
const rawTap = delayLine(model.audioDelay);
for await (const block of mic) { // float32 mono at 16 kHz
const enhanced = proc.process(block);
const raw = rawTap.push(block); // sample-aligned with `enhanced`
// Clean audio drives the conversation's timing…
turnDetector.push(enhanced);
// …and the original drives what the words were.
stt.push(raw);
}The default is an argument, not a law. Send enhanced audio to your transcriber when your own measurements say it helps. On our test set that was almost nowhere, and on band-limited speech nothing produced a meaningful gain in either direction. But that is one recogniser on one corpus. It is a measurement, not a rule, and it is different for every stack.
stt.push(enhanced); // you measured it; it helps on your data
recorder.push(raw); // compliance still archives the captureBefore you flip it, run the same comparison we run: the comparator shows the shape of the effect, and Null Test shows it across the full matrix.
- One extra buffer per stream for the delay line. Memory is a few kilobytes; there is no additional algorithmic latency, because the enhanced path already paid it.
- Two consumers to feed instead of one. The fork is the loop above — a few lines in your own process. Because the processor is a plain function from block to block, it sits in front of whatever framework routes your audio, rather than needing a plugin for each one.
- One more thing to measure. That is the point — the routing is now a decision you can put a number on instead of a default you inherited.