Skip to content
anecho.ai
Guide · architecture

Your VAD and your transcriber want different audio.

Most voice stacks apply one filter to one track and hand the result to everything downstream. That is a routing decision disguised as a default, and it is usually the wrong one. Here is the alternative, and how to wire it.

Anecho in a voice-agent pipeline, with a split outputInbound audio enters Anecho. Two taps leave it. The enhanced tap feeds voice-activity detection for turn-taking and barge-in. The raw tap, delayed to stay sample-aligned, feeds your speech-to-text engine, then your language model, then text-to-speech, and back to the caller.INBOUNDSIP · WebRTCAnechoFocus · 16 kHz15 ms · two tapsCH2 ENHANCEDCH1 RAWVADendpoint · barge-inYour STTunprocessed by defaultYour LLMYour TTSto callerinterrupt / endpoint
One capture, two taps. Turn-taking gets the clean tap because a suppressor removes exactly the background speech that false-triggers barge-in. Transcription gets the unprocessed tap because the same suppressor removes acoustic detail the recogniser was trained on. Either route is one line in your own loop to override — but this is the default, and defaults are the argument.

This is not a hunch. Every engine below raised word error rate on our test set and several of them cut the voice-activity false-alarm rate by twenty to thirty points in the conditions that break barge-in. Same audio, two consumers, opposite verdicts.

Split pipeline · measuredSame audio, two consumers, opposite verdicts
Effect of enhancement on transcription versus on voice activity detection, per engine.
EngineTranscriptionTurn-taking (Onset)
WER raw → enhΔWERIns raw → enhF1 raw → enhFA Babble 5 dBFA Competing 5 dB
ai-coustics Quail L14.3→15.2+0.86103→1280.948→0.94398→9658→58
anecho.ai_focus_model_16khz_v8_214.3→16.6+2.27103→1190.948→0.94198→9258→53
ai-coustics Quail VF 2.2 L14.3→17.6+3.27103→1630.948→0.94598→7558→37
Resample-only 16→8→16 (control)14.8→18.6+3.86101→1730.950→0.94898→9758→55
GTCRN14.3→20.1+5.82103→2230.948→0.94598→9258→57
FastEnhancer-S14.3→20.3+6.01103→1730.948→0.94398→8558→56
FastEnhancer-L14.3→20.4+6.04103→1710.948→0.94198→6858→39
FastEnhancer-B14.3→20.6+6.32103→1630.948→0.94398→8758→56
FastEnhancer-T14.3→21.7+7.34103→1460.948→0.94098→8958→57
FastEnhancer-M14.3→22.1+7.81103→2340.948→0.94098→7558→48
DeepFilterNet3 (96-frame warm-up)14.3→30.7+16.37103→3750.948→0.91498→9358→57
DeepFilterNet314.3→34.0+19.72103→3810.948→0.90198→9258→57

Every engine we tested raised word error rate. Several of them cut the VAD false-alarm rate by twenty to thirty points in the conditions that actually break turn-taking. That is not a contradiction — it is two different consumers with two different requirements, and it is why the SDK emits two channels instead of picking a winner on your behalf.

Raw → your transcriberEnhanced → Onset

Who wants what
Default channel routing per consumer
ConsumerDefault channelReasoning
VAD and endpointingEnhancedEndpointers trigger on energy and spectral change. Background speech and broadband noise produce false starts, clipped turns and barge-in that fires at the wrong person.
Barge-in / interrupt detectionEnhancedThe failure you are avoiding is a second talker in the room stopping the agent mid-sentence. Suppressing that talker is exactly the job.
Speech-to-textRaw, by defaultRecognisers are trained on noisy audio and have their own robustness. A suppressor removes acoustic detail the acoustic model was relying on, and the errors show up as deletions of short unstressed words.
Recording / complianceRawProcessed audio is not evidence of what was said in the room. Archive the capture, not the interpretation.
Human listening / QA reviewEnhancedThe one case where sounding better is the requirement.
Wiring it

The split is two taps on one capture, in your own process. The enhanced tap is the processor’s output. The raw tap is your own copy of the input, pushed through a delay line exactly model.audioDelay samples deep so both taps share one sample clock.

node · @anecho-official/sdk
import { Model, Processor, UsageReporter } from "@anecho-official/sdk";

const reporter = new UsageReporter(process.env.ANECHO_API_KEY!);
const model = Model.fromFile("anecho.ai_focus_model_16khz_v4_1.anecho");
const proc = new Processor(model, await reporter.ensureLicense());

// Your own delay line, model.audioDelay samples deep — a ring buffer,
// a dozen lines you write once.
const rawTap = delayLine(model.audioDelay);

for await (const block of mic) {      // float32 mono at 16 kHz
  const enhanced = proc.process(block);
  const raw = rawTap.push(block);     // sample-aligned with `enhanced`

  // Clean audio drives the conversation's timing…
  turnDetector.push(enhanced);

  // …and the original drives what the words were.
  stt.push(raw);
}
When to override the default

The default is an argument, not a law. Send enhanced audio to your transcriber when your own measurements say it helps. On our test set that was almost nowhere, and on band-limited speech nothing produced a meaningful gain in either direction. But that is one recogniser on one corpus. It is a measurement, not a rule, and it is different for every stack.

per-consumer override — the routing is your code, so it is one line
stt.push(enhanced);       // you measured it; it helps on your data
recorder.push(raw);       // compliance still archives the capture

Before you flip it, run the same comparison we run: the comparator shows the shape of the effect, and Null Test shows it across the full matrix.

Cost of the split
  • One extra buffer per stream for the delay line. Memory is a few kilobytes; there is no additional algorithmic latency, because the enhanced path already paid it.
  • Two consumers to feed instead of one. The fork is the loop above — a few lines in your own process. Because the processor is a plain function from block to block, it sits in front of whatever framework routes your audio, rather than needing a plugin for each one.
  • One more thing to measure. That is the point — the routing is now a decision you can put a number on instead of a default you inherited.