Four classes, one contract, two runtimes.
@anecho-official/sdk on npm (Node 18+, pure JavaScript, with the anecho CLI) and anecho on PyPI (module anecho, needs numpy and torch) expose the same Model, Processor, Vad and UsageReporter with each language's casing. The two runtimes are cross-tested against each other to −58 dB, so a result you hear in one is the result you ship in the other.
| Node | Python | Notes |
|---|---|---|
| Model.fromFile(path) | Model.from_file(path) | Loads the .anecho file from the dashboard. The file is the model — anecho.ai_focus_model_16khz_v4_1, 16 kHz, 5.3 MB, 1.3 M parameters. |
| model.audioDelay | model.audio_delay() | Algorithmic delay in samples — 15 ms at 16 kHz. Delay your raw copy by exactly this to keep it sample-aligned with the output. |
The one model is Focus anecho.ai_focus_model_16khz_v4_1: 16 kHz primary-speaker isolation, streaming, real-time on a CPU core. Your anecho.ai_focus_model_16khz_v4_1.anecho download from the dashboard is stamped per customer — treat it like a credential.
| Node | Python | Notes |
|---|---|---|
| new Processor(model, licenseKey) | Processor(model, license_key) | One per stream. Carries filter state across blocks; never share one between two callers. |
| proc.process(block) | proc.process(block) | Float32Array / float32 mono at 16 kHz, any block length. Returns the same number of samples. See the delay contract below. |
| proc.takeProcessedMs() | proc.take_processed_ms() | Drains the processed-duration counter — the number the UsageReporter sends. Call it yourself only if you meter manually. |
| proc.getContext() | proc.get_context() | The parameter surface: setParameter / set_parameter and reset(). Safe to call between blocks on a live stream. |
process(block: Float32Array): Float32Array
// in: float32 mono at 16 kHz, ANY length — buffered to the 320-sample hop internally
// out: exactly block.length samples, 15 ms behind the input
// the first model.audioDelay samples of a stream are silence — the delay, made audible| Parameter | Range | Effect |
|---|---|---|
| EnhancementLevel | 0 – 1 | How much of the model's separation is applied. 1 is full isolation; lower values mix the original back in. CLI: --level. |
| VoiceGain | dB | Output gain applied to the kept voice, after enhancement. CLI: --gain. |
| Bypass | on / off | Pass audio through untouched, same delay, so an A/B flip is click-free. CLI: --bypass. |
from anecho import ProcessorParameter
ctx = proc.get_context()
ctx.set_parameter(ProcessorParameter.EnhancementLevel, 0.8)
ctx.set_parameter(ProcessorParameter.VoiceGain, 3.0) # dB
ctx.set_parameter(ProcessorParameter.Bypass, True) # honest A/B control
ctx.reset() # back to defaultsVad(model, license_key) reports speech probability from the same model file — the signal to drive endpointing and barge-in. It exists so the turn-taking half of the split pipeline does not need a second model or a second download.
| Node | Python | Notes |
|---|---|---|
| ensureLicense() | ensure_license() | Returns a licence token for the Processor. Cached under ~/.anecho and refreshed from app.anecho.ai; works offline until the token's TTL expires — 72 hours of grace. |
| attach(proc) | attach(proc) | Points the reporter at a processor so it can collect processed durations. |
| start() | start() | Begins periodic usage reporting in the background. |
| stop() | stop() | Flushes the final report and stops. Call it on shutdown so the last minutes are counted. |
import { UsageReporter } from "@anecho-official/sdk";
const reporter = new UsageReporter(process.env.ANECHO_API_KEY!);
const proc = new Processor(model, await reporter.ensureLicense());
reporter.attach(proc);
reporter.start();| Command | What it does |
|---|---|
| anecho fetch [alias] | Downloads your model file (sha256-verified). Auth: ANECHO_API_KEY env or --key; keys live in the dashboard. |
| anecho fetch --list | Every model version your key can download, with sizes and hashes. |
| anecho process <model> in.wav out.wav | A file through the model. Flags: --level 0..1, --gain dB, --bypass. |
| anecho mic <model> [--seconds N] | Your own microphone through the model, played back A/B. |
| anecho inspect <model> | Prints what a .anecho file contains — model id, sample rate, delay. |
The anecho binary installs with @anecho-official/sdk. It is the same runtime the Node API uses — not a separate implementation — so what you hear from the CLI is what your integration will produce.
| Version | Status |
|---|---|
| v4_1 | Default — what the live demo and agent run. |
| v7 | Experimental: stronger on synthetic separation, softer on real-mic. |
| v7_2 | Candidate under evaluation — large gains on the noise and telephony benches; speaker separation is still the R&D front. |
| v8_2 | Newest candidate; native-rate siblings (8 and 24 kHz) are in the pipeline. |
Three containers are currently issued from the dashboard, one API. The SDK loads whichever file you pass to Model.fromFile / Model.from_file; the hosted API selects with model=<alias> — a form field on /enhance, ?model= on the stream socket — and defaults to v4_1.
The engine is natively 16 kHz. The SDK expects 16 kHz in and returns 16 kHz out — it does not resample; do that at your edge, statefully across blocks. The hosted API accepts and returns any rate — 8 kHz telephony, 24, 44.1, 48 kHz — through a stateful polyphase resampler in both directions, verified end to end at 8 and 24 kHz. The voice-agent path follows Gemini Live’s contract: 16 kHz uplink, 24 kHz downlink.
| Error | Meaning |
|---|---|
| ModelInvalidError | Model file corrupt, tampered with, or its integrity record stripped — the loader refuses rather than guesses. |
| LicenseFormatInvalidError | Token malformed or the signature does not verify. |
| LicenseExpiredError | Token outside its validity window — call ensureLicense() / ensure_license() and retry. |
| ProcessingNotAllowedError | Licence is for a different product, or the feature is not licensed. |
| AudioConfigUnsupportedError | Sample rate or block size does not match the model; pass variableBlockSize to buffer non-hop-multiple blocks. |
Every failure is a typed refusal at construction or configuration time — the audio path itself does not throw. A processor that constructed successfully keeps processing until you stop it, licence expiry included: renewal happens in the reporter, never mid-stream.