Vocoders

Digital trunked-radio voice traffic is carried by one of three vocoders — two DVSI-derived, plus TETRA’s own ACELP:

  • IMBE — used by P25 Phase 1 LDU1/LDU2 voice frames. Core US patents (filed early-to-mid-1990s, 20-year term) have expired. The algorithm is implementable in pure Go without licence concerns; GopherTrunk ships a pure-Go decoder at internal/voice/imbe.
  • AMBE+2 — used by P25 Phase 2, DMR (Tier II / III), and NXDN. AMBE+2 is patent-encumbered. DVSI sells hardware vocoders (USB-3000 / AMBE-3003) and licences software ports. Open-source software implementations (e.g. mbelib) implement the algorithm; the code is permissively licensed (mbelib is ISC) but the patents are the user’s risk to evaluate.
  • TETRA ACELP — the full-rate speech codec used by TETRA voice (ETSI EN 300 395-2). This is not a DVSI/MBE-family vocoder; it is a conventional ACELP codec (LSP-quantised LPC + adaptive/algebraic codebooks), specified openly by ETSI. GopherTrunk ships a pure-Go decoder at internal/voice/acelp. See the TETRA ACELP section below for the clean-room / patent posture.

Re-implementing AMBE+2 in pure Go does not change the patent posture — the algorithm itself is what the patents cover, regardless of implementation language. Operators in licence-restrictive jurisdictions should evaluate with counsel before deploying. GopherTrunk’s default build ships the pure-Go AMBE+2 decoder default-on; the legal responsibility for operating it falls on the deployer, not the project.

How GopherTrunk handles this

The internal/voice package defines a Vocoder interface and a process-global Registry. Each backend registers a factory at init() time. The set of factories present in a binary is determined by the import set:

type Vocoder interface {
    Name() string
    FrameSize() int
    Decode(frame []byte) ([]int16, error)
    Reset()
    Close() error
}
Backend Build tag Default? Status
null (silence) none yes Always available
imbe (pure-Go, P25 P1) none yes Producing intelligible audio; level calibration pending reference data (voice-calibration.md)
ambe2 (pure-Go) none yes Producing audio; level calibration pending reference data; DTMF tones synthesise, knox tones via ambe2.SetKnoxTone or stay silent
tetra-acelp (pure-Go) none yes Producing audio; TETRA full-rate voice (ETSI EN 300 395-2), verified bit-exact against the reference codec; up to 4 concurrent same-carrier timeslots
dvsi (USB-3000 chip) -tags dvsi no Wire-protocol + Vocoder scaffolding shipping; USB transport stub (returns ErrNoDevice) — hardware integration follows in a separate PR

Live-pipeline auto-decode

When CallStart fires, the recorder maps Grant.Protocol to a vocoder name and instantiates a fresh vocoder per call. Each WriteRawFrame call decodes its frame and appends the resulting PCM to the call’s WAV, alongside the optional .raw sidecar.

Default mapping (see voice.DefaultVocoderForProtocol):

Grant.Protocol Vocoder Notes
p25 imbe P25 Phase 1 LDU1 / LDU2
p25-phase2 ambe2 P25 Phase 2 — AMBE+2 3600x2400
dmr-tier1 ambe2-dmr DMR Tier I direct-mode — AMBE+2 3600x2450
dmr-tier2 ambe2-dmr DMR Tier II conventional
dmr-tier3 ambe2-dmr DMR Tier III trunked
nxdn ambe2-dmr NXDN VCH — AMBE+2 3600x2450 (EHR); unverified on air
dpmr ambe2-dmr dPMR Mode 3 — AMBE+2 3600x2450; unverified on air
dstar ambe2 D-STAR DV — original AMBE 3600x2400 (the base decoder’s codebook); unverified on air
tetra tetra-acelp full-rate ACELP; up to 4 concurrent same-carrier timeslots
tetra-dmo tetra-acelp TETRA Direct Mode — same TCH/S speech frames

Analog protocols (motorola, edacs, ltr, mpt1327, etc.) have no entry — for those, the composer’s FM chain feeds WritePCM directly. EDACS ProVoice (Grant.ProVoice == true) has no in-binary decoder either; it always gets a .raw sidecar regardless of the global WriteRaw flag, so researchers can decode out-of-band.

Operators override the mapping via RecorderOptions.VocoderForProtocol:

voice.NewRecorder(voice.RecorderOptions{
    // …other fields…
    // Replace the IMBE mapping with the silence vocoder for
    // testing, leave AMBE+2 alone:
    VocoderForProtocol: map[string]string{
        "p25":        "null",
        "p25-phase2": "ambe2",
        // …other defaults…
    },
})

Pass an explicit empty (non-nil) map to disable auto-decode entirely — the .raw sidecar then becomes the only audio output for digital calls.

Raw sidecar (escape hatch)

The recorder emits a raw-frame sidecar (.raw next to the WAV) when WriteRaw is enabled or for ProVoice grants, so users can run their own decoder on the captured frames without trusting the in-binary vocoders. This is the escape hatch for operators who want bit-exact mbelib / DSD-FME / OP25 output or who prefer to defer the decoding choice to post-processing.

Decoding a captured .raw sidecar

The daemon ships a decode subcommand that runs the registered in-binary vocoders against a .raw frame stream out-of-band:

gophertrunk decode -in call.raw -out call.wav -vocoder imbe
gophertrunk decode -in dmr.raw  -out dmr.wav  -vocoder ambe2
gophertrunk decode -list-vocoders   # enumerate registered names

Stdin / stdout work for the input via -in -, so capture pipelines can stream into the decoder without a temporary file:

some-source | gophertrunk decode -vocoder imbe -out out.wav

The library function backing this — voice.DecodeStream(in, vocoderName, out) — is exported from internal/voice so other consumers (web UIs, batch processors, post-mortem analysis tools) can reuse the same decode path without spawning a binary. See internal/voice/streamdecode.go.

DSD-FME-playable sidecars (recordings.mbe_files)

With recordings.mbe_files: true the recorder additionally writes a DSD-FME-native container next to each recording, for the protocols DSD-FME can play offline: <name>.imb for P25 Phase 1 IMBE and <name>.amb for DMR / NXDN / P25 Phase 2 AMBE+2. These are DSD-FME’s own cookie-headed format, so they decode directly:

dsd-fme -f1 -w out.wav -r call.imb    # P25 Phase 1 IMBE
dsd-fme -fs -w out.wav -r call.amb    # DMR AMBE+2

TETRA ACELP and EDACS ProVoice have no DSD-FME playback mode, so they produce no MBE sidecar (the flat .raw remains the escape hatch).

dsd-neo (the DSD-FME fork with mbelib-neo) plays the same files, writes its -w WAV at a correct 8 kHz, and is the reference the calibration below was measured against. To batch-decode a directory of sidecars for an A/B against GopherTrunk’s own decode of the same frames:

for f in *.amb; do ./dsd-neo -o null -w "${f%.amb}_dsd.wav" -fs -r "$f"; done   # DMR (-fs = DMR BS/MS simplex)
for f in *.imb; do ./dsd-neo -o null -w "${f%.imb}_dsd.wav" -f1 -r "$f"; done   # P25 Phase 1
gophertrunk decode -in call.raw -out call_gt.wav -vocoder ambe2-dmr          # GopherTrunk, recording defaults

gophertrunk decode applies the daemon’s recording defaults (spec- faithful §6.2 amplitude enhancement + the calibrated unvoiced band level) so the WAV matches what the recorder wrote; pass -legacy-synthesis to hear the raw decoder instead.

Note on DSD-FME’s -w output rate. DSD-FME’s -w single.wav writer stamps an 8 kHz header on synthesis it produces at 12 kHz, so the file plays ~1.5× slow / low-pitched. This is an upstream DSD-FME quirk, not a GopherTrunk one — the .amb/.imb frame counts match GopherTrunk’s own decode exactly. Use DSD-FME’s per-call output (-P) or live audio (-o pulse) to hear it at the right rate, or re-label the -w WAV to 12 kHz.

Implementation notes

  • internal/voice/mbe/ is the shared MBE-family synthesis core: cross-frame log-amplitude prediction, voiced harmonic generator, unvoiced FFT excitation + §6.4 overlap-add window, §6.2 spectral enhancement, and a per-frame fast-attack / slow-release AGC. Consumed by both imbe and ambe2.
  • internal/voice/imbe/ holds the IMBE-specific front half: 88-bit unpack, Golay/Hamming FEC inverse, scrambler, PRBA/HOC + inverse DCTs producing the spectral residuals the shared core consumes.
  • internal/voice/ambe2/ holds the AMBE+2-specific front half: 49-bit unpack, codebook lookups (auto-generated from szechyjs/mbelib’s ambe3600x2400_const.h under ISC, regenerable via scripts/gen-ambe2-tables.sh), inverse DCTs producing the spectral residuals, and the cross-frame gamma bookkeeping AMBE+2 requires.

Both decoders share the same constructor surface (New() / NewWithSeed(seed) / NewWithConfig(seed, mbe.AGCConfig)) so operators can pin reproducibility for tests or tune the AGC for their downstream chain.

Why a plugin model

This is exactly what SDR# / OP25 / DSD do. The key benefits:

  1. The default binary has no external library dependencies for voice (no CGO, no system shared library, no install scripts).
  2. Users with DVSI hardware can opt in by building with -tags dvsi. The Vocoder + AMBE-3003 wire protocol + voice.Vocoder interface conformance ship in internal/voice/dvsi/; the USB / FTDI transport that talks to the physical chip is a stub today (returns ErrNoDevice) so the recorder fallback chain activates cleanly. Hardware integration with a real DVSI USB-3000 lands in a follow-up. CI exercises the wire protocol + Vocoder plumbing via the scripted mock Transport and the software-loopback Transport (Options{LoopbackOnly: true}); both live behind the -tags dvsi build tag.
  3. Captures contain raw frames so a researcher can defer the decoding choice to post-processing.

DVSI backend layout (-tags dvsi)

internal/voice/dvsi/:

  • packet.go — AMBE-3003 wire format (sync byte + length + type + payload). Always compiled — no patent surface in describing a serial wire protocol.
  • doc.go — exports VocoderName = "dvsi" so config validation paths can reference the key without -tags dvsi linked in.
  • dvsi_enabled.go (//go:build dvsi) — Vocoder, Transport interface, loopbackTransport, openUSBTransport stub, and the init() registration into voice.DefaultRegistry.
  • dvsi_disabled.go (//go:build !dvsi) — empty; default builds link nothing from the DVSI codepath.
  • dvsi_test.go (//go:build dvsi) — Vocoder interface conformance, loopback round-trip, scripted-mock wire-format verification, frame-size validation, unexpected-reply rejection.

make test-dvsi runs the tagged unit tests; the dvsi CI job runs the same target on Ubuntu.

Voice calibration plumbing

The calibration harness ships end-to-end:

What the 10 Sep calibration pairs showed

Four sidecars (DMR male/female, P25 male/female) decoded by both GopherTrunk and dsd-neo from the SAME .amb/.imb frames, compared harmonic by harmonic (per-frame DFT at each l·ω₀) and by long-term spectrum:

  • Pitch and voiced harmonics agree. f₀ tracks are identical (ratio 1.00), and above ~1.5 kHz the IMBE voiced harmonics match within ±1 dB — the vocoder core is right.
  • Low-frequency excess on the voiced harmonics: GopherTrunk was +2 dB at 1 kHz, +4 dB at 500 Hz, +8 dB at 350 Hz, +12 dB at 175 Hz, +16 dB at 125 Hz over dsd-neo — the shape of a first-order high-pass. dsd-neo runs a digital-voice high-pass by default (use_hpf_d = 1), and a handset’s audio path tilts the same way, which is why the un-tilted decode read as boomy/muffled. This is now the recordings.enhance.tilt_hz stage (default 450 Hz: the corner that minimises the long-term log-spectral distance over 100–3400 Hz across all four pairs once the unvoiced level is also calibrated).
  • Unvoiced (fricative / noise) bands were 6 dB (IMBE) to 13–15 dB (AMBE+2 2450) too quiet. mbelib synthesises an unvoiced harmonic as three random-phase cosines of amplitude ≈1.353·Ml each (≈2.75·Ml² of band power, 5.5× a voiced harmonic of the same amplitude); GopherTrunk’s FFT-noise §6.4 path scaled each noise bin by Ml alone, which for a 125 Hz male voice put ~12 dB less than even the equal-power level into each band and made it pitch-dependent. That is now recordings.unvoiced_gain (default 5.49 = the mbelib level; 1 = equal-power; negative = legacy). Long-term log-spectral distance to dsd-neo over 300–3400 Hz fell from 8.3 → 1.5 dB (P25 male) and 4.9 → 2.1 dB (DMR male) with the calibration alone.
  • Still open: the AMBE+2 2450 (DMR) male voice keeps a −3 to −6 dB VOICED deficit above 1 kHz that IMBE does not show (the female DMR pair goes the other way), so it is in the 3600×2450 parameter reconstruction, not the shared synthesis — the frame-by-frame diff against a wired mbelib-neo 2450 reference is the next instrument.

Reproduce: decode the pairs with both tools as above, then compare octave-band energy fractions / spectral centroid, or run internal/voice/calibrate against the _dsd.wav.

Picking a reference balance: recordings.voice_profile

The two references GopherTrunk has been measured against disagree with each other on the unvoiced level by several dB: mbelib / dsd-neo synthesise an unvoiced band at ≈5.5× a voiced harmonic (unvoiced_gain 5.49, the default), OP25 / trunk-recorder at the spec’s equal-power reading (unvoiced_gain 1). On the 10 Sep P25 A/B (one 7-reply conversation decoded by both GopherTrunk and trunk-recorder) the long-term log-spectral distance to the OP25 decode fell monotonically with the gain — 5.49 → 3.57 dB, 2 → 2.96, 1 → 2.69 — and the shipped enhance chain’s dsd-neo-matched radio tilt / high-pass sat −5 dB against OP25 at 100–300 Hz. So which decode “sounds right” depends on which reference an operator is used to, and the preset makes that one choice instead of three knobs:

recordings:
  voice_profile: op25      # "" / mbelib (default) | op25 (alias trunk-recorder)

op25 sets unvoiced_gain: 1 and, when enhance.enabled is on, disables the radio tilt and moves the rumble high-pass to 100 Hz. It only fills knobs still at their default — an explicit unvoiced_gain / enhance.tilt_hz / enhance.hpf_hz wins. gophertrunk decode -voice-profile op25 applies the same vocoder-level choice to an offline .raw decode for A/B listening.

What the preset does not change, because it is not yet measured to a fix: the residual +5..+7 dB GopherTrunk excess at 3–4 kHz relative to OP25 (it does not move with the unvoiced gain, so it is in the high-harmonic spectral-amplitude reconstruction), and the whole-LDU head loss on weak P25 calls (receiver acquisition; needs a voice IQ capture). See CLAUDE.md “P25 IMBE vs trunk-recorder/OP25”.

Knox / call-alert extension hook

AMBE+2 tone frames with b1 ∈ [144, 163] are vendor-specific knox / call-alert pairs. The public spec doesn’t document them; without registration, the decoder routes those frames through silence.

Single-entry registration

Operators with a per-vendor reference can register one (freqA, freqB) pair at a time via ambe2.SetKnoxTone (typically from a per-vendor sub-package init()):

import "github.com/MattCheramie/GopherTrunk/internal/voice/ambe2"

func init() {
    // Hypothetical Motorola Trbo call alert (frequencies illustrative).
    _ = ambe2.SetKnoxTone(150, 1100, 1750)
}

Preset bundles

For curated tables (a service-manual extract, an open-source receiver’s vendor table), use ambe2.RegisterPreset instead of repeated SetKnoxTone calls. The preset name shows up in ambe2.ListPresets() for operator diagnostics:

import "github.com/MattCheramie/GopherTrunk/internal/voice/ambe2"

func init() {
    // Hypothetical Vendor X call-alert table. Replace with values
    // sourced from a verifiable reference (DSDcc, DSD-FME, vendor
    // service manual) — the public AMBE+2 spec does not document
    // these frequencies.
    _ = ambe2.RegisterPreset(ambe2.KnoxPreset{
        Name: "vendor-x-call-alerts",
        Entries: map[int][2]float64{
            150: {1100, 1750},
            151: {1200, 1800},
        },
    })
}

Registered indices synthesise through the same summed-sinewave dual-tone path as DTMF — phase-continuous across consecutive tone frames, AGC-scaled, click-free.

Sourcing vendor frequencies

The in-tree code ships no vendor presets because the AMBE+2 public spec does not document the [144, 163] frequency range and the maintainers have not had operator-confirmed reference data to copy from. Contributors with a verifiable source (cite the file + commit or document section in the PR) can land a per-vendor preset under internal/voice/ambe2/presets/<vendor>/. Open-source receivers worth checking: szechyjs/mbelib, DSDcc, DSD-FME, OP25 — search the tone-decode paths for the b1 index range.

TETRA ACELP (internal/voice/acelp)

TETRA full-rate voice uses an ACELP speech codec specified openly in ETSI EN 300 395-2 (13.167 kbit/s, 137-bit frames / 30 ms → 240 samples at 8 kHz). GopherTrunk ships a pure-Go, clean-room decoder.

Chain. A granted call’s traffic bursts are recovered by the composer’s TETRA voice chain, TCH/S channel-decoded (EN 300 395-2 §5.5: RCPC + interleave + class-2 CRC) into 137-bit speech frames, and rendered to PCM by the ACELP decoder registered as the tetra-acelp vocoder. The decoder is the usual ACELP pipeline: parameter unpack → LSP split-VQ dequantisation → per-subframe adaptive (pitch) + algebraic (4-pulse) codebooks → log-energy gain dequantisation → LSP→LPC conversion → synthesis filter. Bit-exact fixed-point arithmetic (ITU-T-style 16/32-bit saturating operators) underlies every block.

Slot isolation. The traffic extractor emits a burst for every TDMA timeslot on the carrier and tags each with its TDMA timeslot (1..4), numbered from the synchronisation burst that anchors the 255-dibit slot grid. The composer’s TETRA chain keeps only bursts on its granted timeslot, so up to four concurrent calls on one carrier — one per slot, each on its own cc:same-carrier:N tap — decode into four independent recordings. The TCH/S class-2 CRC still gates what counts as speech (signalling bursts and encrypted TEA1-4 traffic fail it); until the slot grid anchors, or on a traffic-only carrier with no synchronisation burst, the chain falls back to CRC-gated single-call isolation.

Provenance / licensing. The decoder was ported from the algorithmic description in EN 300 395-2 and the ITU-T fixed-point operator definitions; the fixed quantiser tables (LSP codebooks, energy codebook, interpolation filters) are numeric constants sourced from the ETSI reference codec (via github.com/curlyboi/libtetradec). ACELP’s core algebraic-CELP patents (Sherbrooke/VoiceAge, filed late-1980s to mid-1990s, 20-year term) have expired, so ACELP is far less patent-encumbered than the DVSI/MBE family. As with IMBE/AMBE+2, the legal responsibility for operating the codec falls on the deployer, not the project; operators in licence-restrictive jurisdictions should evaluate with counsel. The decoder lives behind an explicit vocoder registration (tetra-acelp) so builds can omit it.

Validation. Each block carries unit tests validated independently of the reference where possible (e.g. the LSP→LPC conversion is checked against the LSP root property; log2/pow2 against math; the synthesis filter against a float recursion). Beyond that, the decoder is bit-exact to the ETSI EN 300 395-2 reference codec: feeding both the same 137-bit bitstream yields identical PCM sample-for-sample (0 / 96000 mismatches over 400 frames, including erasure/BFI frames). The TCH/S channel decode was validated the same way against the ETSI reference channel codec — GT’s type-3 frame is bit-identical, which is how the class-2 CRC error that had been dropping every on-air burst was found and fixed. The ACELP check is reproducible via the skip-guarded harness internal/voice/acelp/etsi_reference_test.go; the class-2 CRC is guarded in-tree by TestTCHClass2CRCMatchesETSI (two reference-verified vectors). The ETSI sources/vectors themselves are copyrighted and not committed.

Future work

  • Absolute-level calibration thresholds documented; reference data (operator-supplied DSD-FME / OP25 decoded WAVs) is the remaining blocker. AGC per-frame gain tweaks if real frames show systematic level offset.
  • Per-vendor knox tone tables (Motorola Trbo, Hytera, generic AMBE+2) — the extension hook ships; vendor reference data is the remaining piece.
  • DVSI USB-3000 / AMBE-3003 USB / FTDI transport implementation — the wire-protocol + Vocoder + interface conformance ship now; the actual USB bulk-IN / bulk-OUT plumbing follows when a chip is available for round-trip testing.
  • Optional Opus / FLAC re-encoding of the recorded WAVs to shrink long-running archives.
  • Plain AMBE decoder for D-STAR voice (different algorithm from AMBE+2; same DVSI patent family).
  • TETRA ACELP concurrent-call decoding runs four independent TETRA receivers (one per same-carrier tap) over the same IQ; sharing one receiver’s slot-tagged bursts across the per-slot chains would cut that CPU. Per-burst talkgroup gating is also still disabled for TETRA (the boundary tracker uses grantTG 0) because TCH/S carries no in-band talkgroup identity.