Part 2 of Recording, Composition & Streaming, tracing the seam between the demod composer and the recorder. In Part 1 we followed our 3 p.m. dispatch on talkgroup 101 through four bus-subscribing subsystems. Now we zoom into the one hop that carries the call’s actual audio: how the composer hands samples to the recorder without either package importing the other’s concrete types, and why the recorder — not the composer — is the single place a digital call gets decoded.
TL;DR: The composer knows the recorder only through three tiny, consumer-declared interfaces —
IQSource,PCMSink,EngineHooks. It never importsvoice.Recorder. Analog FM chains demodulate to PCM in the composer and push it viaWritePCM; digital chains emit undecoded vocoder frames viaWriteRawFrame, and the recorder decodes them. That makes the recorder the lone decoder: one vocoder pass fills the WAV and fans the same PCM to every live listener, so a call is never decoded twice.
Key takeaways
- The composer package depends on interfaces it defines itself
(
PCMSinkisWritePCM(serial, samples)and nothing more), not on the recorder. Dependency arrows point into the composer, not out of it. - Two paths cross the seam. Analog FM → real PCM via
WritePCM. Digital (P25/DMR/…) → raw IMBE/AMBE frames viaWriteRawFrame, which the recorder decodes with a per-call vocoder. - The recorder is the lone decoder for digital voice.
writeRawFramedecodes once, then fans out:.rawsidecar, live raw tap, WAV, and the decoded-PCM live tap all come from that single pass. - The daemon glues the two packages together with a
fanoutSink— one object that satisfiesPCMSinkfor the composer while multiplexing to the recorder, the host player, and the tone-out detector.
Cheat sheet
| Thing | What it is | Where in code |
|---|---|---|
IQSource |
The slice of an SDR the composer needs (StreamIQ + SampleRateHz) |
internal/voice/composer/composer.go |
PCMSink |
WritePCM(serial, []int16) — the composer’s whole view of the recorder |
internal/voice/composer/composer.go |
EngineHooks |
Touch / EndCall / UpdateSignal / UpdateDemod back to the engine |
internal/voice/composer/composer.go |
WritePCM |
Recorder entry for analog PCM; appends to the open WAV | internal/voice/recorder.go |
WriteRawFrame |
Recorder entry for a digital vocoder frame; decode-and-fan | internal/voice/recorder.go |
DecodedPCMSink / RawFrameSink |
Live taps the recorder feeds after decode | internal/voice/recorder.go |
fanoutSink |
Daemon glue: one PCMSink that multiplexes to many |
cmd/gophertrunk/daemon.go |
In this post
- Consumer-owned interfaces — why
PCMSinklives in the composer, not the recorder. - The two crossings — the analog PCM path versus the digital raw-frame path.
- The lone decoder — how
writeRawFramedecodes once and fans out four ways. - The daemon glue —
fanoutSink,playerSink, and how the wiring is assembled. - Where the shape of a call (overs, hangtime, segments) comes from — the teaser for Part 3.
Who imports whom
The interesting fact about the composer package is what it does not import.
It builds per-call demod chains that end by handing audio to the recorder, yet it
has no compile-time knowledge of voice.Recorder at all. Instead it declares the
narrow slice of behaviour it needs and accepts anything that satisfies it. This
is the classic Go move — interfaces belong to the consumer — and the composer
uses it three times.
// internal/voice/composer/composer.go (shape)
// The subset of an SDR device the composer consumes.
type IQSource interface {
StreamIQ(ctx context.Context) (<-chan []complex64, error)
SampleRateHz() uint32
}
// The subset of voice.Recorder we touch. WritePCM matches it exactly.
type PCMSink interface {
WritePCM(deviceSerial string, samples []int16) error
}
// The engine calls the chain uses to keep the engine in sync.
type EngineHooks interface {
Touch(deviceSerial string)
EndCall(deviceSerial string, reason trunking.EndReason) bool
UpdateSignal(deviceSerial string, dbfs float64)
UpdateDemod(deviceSerial string, evmPct, snrDB float64)
}
PCMSink is the whole of the composer’s view of the recorder: a single method
that takes a device serial and a slice of int16. The doc comment is blunt about
it — “Recorder.WritePCM matches this signature exactly” — but the point is that
the composer doesn’t know that. It knows PCMSink. The recorder could be
swapped for a test double, a null sink, or a fanoutSink (below) and the composer
would neither notice nor need recompiling.
The Options struct the daemon fills in is entirely interface-typed on these
seams:
// internal/voice/composer/composer.go (shape)
type Options struct {
Bus *events.Bus
Devices Devices // resolves an IQSource by serial
Sink PCMSink // typically *voice.Recorder, via a fanoutSink
Engine EngineHooks // typically *trunking.Engine
// …IQSampleRate, PCMSampleRate, VoiceHangtime, SplitPerTransmission, DSP knobs
}
Devices is a second consumer interface — FindBySerial(serial) IQSource — so
even the SDR pool reaches the composer through an abstraction. The concrete
sdr.Pool, the concrete trunking.Engine, and the concrete voice.Recorder all
live in other packages; the composer holds only their shadows. That is why the
whole demod pipeline can be exercised in a unit test with an in-memory IQ channel
and a map-backed Devices — no SDR, no daemon.
Two ways audio crosses the seam
Once a chain is running, how audio reaches the recorder depends on whether the protocol is analog or digital — and the difference is the reason there are two recorder entry points.
Analog FM (Motorola Type II, EDACS clear, LTR, MPT-1327, conventional FM)
carries voice as plain narrowband FM. The composer’s runFMChain does the whole
job: low-pass the IQ, decimate, quadrature-demodulate, convert to int16, and
call WritePCM. By the time audio crosses the seam it is finished PCM — the
recorder just appends it to the WAV. There is nothing left to decode.
Digital protocols (P25 Phase 1/2, DMR, TETRA, …) are the opposite. Their
chains do the DSP — carrier recovery, symbol timing, FEC — and produce not audio
but vocoder frames: 88-bit IMBE codewords for P25 Phase 1, AMBE+2 for DMR and
Phase 2, and so on. Those chains never call WritePCM. They call WriteRawFrame
with the raw codec bytes, and the recorder runs the vocoder. (How each chain is
selected and clocked — C4FM vs CQPSK, the FIR decimation, the boundary tracker —
is the subject of Voice Coding Part 9;
here we care only about which recorder method the chain calls and what the
recorder does with it.)
The recorder exposes a small family of raw-frame entry points, differing only in how much side information the chain can supply:
// internal/voice/recorder.go (shape)
func (r *Recorder) WriteRawFrame(serial string, frame []byte) error
func (r *Recorder) WriteRawFrameWithErrors(serial string, frame []byte, correctedBits int) error
func (r *Recorder) WriteRawFrameForCall(serial string, callID uint64, frame []byte, correctedBits int) error
All three funnel into one unexported writeRawFrame. WriteRawFrameWithErrors
adds the channel-FEC corrected-bit count, which the pure-Go IMBE decoder folds
into its adaptive smoothing. WriteRawFrameForCall adds the Grant.CallID, which
lets the recorder fence a stale frame from a call that previously held a
reused voice-tap serial — the cross-call audio-bleed guard we’ll return to in
Part 7. A P25 Phase 1 chain that knows all of this calls the richest variant; a
simple test harness calls plain WriteRawFrame.
The lone decoder
Here is the design decision the whole series leans on: a digital call is decoded
in exactly one place. Not once for the recording and again for live audio — once,
in the recorder, with the result fanned to every consumer. The writeRawFrame hot
path is where that fan-out happens.
// internal/voice/recorder.go (shape)
func (r *Recorder) writeRawFrame(serial string, callID uint64, frame []byte,
correctedBits int, haveErrs bool) error {
s := r.sessionForWrite(serial, callID) // nil ⇒ no session (or CallID fenced)
if s == nil {
return nil
}
if s.raw != nil {
s.raw.Write(frame) // (1) verbatim to the .raw sidecar
}
if r.rawTap != nil { // (2) verbatim to the live raw tap (gRPC include_raw)
// prefers RawFrameCallSink (CallID-fenced) over plain WriteRawFrame
}
if s.vocoder != nil {
if haveErrs { /* hand correctedBits to an ErrorAware vocoder */ }
samples, err := s.vocoder.Decode(frame) // the one decode
if err != nil { return nil } // log + drop from PCM, keep sidecar
if s.wav != nil {
s.wav.WriteSamples(samples) // (3) into the WAV
}
if r.decodedTap != nil { // (4) to the live decoded-PCM tap
// prefers DecodedPCMCallSink (CallID-fenced) over plain WritePCM
}
}
return nil
}
One frame in, four things out. The raw bytes go to disk and to the raw live
tap before any decode, because they exist even for protocols with no in-process
vocoder (ProVoice, or an encrypted call) — the raw sidecar is often the only
capture of such a call. Then, if a vocoder exists for the protocol, the single
Decode call produces PCM that is written to the WAV and handed to the decoded
live tap. Recording and live listening therefore share one vocoder pass; a
decode error drops the frame from PCM but keeps the sidecar intact.
The two live taps are themselves consumer-owned interfaces, declared in the
recorder for the same reason the composer declares PCMSink — to avoid an import
cycle back to the packages that consume the audio:
// internal/voice/recorder.go (shape)
type DecodedPCMSink interface {
WritePCM(deviceSerial string, samples []int16) error
}
type RawFrameSink interface {
WriteRawFrame(deviceSerial, vocoder string, frame []byte) error
}
func (r *Recorder) SetDecodedPCMSink(s DecodedPCMSink) { r.decodedTap = s }
func (r *Recorder) SetRawFrameSink(s RawFrameSink) { r.rawTap = s }
Both are wired once, at daemon construction, before Run starts — so the hot
path reads r.decodedTap and r.rawTap without locking. Each has a call-aware
sibling (DecodedPCMCallSink.WritePCMForCall, RawFrameCallSink.WriteRawFrameForCall)
that carries the CallID; writeRawFrame prefers the call-aware form via a type
assertion so the live stream is fenced by the same identity the WAV uses.
Why does this matter enough to build a whole discipline around? Because the
composer’s digital chains emit only raw frames — they never produce PCM. If
the recorder weren’t the one to decode them, the live-audio path would be silent
for every digital call while the recordings played back fine. That was a real
bug (issue #598): decoded audio reached the WAV but never reached the
WritePCM-only live sinks, because nothing decoded on their behalf. Making the
recorder the lone decoder and fanning its output is what fixed it.
The daemon glue
The composer wants one PCMSink. The daemon has several things that want the
audio — the recorder, the host speaker player, the tone-out detector. It reconciles
the two with a tiny adapter, fanoutSink, which is a slice of PCMSink and
satisfies PCMSink itself:
// cmd/gophertrunk/daemon.go (shape)
type fanoutSink []composer.PCMSink
func (f fanoutSink) WritePCM(serial string, samples []int16) error {
for _, s := range f {
_ = s.WritePCM(serial, samples)
}
return nil
}
That is the whole trick behind “one composer, many consumers.” The composer holds
a fanoutSink as its Sink and calls WritePCM once; the fan-out loops. But
fanoutSink also implements the richer shapes — WritePCMForCall (CallID-fenced
for the reused voice taps), and crucially the raw-frame methods WriteRawFrame,
WriteRawFrameWithErrors, and WriteRawFrameForCall. Each uses a type assertion
to forward only to the contained sinks that understand that shape: raw frames
reach the recorder (which understands them) and are silently skipped for the
player and tone-out (which don’t).
⚠ That raw-frame forwarding is not optional politeness. Before it existed (issue #356), the digital chains’
c.sink.(rawFrameSink)assertion failed against afanoutSink, so every IMBE/AMBE frame was dropped before reaching disk — while the activity counter still ticked, producing healthy-looking call logs next to 0-byte.rawfiles and 44-byte header-only WAVs. The daemon glue must implement every shape the chains might assert, or the fence and the frames both go silently inert.
The other adapter, playerSink, wraps the host audio player and adds a
per-talkgroup mute check before writing to the speakers — a call for a muted
talkgroup is still recorded and streamed, just not played aloud:
// cmd/gophertrunk/daemon.go (shape)
type playerSink struct {
p *player.Player
engine *trunking.Engine
}
func (s playerSink) WritePCM(serial string, samples []int16) error {
if s.engine != nil {
if tg := s.engine.TalkgroupForDevice(serial); tg != nil && tg.Mute {
return nil // muted for the speakers only
}
}
return s.p.WritePCM(serial, samples)
}
So for our 3 p.m. dispatch: if it’s a P25 call, the composer’s Phase 1 chain
demodulates the IQ into IMBE frames and calls WriteRawFrameForCall. The daemon’s
fanoutSink forwards each frame to the recorder, which decodes it once, writes
PCM to talkgroup 101’s WAV, and simultaneously feeds the browser stream and the
host player. If it were an analog Motorola call instead, the composer would have
FM-demodulated to PCM itself and called WritePCM — same fan-out, no vocoder.
Where this goes next
Part 3
answers a question this post skipped: audio crosses the seam frame by frame, but
what decides where one call — one continuous recording — begins and ends? That’s
the “composition” in the series title: how a burst of overs separated by silence
becomes one file (or many), how hangtime and talkgroup gating draw the
boundaries, and how a transmission boundary becomes the KindCallSegment event
that rolls the recorder to a new file.
FAQ
Why does the composer define PCMSink instead of importing voice.Recorder?
Because the interface belongs to the consumer. The composer needs exactly one
method — WritePCM(serial, samples) — so it declares that and accepts anything
satisfying it. This keeps the composer free of a hard dependency on the recorder,
avoids an import cycle, and lets tests drive the demod pipeline with an in-memory
sink instead of a real recorder.
What’s the difference between WritePCM and WriteRawFrame?
WritePCM takes finished PCM and appends it to the WAV — analog FM chains produce
audio directly and use it. WriteRawFrame takes an undecoded vocoder frame
(IMBE/AMBE) from a digital chain; the recorder runs the vocoder to turn it into
PCM. Analog calls never touch WriteRawFrame; digital calls never call WritePCM
from the composer.
Why is the recorder the “lone decoder” rather than decoding in the composer?
So a digital call is decoded exactly once. writeRawFrame runs the vocoder a
single time and fans the resulting PCM to both the WAV and every live listener. If
the composer decoded instead, the live WritePCM-only sinks would never see
digital audio (they don’t handle raw frames) — the exact silence bug #598 fixed by
centralising decode in the recorder.
What does fanoutSink do that a plain slice couldn’t?
It implements every sink shape the composer’s chains might assert for —
WritePCM, WritePCMForCall, and the three raw-frame methods — and forwards each
call only to the contained sinks that understand it. A missing method there
doesn’t fail loudly; the chain’s type assertion just falls back and silently drops
frames, which is how issue #356 produced empty recordings.
Series navigation
Part 2 of 14 · ← Part 1: The Output Half · Next → Part 3: Assembling a Call