APU — Ricoh 2A03 audio unit¶
References: ref-docs/research-report.md §Technical deep-dive → APU;
ref-docs/nesdev-wiki-technical-report.md §APU; Nesdev
APU,
APU Frame Counter,
APU DMC,
DMA, and
Controller reading.
Purpose¶
Implement the 2A03 APU in crates/rustynes-apu: five sound channels (pulse 1, pulse 2, triangle, noise, DMC), the 4-step or 5-step frame counter that drives sub-channel events, the nonlinear mixer, and the analog-style high-pass / low-pass filter chain. Output is band-limited (blip_buf-style) to a configurable host sample rate (typically 44.1 or 48 kHz).
Interfaces¶
The implementation that landed in Phase 3 polled the bus differently from
the original sketch: rather than a callback-style ApuBus trait, the APU
exposes dmc_dma_pending() / dmc_dma_addr() / complete_dmc_dma(byte) that
the bus polls and services on its halt cycles. The unused ApuBus trait was
deprecated at v2.7.5 and removed at v2.9.8 (ADR 0042).
pub struct Apu { /* opaque */ }
impl Apu {
pub fn new(region: Region, sample_rate: u32) -> Self;
pub fn reset(&mut self);
pub fn tick(&mut self); // 1 CPU cycle of APU work
pub fn read_status(&mut self) -> u8; // $4015 with side effects
pub fn write_register(&mut self, addr: u16, value: u8); // $4000-$4017
// DMC DMA cooperation with the bus.
pub fn dmc_dma_pending(&self) -> bool;
pub fn dmc_dma_addr(&self) -> u16;
pub fn complete_dmc_dma(&mut self, byte: u8);
// Audio drain (host sample rate).
pub fn drain_audio(&mut self) -> Vec<f32>;
pub fn drain_audio_into(&mut self, out: &mut [f32]) -> usize;
pub fn frame_irq_pending(&self) -> bool;
pub fn dmc_irq_pending(&self) -> bool;
pub fn irq_line(&self) -> bool; // either source asserting
// Per-channel raw outputs for tests.
pub fn pulse1_out(&self) -> u8;
pub fn pulse2_out(&self) -> u8;
pub fn triangle_out(&self) -> u8;
pub fn noise_out(&self) -> u8;
pub fn dmc_out(&self) -> u8;
}
The DMC sample DMA path is intentionally a polling protocol on the
Apu, not a callback trait. When the DMC bit-shift register empties,
Apu::dmc_dma_pending() returns true and Apu::dmc_dma_addr() exposes
the target address; the SystemBus polls these on its halt cycles,
performs the read (which can stall the CPU for the documented 1-4 cycles
depending on what the CPU was doing), and feeds the byte back via
Apu::complete_dmc_dma(byte). This keeps the rustynes-apu crate from
needing any reference (trait object or otherwise) to the bus, which in
turn keeps the workspace dep graph one-directional (rustynes-apu is a leaf;
see CLAUDE.md §"Workspace dependency graph is one-directional"). An
earlier sketch of an ApuBus { fn dmc_read(...) } callback trait was
considered but never wired in production — the polling shape is simpler
and avoids the trait-object indirection on the DMA-read hot path. The trait
itself survived, unimplemented, until v2.9.8 removed it (ADR 0042).
The APU is clocked by the master scheduler at CPU cadence (every other PPU dot triple on NTSC). The triangle wave timer runs at CPU clock; pulses, noise, and DMC timer-divide at half CPU clock. The frame counter divides further to ~240 Hz.
State¶
- Per channel: 11/12-bit timer (counts down to reload), sequencer (4 step for pulse, 32 step for triangle, 1-bit LFSR for noise), length counter (5-bit, with halt flag), envelope (4-bit volume + decay), sweep (pulse only), linear counter (triangle only), DMC bit-shift register + sample buffer + memory reader.
- Frame counter: 4-step or 5-step mode, internal cycle counter (CPU clock granularity), IRQ inhibit flag, IRQ pending flag.
- Mixer state: high-pass filter state (two stages), low-pass filter state (one stage), output accumulator.
- Sample emitter: blip_buf-style ring of pending step responses + windowed-sinc kernel cache.
Save-state restore validation¶
A save state is untrusted input (v2.7.0, core audit IMP-02). Apu::restore
rejects values the hardware registers cannot hold, and resampler values that
would hang or poison the audio, with a typed ApuSnapshotError:
| Field | Legal range | Error |
|---|---|---|
pulse duty / step |
0..=3 / 0..=7 | FieldOutOfRange |
pulse sweep_period / sweep_shift / sweep_divider |
0..=7 each | FieldOutOfRange |
envelope volume_or_period / divider / decay |
0..=15 each | FieldOutOfRange |
triangle step |
0..=31 | FieldOutOfRange |
DMC rate_index / bits_remaining / dac |
0..=15 / 0..=8 / 0..=127 | FieldOutOfRange |
resampler sample_rate |
non-zero | InvalidResampler |
resampler cpu_rate |
finite, positive | InvalidResampler |
sample_rate / cpu_rate |
at most one host sample per CPU cycle | InvalidResampler |
resampler phase |
[0, 1) |
InvalidResampler |
filter coeff |
finite, [0, 1] |
InvalidResampler |
| filter stage kind | high-pass, high-pass, low-pass, in that order | InvalidResampler |
filter prev_in |
magnitude at most the stage's input bound X: 16, 33, 67 for hp1, hp2, lp |
InvalidResampler |
high-pass prev_out |
magnitude of prev_out - coeff * prev_in at most coeff * X + 1 |
InvalidResampler |
low-pass prev_out |
magnitude at most X + 1 |
InvalidResampler |
resampler held_value |
finite, magnitude at most 4 (add_sample's clamp) |
InvalidResampler |
resampler integrator, and every partial sum integrator + window[0..=i] |
finite, magnitude at most 16 | InvalidResampler |
integrator + sum(window) - held_value |
magnitude at most 1 | InvalidResampler |
resampler ring head |
any u16; masked to the ring size on install |
— |
The register-width rows prevent out-of-bounds indexing on the next tick (the
duty and triangle tables, and the mixer's 31- and 203-entry lookup tables that
decay, a constant volume_or_period and the DAC feed). The resampler rows
never panicked: a zero, non-finite or merely huge rate ratio hung
BlipBuf::add_sample's while phase >= 1.0 loop, and a high-pass coefficient
above 1 diverges to NaN.
The signal-state bounds (v2.9.9, core re-audit NC-09). Until v2.9.9 the
filter state and held_value were checked for finiteness only. A finite value
near f32::MAX (-2.05e38, say) overflowed within a frame, the audio stayed
NaN for the session, and the machine's next own snapshot was refused by the
same validator, which also broke Nes::restore's rollback, since that restores
the machine's own snapshot. The bounds are chosen to be closed under the
resampler's motion, so a state inside them can only produce states inside
them: held_value is clamped by add_sample; once every delta in flight is
integrated the integrator equals held_value, so integrator + sum(window) -
held_value is a constant of the motion (zero, up to rounding, for a state this
emulator produced); and each filter bound holds a PAIR, not a field: for a high-pass y' = c(y + x' - x)
whose input stays within X, the quantity y - c*x is closed, since
y' - c*x' = c((y - c*x) - (1 - c)x); so |y - c*x| <= c*X + 1 stays true
for any coefficient in [0, 1], and gives |y| <= 2X + 1, the next stage's
input bound. The low-pass is a convex step towards its input, so |y| <= X + 1
is closed directly. The input bounds chain from the integrator's 16: hp1 16,
hp2 33, low-pass 67. Two earlier forms were not closed. A cap of 16 on every
value let a state at the cap step past it on the next sample. Separate caps of
1024 on prev_in and prev_out let the pair (-1024, 1024) step to about
2022 (#583 review), and the 8a / (1 - a) argument behind 1024 failed for the
10 Hz Clean stage, where that figure is about 5,700. Pinned by
every_accepted_filter_state_keeps_its_own_snapshot_loadable, which checks the
snapshot after each of the first fifty samples from every accepted state on a
grid, and a_filter_stage_of_the_wrong_kind_is_refused.
The resampler's live state is saved (APU snapshot v5, v2.9.9, NL-12).
Before v5 a restore rebuilt the band-limited resampler from scratch, losing its
integrator, its warm-up flag and the 32 delta-ring slots still in flight. A
load then delivered about 17 fewer samples (717 instead of 734 in the audit's
probe), stepped the output by the lost integrator value (an audible click), and
the filter state never re-converged, so a state saved after a round trip
differed from a straight run's in 24 bytes. v5 carries the ring head, the
warm-up flag, the integrator and the 32-slot window (fixed size, so libretro's
retro_serialize_size, read once at load, still covers it), and a restore
resumes the exact stream: rollback, run-ahead and netplay stay bit-for-bit. The
undrained output queue is still not saved; it is output, not machine state,
its length varies, and every host drains it at frame end, where snapshots are
taken. A v4 APU section is refused. Pinned by
a_restore_resumes_the_exact_audio_stream and, through the C ABI, by
libretro's a_mid_run_round_trip_serializes_like_a_straight_run.
Behavior¶
Register map¶
Per ref-docs/research-report.md §APU:
| Addr | Name | Purpose |
|---|---|---|
| $4000 | PULSE1_DDLC.NNNN | Duty (DD), envelope loop / length halt (L), constant volume (C), volume / envelope period (NNNN) |
| $4001 | PULSE1_EPPP.NSSS | Sweep enable (E), period (PPP), negate (N), shift (SSS) |
| $4002 | PULSE1_LLLL.LLLL | Timer low |
| $4003 | PULSE1_lllL.LHHH | Length counter load (lllL.L), timer high (HHH) |
| $4004-$4007 | PULSE2 | Same layout as Pulse 1 |
| $4008 | TRI_CRRR.RRRR | Length counter halt / linear counter control (C), linear counter reload (RRRR.RRR) |
| $400A | TRI_LLLL.LLLL | Timer low |
| $400B | TRI_lllL.LHHH | Length counter load + timer high |
| $400C | NOISE___LC.NNNN | Length halt (L), constant volume (C), volume / envelope period (NNNN) |
| $400E | NOISE_M___.PPPP | Mode (M, 0=15-bit / 1=6-bit), period index (PPPP) |
| $400F | NOISE_lllL.L___ | Length counter load |
| $4010 | DMC_IL__.RRRR | IRQ enable (I), loop (L), rate index (RRRR) |
| $4011 | DMC_.DDDD.DDDD | Direct DAC value (7-bit) |
| $4012 | DMC_AAAA.AAAA | Sample address ($C000 + A*64) |
| $4013 | DMC_LLLL.LLLL | Sample length (L*16+1) |
| $4015 | STATUS_IF__.DNT21 | IRQ flags (read), enable bits (write) |
| $4017 | FRAME_MI__.____ | Mode (M, 0=4-step / 1=5-step), IRQ inhibit (I) |
Frame counter¶
Per ref-docs/research-report.md §Frame counter:
- 4-step (mode 0): clocks envelope+linear at every step, length+sweep at steps 2 and 4, frame IRQ at step 4 (if not inhibited). Total 14914 CPU cycles per loop NTSC.
- 5-step (mode 1): clocks envelope+linear at steps 1,2,3,5; length+sweep at 2 and 5; never sets frame IRQ. Total 18640 CPU cycles per loop NTSC.
- Writing
$4017resets the counter with a 3- or 4-CPU-cycle delay (depending on whether the write happened on an even or odd CPU cycle); if mode 1 selected, immediately clocks the half-frame and quarter-frame events. - The write's clock and the sequencer's step can be one pulse (v2.9.5). A mode-1 write whose reset matures on the CPU cycle after the sequencer fired a quarter- or half-frame step does not clock that unit again. The triggers are emitted on APU-cycle boundaries, and both land in the same one.
FrameCounter::prev_tick_stepderives "the previous tick fired a step" from the sequencer position, so no state is added. The spec is blargg'stests/roms/extra/apu/apu_test_{1..10}: 1, 2, 5 and 6 require one decrement at deltas 29830/29831 (4-step) and 37282/37283 (5-step), and 3, 4, 7 and 8 require two, one cycle later (apu_frame_clock_coincidence.rs). Only the half-frame side is observed by those ROMs; the quarter-frame side follows the same mechanism.
Nesdev's frame-counter timing is expressed in APU get/put cycle terms:
the reset side effects occur 3 CPU clocks after the $4017 write if the write
lands during an APU cycle and 4 CPU clocks otherwise. The frame IRQ line is
connected to CPU IRQ; reading $4015 returns the old frame IRQ status and then
clears the frame IRQ flag, while setting $4017 bit 6 clears it immediately.
The DMC IRQ flag is not cleared by reading $4015.
PAL has separate frame-counter step positions. Do not derive PAL frame-counter timing by scaling NTSC sample rates; use region tables.
The PAL (2A07) sequencer positions are (in CPU cycles since sequencer reset):
- 4-step (mode 0): 8313 / 16627 / 24939 / 33252 / 33253 / 33254 — quarter at 8313 / 16627 / 24939 / 33253, half at 16627 / 33253, frame IRQ at 33252 / 33253 / 33254 (if not inhibited).
- 5-step (mode 1): 8313 / 16627 / 24939 / 41565 / 41566 — quarter at 8313 / 16627 / 24939 / 41565, half at 16627 / 41565, no IRQ.
PAL frame-counter step positions are modeled (v2.1.5).
crates/rustynes-apu/src/frame_counter.rsselects the PAL positions above via theFrameCounter::palselector, whichApu::newderives from the consoleRegion(true only forRegion::Pal; NTSC and Dendy keep the NTSC positions 7457 / 14913 / 22371 / 29828-29830, and 37281-37282 for mode 1). The NTSC arms are unchanged, so the default build and every NTSC/Dendy tick is byte-identical to the pre-v2.1.5 model — AccuracyCoin APU Frame-Counter-IRQ holds 141/141 andapu_testholds 8/8. The mode-0 IRQ-flag-visibility /irq_line_activesplit is replicated verbatim at the PAL terminal steps (33252 / 33253 / 33254). The blarggpal_apu_testsoracle (see §Test plan) validates this: all 10 sub-ROMs pass, including all five PAL frame-counter-timing checks (clock jitter, mode-0/1 length timing, the two frame-IRQ timing checks) and — since the length halt/reload ordering fix below —10.len_halt_timingand11.len_reload_timing.
Length halt/reload ordering vs the half-frame clock (v2.1.5)¶
The 2A03 applies a length-counter halt change ($4000/$4004/$4008/$400C
bit) and a length reload ($4003/$4007/$400B/$400F load) one step
behind the frame sequencer's half-frame length clock:
- Halt takes effect after clocking length, not before. A halt write on the exact CPU cycle of a half-frame length clock does not suppress that cycle's clock; it governs the next one.
- A reload is ignored during a non-zero length clock. A load on the half-frame-clock cycle is honoured only when the counter was not clocked this cycle (it was already zero, so the decrement was a no-op); if it was clocked from a non-zero value the load is dropped.
crates/rustynes-apu/src/length.rs models this with the deferral fields
new_halt, reload_val and previous_count: set_halt / load latch the
written values, and LengthCounter::reload — which Apu::tick_with_external
calls on all four length channels once per CPU cycle, after the half-frame
clock and before the mixer samples the channels — promotes the halt and
applies (or drops) the reload. This mirrors TetaNES LengthCounter::reload
and Mesen2's _newHaltValue + reload-request. The change is region-agnostic
and byte-identical on NTSC: on the common write cycle with no coincident
half-frame clock the reload settles in-cycle (identical to an immediate load),
and halt does not affect output() directly — so it only alters the exact
write-on-the-clock-cycle coincidence the ROMs probe. blargg's PAL
10.len_halt_timing / 11.len_reload_timing flipped from FAILED: #3 / #4
to PASSED; NTSC AccuracyCoin (141/141), blargg_apu_2005 (11/11) and the
f2a_* length-race pins (f2_accuracy_audit.rs) are all unchanged.
How the NTSC 11/11 is read matters, and until v2.6.2 it was not read at all. The NTSC suite asserted
$6000 == 0throughrun_nes_blargg. These ROMs are plain NROM with no PRG-RAM, so$6000reads back0forever — and0is blargg's success code, making all eleven assertions vacuous. The PAL half of the corpus had the identical defect fixed in v2.1.5; the NTSC half was never migrated. Nor does the PAL fix transfer: these ROMs report a numeric result code, notPASSED/FAILED, so the screen decoder calls all elevenUnresolved.run_nes_result_codedecodes the code per the corpus's owntests.txt("a result code of 1 always indicates that all tests were passed"). The figure is unchanged; it is now earned, and demonstrated to fail — moving the 4-step half-frame clock is caught by two ROMs and the frame IRQ by seven.
Reset behavior (v2.0.0 "Timebase", promoted in beta.4)¶
Per the blargg apu_reset spec and nesdev ("At reset, $4017 mode is
unchanged, but IRQ inhibit flag is sometimes cleared"): the frame counter
retains the last value written to $4017 (FrameCounter::last_4017), and a
warm reset behaves as if that value were written AGAIN — the reset zeroes the
sequencer + IRQ flags, cancels any in-flight pre-reset $4017 write still in
its 3/4-cycle maturation window, and SCHEDULES a re-write of
last_4017 & 0x80 (mode bit retained, IRQ-inhibit bit cleared) landing 2
clocked cycles into the CPU's 8-cycle reset sequence. The re-write flows
through the normal $4017 write path (the 3/4-cycle aligned delay + the
mode-1 immediate quarter/half clock), so execution resumes ~9–12 cycles after
the effective write — blargg 4017_timing measures 8 (its accept window is
6..=12; hardware-typical is 9). $4015 is cleared at reset (channels
disabled); the channel registers — including the halt/duty bits — survive.
This closes plan-residual R4 (apu_reset/4017_written): all six blargg
apu_reset ROMs pass strictly.
Save-state coverage (APU_SNAPSHOT_VERSION v4). The scheduled re-write is
live state for the two CPU cycles between Apu::reset arming it and
tick_with_external firing it, so reset_4017_delay + reset_4017_value are
serialized. They were not before v4: a snapshot landing in that window restored
delay = 0, cancelling the re-write, and the restored frame counter kept the
sequencer phase the re-write exists to reset. The window is narrow and no
user-visible symptom was ever attributed to it — unlike the PPU's v5/v6/v8
tails, this one was found by the standing schema audit
(crates/rustynes-test-harness/tests/snapshot_schema_audit.rs) rather than by a
bug report. Pinned behaviourally by
a_reset_survives_a_snapshot_restore_taken_mid_countdown, which compares
frame_counter.cycle across a mid-countdown round trip; note that
frame_counter.mode cannot serve as the oracle, since reset_rewrite_4017
retains bit 7 and the re-write therefore restores the mode already in effect.
v1..=3 blobs upconverted to "no re-write pending", the resting value, until
v2.9.8; since then Apu::restore reads v4 only, with every field required
(ADR 0042).
DMC channel¶
- Memory reader: when sample buffer is empty and bytes-remaining > 0, request DMA. Bus halts CPU and reads 1 byte from
$C000-$FFFF. Halt cost: 3 or 4 CPU cycles perref-docs/research-report.md§DMA. Read advances address (wraps$8000after$FFFF) and decrements bytes-remaining. - Output unit: shift register bits modify the 7-bit output: bit 1 → +2, bit 0 → -2, clamped 0..=127.
- IRQ: when bytes-remaining reaches 0 and IRQ enable is set (and loop is not), assert DMC IRQ. Cleared by writing
$4015. - Direct write to
$4011sets the DAC immediately, useful for raw PCM.
DMC DMA has two scheduling classes. Load DMA follows enabling playback through
$4015 and is scheduled around the second APU cycle after the write. Reload DMA
follows the sample buffer emptying during playback and schedules on the opposite
get/put phase. Both perform a dummy cycle after halting the CPU and may need an
alignment cycle before the memory read. This distinction is observable through
CPU stalls and repeated side-effect reads.
DMC load-DMA even/odd-cycle delay (v1.7.0 F2b)¶
A DMC sample-buffer LOAD DMA that begins on a "get" (odd) CPU cycle is
deferred one extra cycle relative to one that begins on a "put" (even) cycle —
the load only takes effect on its put half. This is modelled by
Bus::dmc_dma_defer_load_entry in crates/rustynes-core/src/bus.rs, which gates
the load entry on the APU's put_cycle() parity (on the current dot-lockstep
scheduler the pre-cycle parity is read flip-invariant; see the in-source note for
why the predicate is put_cycle rather than !put_cycle). The behavior is
already implemented and verified, not new in v1.7.0; the
f2b_* tests in crates/rustynes-test-harness/tests/f2_accuracy_audit.rs pin it
end-to-end via dmc_tests/latency.nes (a deterministic DMC fetch-latency audio
signature) and the strictly-passing sprdma_and_dmc_dma alignment ROM.
A load DMA refused by a write takes four cycles (v3.1.0)¶
RDY cannot halt a write. When a pending LOAD DMA reaches the get half on which
it would enter and that cycle is a CPU write, the load is refused and enters
on the very next read whichever half that is: refused by one write it
lands on a put half and takes four cycles ([Put (halt)] [Get] [Put] [Get]);
refused by two consecutive writes it lands on a get half and takes three. The
get-half deferral above therefore does not apply a second time to a
write-refused load. The bus records the refusal in a one-shot latch
(dmc_load_write_delayed, crates/rustynes-core/src/bus.rs), set in
Bus::write, consumed by the DMC entry in unified_dma_cycle_impl and cleared
by the next CPU read; it is in the BUS save-state section (version 3), because
the refusing write is the last cycle of a store and a snapshot can fall between
it and the next opcode fetch. Written from AccuracyCoin DMA Landing on Write
test 9 (upstream f5f41dc2) and its cycle comments, and confirmed by a
black-box per-cycle comparison with TriCNES's output at the test's
STA $5000: before the fix the CPU ran the opcode fetch the hardware spends
halted, one cycle ahead from then on.
Mixer¶
Per ref-docs/research-report.md §APU Mixer, two implementations:
// Linear (first cut, fails apu_mixer test ROM)
pulse_out = 0.00752 * (pulse1 + pulse2);
tnd_out = 0.00851 * triangle + 0.00494 * noise + 0.00335 * dmc;
output = pulse_out + tnd_out;
// Lookup-table (~4% accurate, default)
pulse_table[n] = 95.52 / (8128.0 / n as f32 + 100.0); // n=0 -> 0
tnd_table[n] = 163.67 / (24329.0 / n as f32 + 100.0);
output = pulse_table[(pulse1 + pulse2) as usize]
+ tnd_table[(3 * triangle + 2 * noise + dmc) as usize];
After mixing, apply: 90 Hz first-order high-pass, 440 Hz first-order high-pass, 14 kHz first-order low-pass.
Filter model (v2.1.3). The three-stage chain above is the NES front-loader
(RF/composite) circuit and is the default (FilterModel::NesRf, byte-identical to
earlier builds — it matches ares/tetanes). Because that 440 Hz high-pass rolls off
the bass/triangle register hard (an authentic but thin sound), Apu::set_filter_model
also offers two softer, hardware-grounded models: Famicom (a single ~37 Hz
high-pass — the nesdev Famicom spec, fuller low end) and Clean (only a ~10 Hz
DC-block — fullest, the character Mesen2 / FCEUX / Nestopia produce by omitting the
high-pass cascade). The model is tonal only — channel content is identical, it is
never written into the save state, and the frontend re-applies it at ROM load — so
determinism and the audio oracle hold on the default. Frontend selector: Settings
→ Audio → Filter model ([audio] filter_model = nes / famicom / clean).
Per-channel gain. Apu::set_channel_gain scales each channel's integer
output before the lookup mixer. Each gain is clamped to 0.0..=2.0, and since
v2.9.3 a NaN gain falls back to unity (1.0) instead of reaching the mixer,
where NaN.clamp would have stayed NaN and the rounded sample index would have
been meaningless. Infinities clamp to the range ends. Pinned by
channel_gain_rejects_nan_and_clamps_infinities.
Settings across a power cycle (v2.9.8). A power cycle rebuilds the APU from
Apu::new, and Apu::adopt_settings_from carries the host's settings onto the
new one: the channel mask, the per-channel gain and the filter model. The APU
keeps the selected model as a plain value (Apu::filter_model) beside the
built chain, since a chain of coefficients does not say which model made it;
the carried model is rebuilt as a fresh chain at the APU's sample rate, with no
IIR history from the old timeline. Until v2.9.8 all three reverted to their
defaults in the cycle, and every host had to push them again. Pinned by
a_power_cycle_keeps_every_ppu_and_apu_setting and, for any future
configuration field, every_config_field_survives_a_power_cycle.
Band-limited sample emission¶
Naive sample-rate conversion produces aliasing. Use a blip-buf-style ring buffer:
- Each "step" (channel transition) is registered with the time-of-step at CPU-cycle resolution.
- The buffer convolves each step against a windowed-sinc kernel into the host-sample-rate output buffer.
drain_samples()returns finalized samples; the buffer slides forward in time.
Implementation: blip_buf-rs crate or hand-rolled equivalent (~200 LOC).
$4015 semantics¶
- Read: returns frame IRQ (bit 6), DMC IRQ (bit 7), DMC bytes-remaining > 0 (bit 4), pulse 1 / 2 / triangle / noise length-counter > 0 (bits 0-3). Reading clears the frame IRQ flag. Does not clear the DMC IRQ flag.
- Write: bit 4 set enables DMC (initiates sample if buffer empty); bit 4 clear silences DMC. Bits 0-3 enable channels (clearing forces length counter to 0).
$4015 is internal to the CPU/APU package rather than an external-bus device.
When refining open-bus behavior, do not assume $4015 reads update the same
external open-bus latch used by cartridge or PPU register accesses.
Bit 5 is open bus, from the CPU's internal data bus (internal_data_bus in
crates/rustynes-core/src/bus.rs): nesdev's APU page says the value "comes from
the last cycle that did not read $4015". Every CPU write, and every CPU read
other than $4015 itself, sets that latch (the $4015 read returns before the
latch is assigned, which is what keeps bit 5 on the older value), including a read of an address nothing decodes, where the CPU latches the
floating value it sees. A DMC DMA fetch drives only the external bus
(AccuracyCoin Internal Data Bus Test 2), and so does an OAM-DMA put while the
6502 bus is parked in $4000-$401F (oam_dma_put), so after either the two
latches differ. Until v2.9.2 an undecoded cartridge-space read ($4020-$FFFF)
returned early and skipped the internal update, while the undecoded
$4000-$401F arm always made it; after either of those that left bit 5 on an
older value (core audit v2.9.2 AUD-03). Pinned by
an_unmapped_cartridge_read_latches_the_floating_value_onto_the_internal_bus.
Through the ordinary instruction stream this is not observable, since
LDA $4015's own operand fetches overwrite the latch first; the paths that can
see it are the OAM-DMA reads with the 6502 bus parked in $4000-$401F, whose
$4015 mirror composes bit 5 in the same cycle as an undecoded source read.
AccuracyCoin stays 144/144 and nestest 0-diff with the change.
Edge cases and gotchas¶
- DMC DMA stalls CPU mid-instruction. Per
ref-docs/research-report.md§DMA, halt only on read cycles. The 2A03 register-readout bug (extra reads of$2007,$4015-$4017while halted) must be reproduced — required bydmc_dma_during_read4. - Frame counter write jitter. Writing
$4017with a value that includes IRQ inhibit set clears any pending frame IRQ flag — on the write cycle itself (both$4015bit 6 and the CPU /IRQ line); only the timer reset waits the 3-4 cycles. The wiki states the two separately. Until v2.9.8 the clear waited for the timer reset, and Nintendo World Championships 1990 (STA $4017with$40, thenCLI, with the frame flag set since cycle 29,828 and its IRQ vector in uninitialised WRAM) took the IRQ and never drew a frame. Pinned bywrite_4017_inhibit_drops_the_irq_on_the_write_cycleincrates/rustynes-apu/src/frame_counter.rs. The inhibit itself takes effect with the write as well (v2.9.8, found in review): with the flag cleared on the write but the inhibit left to the timer reset, a write 1-3 cycles before step 29,828 let the OLD sequence raise the IRQ again inside the reset delay, an interrupt the program had just inhibited. No source states the cycle directly. This reading is the wiki's, which ties the flag clear to the inhibit bit and not to the reset. No ROM in the suite reaches that window: the APU and AccuracyCoin suites pass under either model. Pinned bywrite_4017_inhibit_holds_through_the_reset_delay. Since v2.9.9 the CLEAR direction is immediate too (re-audit NC-16): v2.9.8 set the inhibit on the write but cleared it only at the timer reset, so a$4017 = $00written 1-3 cycles before step 29,828 with the inhibit set kept the old sequence's IRQ masked. The bit is now a latch the write sets or clears at once; the old sequence's IRQ is raised whenever it reaches 29,828 before the reset restarts it (the write lead shorter than the 3- or 4-cycle delay). Same evidence class: the wiki's reading, no ROM in the window. Pinned bywrite_4017_inhibit_clear_unmasks_on_the_write_cycle. - Length counter halt / reload race (v1.7.0 F2a; ordering fixed v2.1.5). The effective halt flag is consulted at the half-frame length clock; a
$400xhalt-bit write — or a length reload — on the CPU cycle of that clock races over whether the counter is clocked this step. Silicon resolves the halt change after the clock and drops a reload that lands on a non-zero clock. This is modeled by the deferral mechanism inlength.rs(new_halt/reload_val/previous_count, promoted byLengthCounter::reloadafter the half-frame clock and before the mixer sample — see §Length halt/reload ordering above). blargg10.len_halt_timing+11.len_reload_timingbracket the exact cycle and pass strictly on both the NTSC (blargg_apu_2005.07.30) and PAL (pal_apu_tests) builds. Thef2a_*tests incrates/rustynes-test-harness/tests/f2_accuracy_audit.rsare the named NTSC regression pin. - Triangle disabled silently when length counter or linear counter reaches 0. Holds the last sequencer step (does not produce a click).
- Ultrasonic silence (timer period < 2). When the triangle timer period is below 2 (frequency above ~55.9 kHz), real hardware cannot follow the sequencer and the channel effectively halts. We freeze the sequencer in
Triangle::clock_timer(the step does not advance and the output holds its current value) rather than emitting the aliasing tone, matching the common-emulator convention; Mega Man 2's "Crash Man" stage relies on this to silence the triangle. The threshold is strictly< 2(period 2 still clocks). Seecrates/rustynes-apu/src/triangle.rs. - Pulse duty-sequencer phase reset on
$4003/$4007. Writing the length/timer-high register resets the pulse duty sequencer to step 0 (and sets the envelope-restart flag) but does not reset the timer divider. Implemented inPulse::write_timer_hi(crates/rustynes-apu/src/pulse.rs). - DMC playback stops mid-scanline? Yes;
$4015write to clear bit 4 silences the channel after the current sample byte completes. - Sweep mute. A pulse channel is muted when its CURRENT period is below 8, or when the sweep's target period is above
$7FF. Both conditions are evaluated continuously, whether or not the sweep is enabled. A negated target can never exceed$7FF, so negate mode never mutes through the target: a target that would go negative clamps to zero (NESdev "APU Sweep"). Until v2.7.0 the oracle computed pulse 1'sc - (c >> 0) - 1in wrappingu16arithmetic, got$FFFF, and muted pulse 1 on the$4001 = $08idiom, the documented way to disable the sweep (negate on, shift 0). Pinned bypulse1_negate_shift0_clamps_to_zero_and_does_not_muteincrates/rustynes-apu/src/pulse.rs; the companion test sweeps every shift 1-7 and every period 8-$7FFand shows the clamp changes nothing else. - Pulse 1 sweep negation off-by-one. Pulse 1 negates by
~target(one's complement); Pulse 2 negates by-target. This produces audible difference at certain frequencies. - Controller conflict is APU-owned timing. The standard controller code lives in the input subsystem, but DMC DMA is the root of the classic joypad bit deletion/duplication bug. APU/DMA changes must rerun controller-read coverage, not only APU audio ROMs.
Test plan¶
apu_test(8 sub-ROMs) — register I/O, frame counter, length counter halt timing.pal_apu_tests(10 sub-ROMs, PAL region) — blargg's PAL-calibrated rebuild of the 2005-era APU length/frame-IRQ/timing checks. Wired in v2.1.5 as the first PAL-region APU oracle (tests/pal_apu_tests.rs). These predate the$6000protocol and report on-screen (plain NROM, no PRG-RAM), so the suite decodes the renderedPASSED/FAILED: #<n>verdict via therun_nes_screenharness runner rather than the (vacuous, for these ROMs)$6000check. Current state: 10/10 pass —01.len_ctr/02.len_table/03.irq_flag(region-independent) plus04.clock_jitter,05/06.len_timing_mode0/1,07.irq_flag_timing,08.irq_timing(the PAL frame-counter-timing checks, passing since the v2.1.5 PAL step positions), and10.len_halt_timing/11.len_reload_timing(passing since the v2.1.5 length halt/reload ordering fix documented above and indocs/accuracy-ledger.md).apu_mixer— confirms lookup-table mixer matches reference within 4%.dmc_dma_during_read4— DMC DMA stalls + register read crosstalk.- Audio capture comparison: emit 60 frames of audio for a curated set of demo ROMs, compare PSNR against a Mesen-generated reference. (Not a strict pass/fail but a regression detector.)
- Property test: random
$4017writes interleaved with channel writes; assert frame counter cycle accounting matches a hand-rolled reference.
Expansion-chip audio¶
Six on-cart expansion sound chips are synthesized and summed into the external-audio mix via the Mapper::mix_audio(&mut self) -> i32 hook (default 0; i16 until v2.2.3). Each synth core lives in the owning mapper crate, not the 2A03 APU crate, because they are cartridge hardware:
| Chip | Mapper(s) | Synth core | Clock cadence |
|---|---|---|---|
| VRC6 | 24 / 26 | Vrc6Pulse x2 + Vrc6Saw (crates/rustynes-mappers/src/m024_vrc6.rs) |
every CPU cycle ($9003 halt + freq-scale shift) |
| VRC7 | 85 | rustynes_apu::Opll (emu2413-derived, MIT) |
OPLL calc() every 36 CPU cycles (49,716 Hz) |
| FDS | 20 (FDS device) | FdsAudio wavetable + FM (crates/rustynes-mappers/src/fds.rs) |
wave/mod every 16 CPU cycles; envelopes per cycle |
| MMC5 | 5 | Mmc5Audio (2 pulse + 7-bit PCM, crates/rustynes-mappers/src/m005_mmc5.rs) |
pulse timer every other CPU cycle; envelope/length on 2A03 frame events |
| Namco 163 | 19 / 210 | Namco163Audio (1-8 time-multiplexed wavetable channels) |
round-robin channel update every 15 CPU cycles |
| Sunsoft 5B | 69 (FME-7) | Sunsoft5BAudio (3 tone + noise + envelope) |
every CPU cycle |
All synth cores are behind the default-on mapper-audio Cargo feature; when it is off (e.g. the no_std build) the register decoders still latch (save-state round-trip preserved) but clock/mix are no-op shims that return silence. The VRC7 OPLL core is deliberately the MIT emu2413 lineage — not Nuked-OPLL (GPL/LGPL, license-incompatible).
Expansion-audio levels (v2.1.6 "Expansion Audio")¶
Each chip's mix_audio() is scaled so its full-volume square sits at the relative loudness the hardware produces vs the 2A03 pulse, calibrated against the reference-emulator field (Mesen2 was RustyNES's historical accuracy bar, but VRC6 was recalibrated away from it in v2.2.7 — Mesen2 is the loud outlier for VRC6; see the v2.2.7 note below), measured by the bbbradsmith db_* decibel-comparison ROMs. The reference is Mesen2 NesSoundMixer::GetOutputVolume (2A03 pulse peak 95.88*5000/(8128/15+100) ≈ 746.9; linear expansion weights VRC6 ×5·internally-×15, MMC5 ×43, N163 ×20, 5B ×15, VRC7 ×1), cross-checked against nestopia / puNES / fceux / tetanes. v2.2.7 "Timbre II" re-corrected the VRC6 target away from that Mesen2 weighting — a NESdev-forum reviewer flagged VRC6 as too loud, and a cross-reference across the eleven reference emulators surveyed as oracles plus the NESdev wiki confirmed Mesen2's ×5 is the outlier, not the field: the wiki states that "at maximum volume, the pulse channels of the VRC6 are roughly equivalent to the pulse channels of the 2A03," and rustico / tetanes / BizHawk each encode a VRC6 pulse as exactly a 2A03 pulse (ares / higan / nestopia reach the same figure via a sum/61 normalization). MMC5 / N163 / 5B keep their Mesen2-derived targets, which the same cross-reference corroborates. The crates/rustynes-test-harness/tests/audio_expansion.rs level_db_* oracle asserts the measured expansion-vs-reference ratio from each ROM's rendered waveform:
| Chip (ROM) | Target ratio vs APU square | RustyNES scale (mix_audio) |
Status |
|---|---|---|---|
APU triangle (db_apu) |
≈ 0.524 (fixed 2A03 DAC balance) | pulse_table / tnd_table LUT |
Asserted |
VRC6 (db_vrc6a/b) |
≈ 1.000 | VRC6_MIX_SCALE = 650 (m024_vrc6.rs; was 979) |
Asserted (recalibrated v2.2.7) |
MMC5 (db_mmc5) |
≈ 1.000 ("equivalent to APU") | pulse ×650 / PCM ×40 (m005_mmc5.rs; was 256/16) |
Asserted (v2.1.6) |
Namco 163 1-ch (db_n163) |
≈ 6.02 | NAMCO163_MIX_SCALE = 261 (m019_namco163.rs; was 64) |
Asserted (v2.1.6) |
Sunsoft 5B (db_5b) |
≈ 1.265 (vol-12) / 3.554 (vol-15) | shape SUNSOFT5B_LOG_VOL + level SUNSOFT5B_MIX_SCALE_NUM/DEN = 2549/138 |
Asserted (v2.2.3) |
VRC7 (db_vrc7) |
≈ 2.7 peak (patch-dependent) | raw Opll::calc() (±4095) |
Snapshot-guarded — see below |
VRC6 (1.506), MMC5 (1.0) and N163 (6.02) were the v2.1.6 level corrections; MMC5's mix_audio bias moves to -12290 accordingly. v2.2.7 "Timbre II" superseded the VRC6 figure: the ~1.506× target mirrored Mesen2's specifically louder ×5 mixer weight, but the wider field (rustico / tetanes / BizHawk / ares / higan / nestopia) and the NESdev wiki agree VRC6's pulse is loudness-equivalent to the 2A03 pulse, so the target moved to ≈1.000 and VRC6_MIX_SCALE moved 979 → 650. The db_vrc6a/b snapshots were re-blessed for the new waveform amplitude (audio-only; framebuffer and cycle count are byte-identical). VRC6's per-channel balance — saw 0-31 linearly summed against pulse 0-15 — was already correct and is unchanged. VRC6/MMC5/N163 fixes touch only the expansion channel — the base 2A03 mix is a separate additive term (mix_audio() == 0 for non-expansion mappers), so AccuracyCoin / blargg / nestest stay byte-identical.
Sunsoft 5B absolute level — closed in v2.2.3 (A1). The log-volume DAC shape was always hardware-exact (×1.4126/step, verified by sunsoft5b_volume_dac_follows_logarithmic_step_law); the level was deferred for one reason only — Mapper::mix_audio returned i16, and a full-volume tone at the db_5b level is 1882 × 18.471 = 34,761, past i16::MAX for a single channel (three simultaneous tones ≈104 k, 3.2× over). The trait return is now i32, and the level is calibrated by SUNSOFT5B_MIX_SCALE_NUM/DEN = 2549/138 ≈ 18.471: measured 0.0685× before (~23 dB too quiet), 1.2651× after, against the Mesen2-derived target LUT[12]=63 × weight 15 / 746.9 = 1.265. Asserted by level_db_5b. Shape and level are now separately pinned, each by its own oracle. The widening is representational for every other board — they return the values they always did. v2.2.7 "Timbre II" completed the envelope-mode DAC: the 5-bit hardware envelope previously truncated to 4 bits (>> 1) before indexing the 16-entry SUNSOFT5B_LOG_VOL table — the wiki-documented 3 dB/step approximation — and now indexes a new 32-entry SUNSOFT5B_LOG_VOL32 table (×1.1885/step = 1.5 dB, matching nestopia / rustico) with the full 5-bit level, so odd envelope levels get their own DAC step instead of sharing their neighbor's. LOG_VOL32's odd entries equal LOG_VOL exactly, guarded by log_vol32_odd_entries_match_4bit. Fixed 4-bit volume (already the correct 3 dB/step for a 4-bit register) and the absolute level (1.265×) are unchanged; every committed 5B snapshot stayed byte-identical since the extant 5B test ROMs don't exercise odd envelope levels.
One level remains an honest documented gap (docs/accuracy-ledger.md §Expansion-audio levels):
- VRC7 FM level. The OPLL FM synthesizer is implemented (emu2413 port) and its instrument ROM is verified canonical (
vrc7_all_15_melodic_patches_match_nuke_ykt_canonicalinrustynes_apu::opll— that is thepatch_vrc7criterion). The absolute FM output vs the APU square is a pseudo-sine (not a square) and patch/TL/feedback-dependent, so it is not cleanly oracle-pinned; thedb_vrc7/clip_vrc7ROMs stay byte-exact snapshot regression guards.
NSF expansion-audio routing (v1.7.0 "Forge" G2/G3)¶
A classic .nsf may declare expansion audio in the $07B bitfield (bit 0 VRC6, 1 VRC7, 2 FDS, 3 MMC5, 4 N163, 5 5B). The NSF player (crates/rustynes-mappers/src/nsf.rs) does not reimplement any synthesis: crates/rustynes-mappers/src/nsf_expansion.rs (NsfExpansion) owns instances of the exact same cores listed above and routes the NSF register windows into them — $9000-$B002 (VRC6), $9010/$9030 (VRC7), $4040-$408A (FDS), $5000-$5015 (MMC5), $4800/$F800 (N163), $C000/$E000 (5B) — clocking on notify_cpu_cycle and fanning APU frame events (MMC5 envelope/length) on notify_frame_event. Because the bit-for-bit math is shared with the cartridge path, an NSF VRC6 tune sounds identical to a VRC6 cartridge. The $5FF8-$5FFF bank registers retain priority over the overlapping expansion windows. NsfExpansion is constructed only for NSF files and is unreachable from any oracle cartridge ROM, so it cannot perturb existing AccuracyCoin / blargg / kevtris audio.
MMC5 expansion audio (G3) was the one chip whose synthesis was started-but-deferred for NSF use; the cartridge Mmc5Audio core (2 pulse + raw PCM) is now driven for both cartridge-MMC5 and NSF-MMC5 playback through the shared router.
Open questions¶
- Sample-rate conversion: blip_buf-rs vs. hand-rolled. blip_buf-rs is a thin wrapper; we may inline it for fewer dependencies.
- Audio API choice in cpal:
f32vsi16output streams. cpal supports both; default device choice depends on platform. Architecture: emit i16 internally, convert to f32 in the cpal callback if needed.