Scheduler¶
References: ref-docs/research-report.md §Architecture options;
ref-docs/nesdev-wiki-technical-report.md §Emulator Architecture Guidance;
docs/architecture.md §Scheduling model; Nesdev
DMA and
APU Frame Counter.
Purpose¶
The scheduler is the heart of the cycle-accurate emulator: it advances the PPU, CPU, APU, mapper IRQ counters, and DMA controller in tight lockstep at PPU-dot resolution. It lives in crates/rustynes-core and is the single owner of the Nes run loop.
v2.0.0 "Timebase" (promoted in beta.4): the shipped scheduler is the one-clock, every-cycle-bus-access model —
Cpu::master_clockadvances by the region divider per CPU cycle (asymmetric read 5/7, write 7/5 φ1/φ2 split on NTSC; PAL 16, Dendy 15) with the PPU pulled tomaster_clock − PPU_OFFSETat both half-cycles (run_ppu_todouble catch-up);SystemBus::cycleis the ONE canonical per-cycle counter (Cpu::cycles/Apu::cpu_cycleare assigned from it, never independently incremented); every instruction cycle is a real bus access (no busless cycles); DMA is the per-cycle interleaved unified engine; and the warm reset is a clocked sequence with the$4017re-write (seedocs/cpu-6502.md+docs/apu-2a03.md). Thetick_one_dotprimer below describes the original dot-lockstep model and the DMA controller's rules, which the unified engine preserves semantically; the full doc re-baseline lands with the v2.0.0 rc (ADR 0017).
Design¶
One tick = one PPU dot (historical — see the v2.0.0 note above)¶
fn tick_one_dot(&mut self) {
self.ppu.tick(&mut self.ppu_bus);
self.dot_count += 1;
if self.dot_count % 3 == self.cpu_phase {
self.cpu_tick();
}
}
cpu_phase is a per-power-on offset (0, 1, or 2) representing the random initial CPU/PPU alignment per real hardware. Reset does not change it. Cold power-cycle re-rolls it from a deterministic PRNG seeded by the user (default 0).
CPU tick (historical — see the v2.0.0 note above)¶
fn cpu_tick(&mut self) {
if self.dma.cycles_remaining > 0 {
self.dma.tick(&mut self.bus); // owns the bus during DMA
return;
}
if let Some(req) = self.dma.scheduled_request() {
if self.cpu.next_cycle_is_read() {
self.dma.start(req);
return;
}
}
self.cpu.tick(&mut self.bus);
if self.cpu_cycle.is_apu_cycle() {
self.apu.tick(&mut self.apu_bus);
}
}
Bus design¶
The bus (SystemBus) owns: PPU, APU, mapper (via cart), WRAM, controllers, open-bus latch. The CPU borrows it for each instruction (Cpu::step(&mut bus), generic over rustynes_cpu::Bus). The PPU gets its own bus trait for what it needs:
PpuBus:ppu_read/ppu_read_sprite/ppu_writeand the nametable accessors (delegated to the mapper, which owns CIRAM mapping); the mapper notificationsnotify_a12,notify_scanline_startandnotify_vblank(mapper IRQs).- The APU has no bus trait. The bus drives the DMC sample fetch itself: it polls
Apu::dmc_dma_pending/dmc_dma_addr, performs the read inside the CPU's unified DMA, and returns the byte withApu::complete_dmc_dma. The APU's IRQ line is read throughbus.irq_level(). (AnApuBustrait existed, unimplemented, until v2.9.8 removed it with ADR 0042; until v2.7.5 this section described it as live — core audit §4.3.) - The /NMI line is a level, too: the CPU reads
bus.nmi_level()every cycle and edge-detects it itself. v2.9.8 removed the bus-side edge detector (sample_nmi_edge, run on every PPU dot inrun_ppu_to) and thepoll_nmi/poll_irqfamily it fed (ADR 0042); none had a caller after v2.0.0.
The mapper sees both buses via separate trait methods (cpu_read/write, ppu_read/write).
DMA controller¶
A small inner struct that tracks:
cycles_remaining: u16— non-zero means CPU is halted by DMA.scheduled: Option<DmaRequest>— pending DMA the controller is waiting to start.- DMA type:
OamDma { src_page: u8 },DmcDma { addr: u16, kind: Load | Reload }.
Scheduling rules per ref-docs/research-report.md §DMA:
- DMA can only halt on a CPU read cycle. A load DMA a write refuses enters
on the next read whichever half it is (four cycles after one refusing
write;
docs/apu-2a03.md, v3.1.0). - DMC DMA gets precedence over OAM DMA.
- OAM DMA: 1 halt + (0 or 1 alignment) + 256 read/write pairs = 513 or 514 cycles.
- DMC DMA: 1 halt + 1 dummy + (0 or 1 alignment) + 1 read = 3 or 4 cycles.
Load and reload DMC DMA are not interchangeable. Load DMA is scheduled after a
$4015 enable write and reload DMA is scheduled when the DMC sample buffer
empties; the two start on different get/put phases before halt delay is
considered. The scheduler must preserve this distinction all the way to the bus
because repeated halted reads of $2007, $4015, $4016, and $4017 are
observable side effects.
The v2.0.0 "Timebase" rewrite collapsed these three drivers (standalone OAM,
standalone DMC, and the DMC-during-OAM overlap) into one per-cycle engine —
SystemBus::unified_dma_cycle_impl in crates/rustynes-core/src/bus.rs, a
direct port of TriCNES's _6502 DMA dispatch table. The 513/514 (OAM) and 3/4
(DMC) spans are emergent from a single get/put parity label
(get = !apu.put_cycle()) rather than an owed-cycle counter. The committed
oracle floor for this engine — the five dmc_dma_during_read4 ROMs, both
sprdma_and_dmc_dma variants, and dma_timing_pin (the AccuracyCoin
CheckDMATiming reload span = 4 + the $50-$5F DMC-during-OAM landing sweep on
KEY) — is all green.
Unexpected DMA (2A03 die revision — the documented frontier)¶
nesdev DMA notes that when a DMC-DMA halt
coincides with an OAM-DMA halt (the "double-halt" overlap) some 2A03 silicon
inserts an extra re-read of the parked 6502 address bus, and this differs by
mask revision (RP2A03G vs RP2A03H). RustyNES exposes this as the additive
Cpu2A03Revision { Rp2A03G (default), Rp2A03H } config (Nes::set_cpu_2a03_revision),
gating the halted-DMC parked-address re-read that fires inside an OAM-owned read
cycle during a DMC+OAM overlap. Rp2A03G performs it (the byte-identical
default); Rp2A03H omits it.
This is an unclosed, honestly-documented frontier (ADR 0033): no public
reference emulator branches DMA behavior on 2A03 die stepping and no test ROM
captures it, and on this engine the gate — while at its mechanism-correct
location — is behaviorally inert, because the parked address during a DMC+OAM
overlap is always the post-$4014 instruction fetch (OAM drains on the next
opcode read), never a $2002/$2007/$4015/$4016/$4017 register. So
Rp2A03H is byte-identical to Rp2A03G on every oracle. The knob is a config
re-applied on load, not part of the save-state (determinism is preserved for
a fixed revision). See ADR 0033 and the cpu_2a03_revision test suite.
Region cadence¶
NTSC and Dendy can use a simple 3 PPU dots per CPU cycle cadence. PAL needs a fractional or master-clock representation because its PPU:CPU ratio is 3.2. APU frame-counter tables also differ by region; do not scale NTSC cycle counts for PAL.
CPU-multiplier overclock (v3.1.0, T-CPU-OVERCLOCK)¶
Nes::set_cpu_overclock(k), k in 1..=4 (MAX_CPU_OVERCLOCK), divides the
region's master-clock CPU divider exactly: NTSC 12 becomes 6, 4 or 3 per
cycle; PAL 16 and Dendy 15 alternate cycle lengths so that every k cycles
take exactly one stock cycle (PAL x3: 5, 5, 6; Dendy x4: 3, 4, 4, 4; see
below). The PPU's divider is untouched, so the CPU gets exactly k times the
cycles per frame on every region. The shortest cycle, 3, still leaves both the
read split (0, 3) and the write split (2, 1) valid.
Everything that measures console time stays at the stock rate: the APU (and
with it the DMC), the mappers' notify_cpu_cycle IRQ counters, and the PPU's
open-bus decay and post-reset timers. Every k CPU cycles make exactly one
stock cycle: the bus counts overclock_phase through 0..k, CPU cycle i of
the group lasts ((i+1)*div)/k - (i*div)/k master clocks (the lengths sum to
the stock divider div), and a stock step (those devices advance once)
runs on the last cycle of each group. The cycle length changes only at a
cycle's end, because the CPU reads the divider once for each half of a cycle.
Until the v3.1.0 review the length was div / k rounded down, exact on NTSC
and not elsewhere (PAL x3 ran 3.2x, Dendy x4 5x). The APU is handed its own counter (apu_cycle)
instead of the CPU's, so its put/get phase advances once per stock step. DMA
follows that phase, so a DMA takes about k times as many CPU cycles and the
same real time.
At k = 1 the branch is never taken: every cycle is a stock step and the APU
gets the CPU counter, so the output is byte-identical (the epoch fingerprint
gate's whole panel runs at k = 1). The phase and apu_cycle are in the BUS
save-state section (version 3), because run-ahead restores mid-run; the
multiplier is configuration, carried by HardwareOptions in movies and the
netplay config_digest. run_frame's cycle budget scales by k. Not hardware
behaviour: no console runs its CPU faster than its APU.
Tests: crates/rustynes-test-harness/tests/cpu_overclock.rs.
Frame complete¶
The PPU reports frame_complete() when it transitions from scanline 240 to 241 (i.e., the start of vertical blank). The frontend consumes the framebuffer at this point.
Audio drain¶
The APU emits samples into a band-limited buffer continuously. The frontend's audio callback (on the cpal thread) reads from a ring buffer the run loop fills via apu.drain_samples(&mut buf) once per frame.
Determinism¶
The scheduler is fully deterministic given:
- A fixed
cpu_phase(chosen at power-cycle from a seedable PRNG). - A fixed initial WRAM pattern (deterministic seeded fill).
- A fixed sequence of controller inputs (recorded by the test harness).
- A fixed initial DMA get/put phase.
- A fixed region timing profile and reset/power-up mode.
This guarantees that save/load round-trips and a re-played input sequence produce bit-identical framebuffer + audio output. Required for movie playback, regression tests, and netplay.
Performance targets¶
These are the original design-phase aspirations, not gates. The frame-cost figure was not met and is knowingly accepted — the implemented cycle-accurate core measures ~3.95 ms (nestest) on the shipped fast dot path, ~23% of the 16.639 ms NTSC budget, and ~2.65 ms on flowing palette, a rendering-disabled control whose fast-path variant never enters that path (its guard bails) (the v2.7.0 core, 2026-09-23; v2.9.7's A12 change added about 1.9% on nestest). See
docs/performance.md§"Current figures" for the measured numbers, the exact-path pair, and why the main optimization levers were measured and rejected.
- Frame cost (single-thread, headless core — no frontend, no present): ≤ 2 ms on a 2018-era laptop x86_64 (Skylake-era) — aspirational; ~3.95 ms measured and accepted (
nes_run_frame_nestest_fast, the shipped fast dot path, which renders; the render-lightflowing_paletteworkload measures ~2.65 ms). - Frame cost including wgpu present + cpal callback: ≤ 5 ms (well under the 16.67 ms budget for 60 fps NTSC).
- Audio callback: lock-free SPSC ring buffer; never block the audio thread.
Open questions¶
- Inline vs trait dispatch. The PPU
tick()and CPUtick()are the hot paths. Initial implementation uses trait objects (Box<dyn Mapper>) for the mapper. If profiling shows mapper-dispatch overhead > 5%, switch to a monomorphized enum for the supported mappers. - SIMD for framebuffer scaling. Initial scaling is GPU-side via wgpu (sampler with nearest filter). If we add per-pixel post-processing (CRT shader, scanline), it's GPU-side as well. No CPU SIMD planned.
- Multithreading. Not planned for v1.0. The single-frame work fits comfortably in one thread.