Skip to content

Scheduler

References: ref-docs/research-report.md §Architecture options; ref-docs/nesdev-wiki-technical-report.md §Emulator Architecture Guidance; docs/architecture.md §Scheduling model; Nesdev DMA and APU Frame Counter.

Purpose

The scheduler is the heart of the cycle-accurate emulator: it advances the PPU, CPU, APU, mapper IRQ counters, and DMA controller in tight lockstep at PPU-dot resolution. It lives in crates/rustynes-core and is the single owner of the Nes run loop.

v2.0.0 "Timebase" (promoted in beta.4): the shipped scheduler is the one-clock, every-cycle-bus-access model — Cpu::master_clock advances by the region divider per CPU cycle (asymmetric read 5/7, write 7/5 φ1/φ2 split on NTSC; PAL 16, Dendy 15) with the PPU pulled to master_clock − PPU_OFFSET at both half-cycles (run_ppu_to double catch-up); SystemBus::cycle is the ONE canonical per-cycle counter (Cpu::cycles / Apu::cpu_cycle are assigned from it, never independently incremented); every instruction cycle is a real bus access (no busless cycles); DMA is the per-cycle interleaved unified engine; and the warm reset is a clocked sequence with the $4017 re-write (see docs/cpu-6502.md + docs/apu-2a03.md). The tick_one_dot primer below describes the original dot-lockstep model and the DMA controller's rules, which the unified engine preserves semantically; the full doc re-baseline lands with the v2.0.0 rc (ADR 0017).

Design

One tick = one PPU dot (historical — see the v2.0.0 note above)

fn tick_one_dot(&mut self) {
    self.ppu.tick(&mut self.ppu_bus);
    self.dot_count += 1;
    if self.dot_count % 3 == self.cpu_phase {
        self.cpu_tick();
    }
}

cpu_phase is a per-power-on offset (0, 1, or 2) representing the random initial CPU/PPU alignment per real hardware. Reset does not change it. Cold power-cycle re-rolls it from a deterministic PRNG seeded by the user (default 0).

CPU tick (historical — see the v2.0.0 note above)

fn cpu_tick(&mut self) {
    if self.dma.cycles_remaining > 0 {
        self.dma.tick(&mut self.bus);   // owns the bus during DMA
        return;
    }
    if let Some(req) = self.dma.scheduled_request() {
        if self.cpu.next_cycle_is_read() {
            self.dma.start(req);
            return;
        }
    }
    self.cpu.tick(&mut self.bus);
    if self.cpu_cycle.is_apu_cycle() {
        self.apu.tick(&mut self.apu_bus);
    }
}

Bus design

The bus (SystemBus) owns: PPU, APU, mapper (via cart), WRAM, controllers, open-bus latch. The CPU borrows it for each instruction (Cpu::step(&mut bus), generic over rustynes_cpu::Bus). The PPU gets its own bus trait for what it needs:

  • PpuBus: ppu_read / ppu_read_sprite / ppu_write and the nametable accessors (delegated to the mapper, which owns CIRAM mapping); the mapper notifications notify_a12, notify_scanline_start and notify_vblank (mapper IRQs).
  • The APU has no bus trait. The bus drives the DMC sample fetch itself: it polls Apu::dmc_dma_pending / dmc_dma_addr, performs the read inside the CPU's unified DMA, and returns the byte with Apu::complete_dmc_dma. The APU's IRQ line is read through bus.irq_level(). (An ApuBus trait existed, unimplemented, until v2.9.8 removed it with ADR 0042; until v2.7.5 this section described it as live — core audit §4.3.)
  • The /NMI line is a level, too: the CPU reads bus.nmi_level() every cycle and edge-detects it itself. v2.9.8 removed the bus-side edge detector (sample_nmi_edge, run on every PPU dot in run_ppu_to) and the poll_nmi / poll_irq family it fed (ADR 0042); none had a caller after v2.0.0.

The mapper sees both buses via separate trait methods (cpu_read/write, ppu_read/write).

DMA controller

A small inner struct that tracks:

  • cycles_remaining: u16 — non-zero means CPU is halted by DMA.
  • scheduled: Option<DmaRequest> — pending DMA the controller is waiting to start.
  • DMA type: OamDma { src_page: u8 }, DmcDma { addr: u16, kind: Load | Reload }.

Scheduling rules per ref-docs/research-report.md §DMA:

  • DMA can only halt on a CPU read cycle. A load DMA a write refuses enters on the next read whichever half it is (four cycles after one refusing write; docs/apu-2a03.md, v3.1.0).
  • DMC DMA gets precedence over OAM DMA.
  • OAM DMA: 1 halt + (0 or 1 alignment) + 256 read/write pairs = 513 or 514 cycles.
  • DMC DMA: 1 halt + 1 dummy + (0 or 1 alignment) + 1 read = 3 or 4 cycles.

Load and reload DMC DMA are not interchangeable. Load DMA is scheduled after a $4015 enable write and reload DMA is scheduled when the DMC sample buffer empties; the two start on different get/put phases before halt delay is considered. The scheduler must preserve this distinction all the way to the bus because repeated halted reads of $2007, $4015, $4016, and $4017 are observable side effects.

The v2.0.0 "Timebase" rewrite collapsed these three drivers (standalone OAM, standalone DMC, and the DMC-during-OAM overlap) into one per-cycle engine — SystemBus::unified_dma_cycle_impl in crates/rustynes-core/src/bus.rs, a direct port of TriCNES's _6502 DMA dispatch table. The 513/514 (OAM) and 3/4 (DMC) spans are emergent from a single get/put parity label (get = !apu.put_cycle()) rather than an owed-cycle counter. The committed oracle floor for this engine — the five dmc_dma_during_read4 ROMs, both sprdma_and_dmc_dma variants, and dma_timing_pin (the AccuracyCoin CheckDMATiming reload span = 4 + the $50-$5F DMC-during-OAM landing sweep on KEY) — is all green.

Unexpected DMA (2A03 die revision — the documented frontier)

nesdev DMA notes that when a DMC-DMA halt coincides with an OAM-DMA halt (the "double-halt" overlap) some 2A03 silicon inserts an extra re-read of the parked 6502 address bus, and this differs by mask revision (RP2A03G vs RP2A03H). RustyNES exposes this as the additive Cpu2A03Revision { Rp2A03G (default), Rp2A03H } config (Nes::set_cpu_2a03_revision), gating the halted-DMC parked-address re-read that fires inside an OAM-owned read cycle during a DMC+OAM overlap. Rp2A03G performs it (the byte-identical default); Rp2A03H omits it.

This is an unclosed, honestly-documented frontier (ADR 0033): no public reference emulator branches DMA behavior on 2A03 die stepping and no test ROM captures it, and on this engine the gate — while at its mechanism-correct location — is behaviorally inert, because the parked address during a DMC+OAM overlap is always the post-$4014 instruction fetch (OAM drains on the next opcode read), never a $2002/$2007/$4015/$4016/$4017 register. So Rp2A03H is byte-identical to Rp2A03G on every oracle. The knob is a config re-applied on load, not part of the save-state (determinism is preserved for a fixed revision). See ADR 0033 and the cpu_2a03_revision test suite.

Region cadence

NTSC and Dendy can use a simple 3 PPU dots per CPU cycle cadence. PAL needs a fractional or master-clock representation because its PPU:CPU ratio is 3.2. APU frame-counter tables also differ by region; do not scale NTSC cycle counts for PAL.

CPU-multiplier overclock (v3.1.0, T-CPU-OVERCLOCK)

Nes::set_cpu_overclock(k), k in 1..=4 (MAX_CPU_OVERCLOCK), divides the region's master-clock CPU divider exactly: NTSC 12 becomes 6, 4 or 3 per cycle; PAL 16 and Dendy 15 alternate cycle lengths so that every k cycles take exactly one stock cycle (PAL x3: 5, 5, 6; Dendy x4: 3, 4, 4, 4; see below). The PPU's divider is untouched, so the CPU gets exactly k times the cycles per frame on every region. The shortest cycle, 3, still leaves both the read split (0, 3) and the write split (2, 1) valid.

Everything that measures console time stays at the stock rate: the APU (and with it the DMC), the mappers' notify_cpu_cycle IRQ counters, and the PPU's open-bus decay and post-reset timers. Every k CPU cycles make exactly one stock cycle: the bus counts overclock_phase through 0..k, CPU cycle i of the group lasts ((i+1)*div)/k - (i*div)/k master clocks (the lengths sum to the stock divider div), and a stock step (those devices advance once) runs on the last cycle of each group. The cycle length changes only at a cycle's end, because the CPU reads the divider once for each half of a cycle. Until the v3.1.0 review the length was div / k rounded down, exact on NTSC and not elsewhere (PAL x3 ran 3.2x, Dendy x4 5x). The APU is handed its own counter (apu_cycle) instead of the CPU's, so its put/get phase advances once per stock step. DMA follows that phase, so a DMA takes about k times as many CPU cycles and the same real time.

At k = 1 the branch is never taken: every cycle is a stock step and the APU gets the CPU counter, so the output is byte-identical (the epoch fingerprint gate's whole panel runs at k = 1). The phase and apu_cycle are in the BUS save-state section (version 3), because run-ahead restores mid-run; the multiplier is configuration, carried by HardwareOptions in movies and the netplay config_digest. run_frame's cycle budget scales by k. Not hardware behaviour: no console runs its CPU faster than its APU. Tests: crates/rustynes-test-harness/tests/cpu_overclock.rs.

Frame complete

The PPU reports frame_complete() when it transitions from scanline 240 to 241 (i.e., the start of vertical blank). The frontend consumes the framebuffer at this point.

Audio drain

The APU emits samples into a band-limited buffer continuously. The frontend's audio callback (on the cpal thread) reads from a ring buffer the run loop fills via apu.drain_samples(&mut buf) once per frame.

Determinism

The scheduler is fully deterministic given:

  • A fixed cpu_phase (chosen at power-cycle from a seedable PRNG).
  • A fixed initial WRAM pattern (deterministic seeded fill).
  • A fixed sequence of controller inputs (recorded by the test harness).
  • A fixed initial DMA get/put phase.
  • A fixed region timing profile and reset/power-up mode.

This guarantees that save/load round-trips and a re-played input sequence produce bit-identical framebuffer + audio output. Required for movie playback, regression tests, and netplay.

Performance targets

These are the original design-phase aspirations, not gates. The frame-cost figure was not met and is knowingly accepted — the implemented cycle-accurate core measures ~3.95 ms (nestest) on the shipped fast dot path, ~23% of the 16.639 ms NTSC budget, and ~2.65 ms on flowing palette, a rendering-disabled control whose fast-path variant never enters that path (its guard bails) (the v2.7.0 core, 2026-09-23; v2.9.7's A12 change added about 1.9% on nestest). See docs/performance.md §"Current figures" for the measured numbers, the exact-path pair, and why the main optimization levers were measured and rejected.

  • Frame cost (single-thread, headless core — no frontend, no present): ≤ 2 ms on a 2018-era laptop x86_64 (Skylake-era) — aspirational; ~3.95 ms measured and accepted (nes_run_frame_nestest_fast, the shipped fast dot path, which renders; the render-light flowing_palette workload measures ~2.65 ms).
  • Frame cost including wgpu present + cpal callback: ≤ 5 ms (well under the 16.67 ms budget for 60 fps NTSC).
  • Audio callback: lock-free SPSC ring buffer; never block the audio thread.

Open questions

  • Inline vs trait dispatch. The PPU tick() and CPU tick() are the hot paths. Initial implementation uses trait objects (Box<dyn Mapper>) for the mapper. If profiling shows mapper-dispatch overhead > 5%, switch to a monomorphized enum for the supported mappers.
  • SIMD for framebuffer scaling. Initial scaling is GPU-side via wgpu (sampler with nearest filter). If we add per-pixel post-processing (CRT shader, scanline), it's GPU-side as well. No CPU SIMD planned.
  • Multithreading. Not planned for v1.0. The single-frame work fits comfortably in one thread.