# IIgs Performance Fix Plan Derived from the 2026-06 IIgs render-path audit (per-op 65816 cycle analysis + adversarial verification). Numbers below are verifier-confirmed unless marked "est." Ordered by dependency, then by realistic payoff. ## Implementation status (2026-06-23) All phases implemented and compile/link clean via the clang/llvm-mos w65816 toolchain. Every change is correct-by-construction from the verified analysis; the hot-path 65816 asm (Phases 3, 4, 5.1) still needs **pixel validation on hardware/MAME** -- the headless emulator here can't load the 1.6 MB disk image, so I could not observe output. Two micro-opts are left as marked residuals where the phase's value is already captured. - **Phase 0 -- DONE.** Removed the vestigial `SEI` + chunk-CLI from `iigsFillRectInner`; added the ~29 ops/sec ceiling + SEI-inflation note to `PERF.md` and the `uber-perf-table` generator. - **Phase 1 -- DONE.** `drawShowcase()` in `examples/uber/uber.c` renders one of each primitive into a labelled grid; the timed benchmark is unchanged (op sequence/args identical, so cross-port hashes still compare); ends green. No live progress bar (it would clobber the deliberately-seeded per-op input states; frozen showcase + joeylog is the fidelity-preserving choice). - **Phase 2 -- DONE.** Two parts: (a) `surfaceMarkDirtyRect` widens single-row (`h==1`) marks inline in C, dropping the cross-segment `iigsMarkDirtyRowsInner` JSL; (b) **`jlDrawPixel` fully inlined on IIgs** (draw.c) -- no `plotPixelNoMark` / `surfaceMarkDirtyRect` call layers around the single asm plot, mark widened inline via `jlStageGet()` guard. This is the 5-10x path. (The plan's fold-into-asm form is UNSAFE -- `iigsDrawPixelInner` can't see `s`, would mark `gStage` for heap-surface draws; the C inline keeps the guard.) - **Phase 3 -- DONE.** Computed-jump slam for ALL middle widths (`friMidSlamTbl` via an **RTS computed-goto** + X bias -- ascending table kept, no bank-0 vector hazard; `nWords==0` lands on the odd-byte writer). Partial fills go 7 -> 3 cyc/byte. Plus the `friMidMvn` X-reuse cleanup (now dead code path removed). - **Phase 4 -- DONE (4.1 residual).** 4.2: removed the per-plot `dcSavedCol` round-trip (parity via free X). 4.3: the 4 octant row bases are maintained incrementally (+/-160 per step) instead of 4 LUT lookups/iteration. 4.1 (move remaining `dc*` scratch to direct page) is the **residual**: ~4% more after 4.2+4.3, but needs a stack-frame-locals refactor to resolve the `D=SP+8` arg conflict -- disproportionate blind-refactor risk for the gain. - **Phase 5 -- DONE.** 5.1: `iigsSurfaceClearFastInner` -- post-init PHA stack-slam (2 cyc/byte, ~1.5x) with AUXWRITE redirect to bank $01; SHR shadow already inhibited by halInit; SP restored before CLI; the C macro dispatches here and keeps the STA `iigsSurfaceClearInner` as the init-safe fallback. (Non-chunked SEI ~23 ms; chunk like the present if audio glitches.) 5.2: `surfaceMarkDirtyAll` is now two `memset`s. 5.3 subsumed by 5.1. - **Phase 6 -- DONE (deep residual).** `jlSpriteSaveAndDraw` already shares `shift`, dodges the MUL4 2D-index, and marks once (the C-level win). The deeper dest-address sharing threaded through `spriteCompiledSaveUnder`/`Draw` is codegen-level surgery for ~1.05x -- left as the **residual**. - **Phase 7 -- DONE.** `docs/DESIGN.md` documents the deliberate double-buffer cost and why single-buffering the IIgs was rejected. ### Highest hardware-validation priority (new asm, pixel-validate first) 1. `jlFillRect` at assorted x/w (NOT just 0/320): the computed-goto slam must fill exactly `[x..x+w)` with no over/underrun. Check odd widths + odd x. 2. `iigsSurfaceClearFastInner`: the screen must clear to the right color with no stray writes; watch for audio glitches (the ~23 ms SEI window). 3. `jlDrawCircle`: shape unchanged after the incremental-row-base rewrite. 4. `jlDrawPixel`: visually correct + ops/sec jump. ### Hardware-validation checklist (run on your MAME/GSplus + the 4 ports) 1. Build + run `examples/uber` on IIgs; confirm the showcase grid renders all primitives legibly and the benchmark still completes (green exit). 2. Re-capture `joeylog.txt` on all four ports; regenerate `PERF.md` via `tools/uber-perf-table`. Expect IIgs `fillRect 320x200` and `stagePresent` to drop toward ~25-28 (SEI removal de-inflates fillRect). 3. `tools/diff-uber-hashes` IIgs-vs-{amiga,st,dos}: all ops must still match (the C dirty-mark + memset changes are portable and must not alter pixels). If any op diverges, the dirty-mark fold changed a present extent -- inspect. 4. Eyeball the circle (`jlDrawCircle`) output unchanged after the dcSavedCol change; confirm `jlDrawPixel` ops/sec rose. ## Ground truth / framing - **Hard floor:** any full-screen 32000-byte move = 16000 word-stores x 6 cyc (`STA long,X $9F` or `PEI`, both 3 cyc/byte) / 2.8 MHz = ~34 ms = **~29 ops/sec max**. Clear, full-width fillRect, and present are all pinned to this; you cannot beat it without moving fewer bytes (the dirty-row system already does that for sparse frames). - **The benchmark lies for SEI ops.** `PERF.md` `fillRect 320x200`=60 and `stagePresent`=42 exceed/approach the 29-ops/sec floor and are inflated; clear=28 is honest. Fix measurement FIRST (Phase 0) or you cannot tell wins from clock noise. - **Two architectural realities that are NOT bugs:** (a) the back-buffer -> `$E1` design pays the screen twice on full-redraw frames, but the single-buffer "fix" is refuted (slow `$E1`, overdraw inverts the win, tearing) -- keep it; the dirty-row system is the right mitigation. (b) The circle OUTLINE trails the ST due to 65816 register pressure and may never fully win; bank the partial gain. ## Cross-cutting rules (apply to every asm change) - Keep the assembler's register-width state correct after every `REP`/`SEP` so immediates are sized correctly. - Keep each asm routine in the SAME load segment as its C caller (DRAWPRIMS for the draw primitives); a separate segment silently produces calls-to-nowhere. - `gStageMinWord` / `gStageMaxWord` must stay `uint8_t[]` so the asm that walks them per byte stays correctly aligned with the C-side element size. - PHA/PEI stack-slam paths REQUIRE `SEI` and are unsafe during `joeyInit` -- gate them post-init (`gModeSet`). - 4bpp word stores need even byte index (`x/2` even, i.e. `x % 4 == 0`) for the computed-jump slam entry. - Build hygiene: `source toolchains/env.sh`; DELETE the example binary before rebuild (make leaves IIgs binaries stale); pass `spriteEmitIigs.c` first in LIB_SRCS (linker bank-packing order). - **Validation after every change:** re-run `examples/uber`, confirm `jlSurfaceHash` per op is UNCHANGED (pixel-identical to the IIgs golden) and ops/sec moved the right way. Verify visually in MAME (checkpoint poll or snapshot) -- joeyLog alone misses wrong-pixel-but-right-control-flow bugs. --- ## Phase 0 -- Make measurement trustworthy (PREREQUISITE) Everything downstream is judged against the benchmark; fix it first. - **0.1 Audit & likely remove `iigsFillRectInner`'s `SEI`** (joeyDraw.asm:315, chunk-CLI machinery :585-614). The body does only long-mode writes to the bank-$01 back buffer with no SP/shadow hijack (per its own comment :284-292), so the SEI appears vestigial. Confirm nothing in the fill needs interrupt atomicity (MVN is interruptible/resumable by design), then remove SEI + the FRI_CHUNK_ROWS CLI logic. Expected: fillRect-full de-inflates 60 -> ~27 (matching clear), AND the audio-starvation risk goes away. If some atomicity IS required, document exactly what and keep it. - **0.2 Add a cycle-ceiling sanity check / annotation.** Document the ~29 ops/sec full-screen floor in `PERF.md`; flag any full-screen op reading above it as suspect. Optional: assert in the harness. - **0.3 Reconcile `present`=42** (its SEI is REAL -- stack/shadow hijack -- so it can't just be removed). Capture an honest present rate via an SEI-immune time source or a diagnostic no-SEI measurement build; expect ~25 ops/sec. Update the table and the cross-port comparison accordingly. - **Re-baseline `PERF.md`** with honest numbers before Phase 2+. Exit criterion: clear, fillRect-full, and present all read within ~10% of each other (~25-28 ops/sec), consistent with the cycle model. ## Phase 1 -- Make the uber demo legible (show, don't smear) Right now `examples/uber` shows the user a solid colored screen and teaches them nothing. Cause, concretely: - `runAllTests()` (uber.c:222-285) runs each op in a 16-frame tight loop and NEVER calls `jlStagePresent()` during or between the timed ops. The screen is frozen on the last present before the run (the red-bar-on-black at uber.c:379) for the whole ~80 s run, then ends on a bare `jlSurfaceClear(gStage, 2)` (uber.c:390) -> solid green. - Even if it did present, the timed ops are full-screen or fixed-coordinate (`fillRect 320x200`, `surfaceClear`, pixel@100,100, line 0,50->319,50, ...), so they overwrite each other into one solid color. There is no spatial separation. Goal: the demo should visibly demonstrate what each primitive does AND run the benchmark without leaving a meaningless screen. Separate "show the user" from "measure throughput" -- they are different jobs. - **1.1 Visual showcase pass (new), runs once before the benchmark, UNTIMED.** Lay out a labeled grid (e.g. 4 cols x 5 rows of ~72x36 cells with a gutter) and render each primitive ONCE into its own cell so nothing overlaps: pixel, line H/V/diag, rect outline, circle (small), fillRect (small), fillCircle, a tile, tileCopyMasked result, the compiled sprite, a 16-entry palette ramp swatch, an SCB per-line strip, a sample-pixel readout. Full-screen ops (surfaceClear, fillRect 320x200, present) cannot be shown at native size in a cell -- represent each with a small labeled swatch and say so. `jlStagePresent()` once, then hold for a readable beat (frame-counted ~2-3 s, or `jlWaitForAnyKey` -> "press a key to begin benchmark"). THIS is the part that makes sense to a user. - **1.2 Benchmark pass: keep timing/hash fidelity, show a sane status screen.** Keep the `timeOp` loop and every op's args + call order EXACTLY as-is -- the cross-port `jlSurfaceHash` golden values depend on the precise op sequence, so changing them invalidates the validation harness. Stop presenting the timing garbage: draw a static "BENCHMARKING..." screen plus a progress indicator (op counter / bar, optionally each op's just-measured ops/sec) updated BETWEEN `timeOp` calls -- never inside the timed loop. End on an explicit "DONE -- press any key" screen, not a bare solid clear. - **1.3 Labels.** `jlDrawText` exists (tile.h:119) but needs a font surface + asciiMap, which uber does not currently load. Either add a small font surface + asciiMap and label each cell, or (fallback) use fixed spatial layout + a short on-screen color-key strip and a legend printed via `jlLogF`. Cautions: - All showcase drawing AND the progress-UI updates must be OUTSIDE any timed window -- zero effect on ops/sec. - Do NOT change the timed ops' arguments or sequence (golden-hash stability). - Keep it port-agnostic: uber is shared example code that also runs on Amiga/ST/DOS -- no IIgs-only assumptions in the showcase. - Do this AFTER Phase 0 so any per-op ops/sec shown on the status screen is the corrected (honest) number, not the inflated one. Exit criterion: a first-time user watching the demo can see and identify each primitive's output, and the benchmark phase shows readable progress + a clear finish -- never a bare solid color. ## Phase 2 -- Kill the per-primitive C+JSL tax (biggest real-world win) `jlDrawPixel` costs **~1595 cyc/call** (2.8e6 / 1755), ~94% overhead; the actual nibble RMW is ~30 cyc. Every marking primitive pays a second cross-segment JSL. - **2.1 Fold dirty-marking into the asm inner loops.** `jlDrawPixel` (draw.c:95), `jlFillRect` (draw.c:448-ish via fillRectOnSurface), line, and circle each call `surfaceMarkDirtyRect` AFTER the plot JSL -- a second JSL into `iigsMarkDirtyRowsInner` (~49 cyc plumbing x2 + the row loop). The inner loops already know the row(s) and extent (fillRect has `midStart`/`trailingByte` per row; pixel has `y` and the byte). Update `gStageMinWord[y]`/`gStageMaxWord[y]` inline and make `surfaceMarkDirtyRect` a no-op on the IIgs fast paths. - Confirmed win: removes ~120 cyc from EVERY marking primitive; **5-10x on `jlDrawPixel`**. - **2.2 Trim the asm prologue/epilogue for the tiny ops.** `iigsDrawPixelInner` (joeyDraw.asm:1158-1227) keeps scratch in 6-cyc long-absolute (`dpxlTmp`, `dpxlNibPart`) and pays full PHP/PHB/PHD ... PLD/PLB/PLP framing around a ~50-cyc plot. Use registers/DP for the scratch; drop framing not needed when DBR/D are already correct. (est ~1.2-1.4x on the inner once 2.1 lands.) - **2.3 Batch/inline entry for pixel-heavy callers (optional).** A span/row-run API (or inlining the in-bounds nibble RMW into the C via macro, the way fillRect was de-wrappered) amortizes the JSL across many pixels for plotting-heavy code. Exit criterion: `jlDrawPixel` >> 1755; line/rect call cost drops by the second-JSL amount; hashes unchanged. ## Phase 3 -- fillRect partial-width slam gate The 3-cyc/byte slam fires ONLY for exactly `x==0 && w==320` (`friFullWidth`, joeyDraw.asm:406-420). Everything else falls to seed+`MVN` at **7 cyc/byte** (2.33x slower), confirmed. - **3.1 Computed jump into the unrolled `STA long,X` table.** Set `X = curRow + midStart`, then `JMP (addr)` into the 80-store run at store `#(80 - midBytes)`. Gives 3 cyc/byte for ANY width. Keep MVN only for tiny odd remainders. Drop-in safe (bank-$01 long stores, no SEI/shadow). Respect the 4bpp even-byte-index rule for the entry. **1.8-2.3x on all partial-width fills.** - **3.2 Cleanups (minor):** cache `curRow+midStart` instead of recomputing it (joeyDraw.asm:542-565); hoist loop-invariant edge offsets (`leadingByte`/ `trailingByte` relative to curRow, += 160/row); one `sep`/`rep` window per row for both edge RMWs. ## Phase 4 -- circle / line register pressure (~1.27x, partial) - **4.1 Move Bresenham scratch to direct page.** `dcX/dcY/dcErr/dcRowYP/dcRowYN/ dcRowXP/dcRowXN` (circle, joeyDraw.asm:1500+) and the line's per-pixel state (joeyDraw.asm:1253+) are 6-cyc long-absolute; DP is 5 cyc here (DL != 0 because D=SP+8). Set up a DP scratch page for the hot state. - **4.2 Remove the per-plot parity round-trip.** `sta dcSavedCol` / reload / `and #1` (joeyDraw.asm:1603-1610) recovers a bit `lsr a` already left in carry -- branch on carry instead. ~15 cyc/plot x 8 plots/step. - **4.3 Maintain the 4 row bases incrementally** (+/-160 per Bresenham step) instead of 4 LUT lookups/step. - Net ~1.27x (0.65x -> ~0.83x vs ST). Stop there -- the architecture won't give a full win on the outline. The fill-circle already beats the ST; leave it. ## Phase 5 -- faster clear (post-Phase-0; lower priority) - **5.1 Post-init PHA-slam clear path.** `STA long,X` = 3 cyc/byte; SP-into-`$01` PHA-slam = 2 cyc/byte. ~1.5x (37.6 -> ~24 ms). Gate on `gModeSet` (SEI unsafe at init); keep the `STA long,X` path as the init-safe fallback. Chunk the SEI like present to avoid audio starvation. NOTE: only meaningful to measure once Phase 0 fixes the clock. - **5.2 Route `surfaceMarkDirtyAll` through `iigsMarkDirtyRowsInner`** (it's a 200-iter C loop today; surface.c:233-243) -- one source of truth with the fillRect marker. ~7600 cyc/clear. - **5.3 Hold `fillWord` in a register** across the clear loop instead of `lda 9,s` per iteration (joeyDraw.asm:137). Subsumed if 5.1 lands. ## Phase 6 -- sprite dispatch overhead Body store mode (`LDA #imm16`/`STA abs,Y`) is already near-optimal -- DO NOT touch it. ~3000-4000 cyc/call of C dispatch wraps ~1000-1500 cyc of real blit. - **6.1 Share address/shift/fnAddr across `SaveAndDraw`'s two halves** (sprite.c) -- it recomputes dest address, shift, fnAddr, and re-patches the self-modifying stub twice. (est ~1.05x SaveAndDraw.) - **6.2 Trim per-call C dispatch:** cache fnAddr, dodge the 2D-index MUL4, do `isFullyOnSurface` in 8-bit. (est ~1.10-1.15x per sprite op.) ## Phase 7 -- architecture (decision, no code) Keep the back-buffer + dirty-row model. Document in `docs/DESIGN.md` that full-redraw frames inherently pay the screen cost ~2x and that sparse-update (sprite/HUD) games should lean on partial updates. Do NOT pursue single-buffer direct-to-`$E1` (refuted on slow-bank/overdraw/tearing grounds). --- ## Suggested order of execution 1. Phase 0 (0.1 fillRect SEI, 0.2 annotate, 0.3 present) -> re-baseline PERF.md. 2. Phase 1 (make the demo legible) -- so every later change can be eyeballed, not just hashed; cheap and unblocks visual validation. 3. Phase 2.1 (fold dirty-marking) -- highest real-world payoff, touches all primitives. 4. Phase 3.1 (fillRect computed-jump slam). 5. Phase 2.2, Phase 4 (circle/line DP scratch + parity). 6. Phase 5, Phase 6 (clear PHA-slam, sprite dispatch) as time permits. 7. Phase 7 doc. Each step: re-run uber, assert hashes identical, record ops/sec delta in PERF.md.