294 lines
17 KiB
Markdown
294 lines
17 KiB
Markdown
# IIgs Performance Fix Plan
|
|
|
|
Derived from the 2026-06 IIgs render-path audit (per-op 65816 cycle analysis +
|
|
adversarial verification). Numbers below are verifier-confirmed unless marked
|
|
"est." Ordered by dependency, then by realistic payoff.
|
|
|
|
## Implementation status (2026-06-23)
|
|
|
|
All phases implemented and compile/link clean via the clang/llvm-mos w65816
|
|
toolchain. Every change is
|
|
correct-by-construction from the verified analysis; the hot-path 65816 asm
|
|
(Phases 3, 4, 5.1) still needs **pixel validation on hardware/MAME** -- the
|
|
headless emulator here can't load the 1.6 MB disk image, so I could not observe
|
|
output. Two micro-opts are left as marked residuals where the phase's value is
|
|
already captured.
|
|
|
|
- **Phase 0 -- DONE.** Removed the vestigial `SEI` + chunk-CLI from
|
|
`iigsFillRectInner`; added the ~29 ops/sec ceiling + SEI-inflation note to
|
|
`PERF.md` and the `uber-perf-table` generator.
|
|
- **Phase 1 -- DONE.** `drawShowcase()` in `examples/uber/uber.c` renders one of
|
|
each primitive into a labelled grid; the timed benchmark is unchanged (op
|
|
sequence/args identical, so cross-port hashes still compare); ends green. No
|
|
live progress bar (it would clobber the deliberately-seeded per-op input
|
|
states; frozen showcase + joeylog is the fidelity-preserving choice).
|
|
- **Phase 2 -- DONE.** Two parts: (a) `surfaceMarkDirtyRect` widens single-row
|
|
(`h==1`) marks inline in C, dropping the cross-segment `iigsMarkDirtyRowsInner`
|
|
JSL; (b) **`jlDrawPixel` fully inlined on IIgs** (draw.c) -- no `plotPixelNoMark`
|
|
/ `surfaceMarkDirtyRect` call layers around the single asm plot, mark widened
|
|
inline via `jlStageGet()` guard. This is the 5-10x path. (The plan's
|
|
fold-into-asm form is UNSAFE -- `iigsDrawPixelInner` can't see `s`, would mark
|
|
`gStage` for heap-surface draws; the C inline keeps the guard.)
|
|
- **Phase 3 -- DONE.** Computed-jump slam for ALL middle widths (`friMidSlamTbl`
|
|
via an **RTS computed-goto** + X bias -- ascending table kept, no bank-0 vector
|
|
hazard; `nWords==0` lands on the odd-byte writer). Partial fills go 7 -> 3
|
|
cyc/byte. Plus the `friMidMvn` X-reuse cleanup (now dead code path removed).
|
|
- **Phase 4 -- DONE (4.1 residual).** 4.2: removed the per-plot `dcSavedCol`
|
|
round-trip (parity via free X). 4.3: the 4 octant row bases are maintained
|
|
incrementally (+/-160 per step) instead of 4 LUT lookups/iteration. 4.1 (move
|
|
remaining `dc*` scratch to direct page) is the **residual**: ~4% more after
|
|
4.2+4.3, but needs a stack-frame-locals refactor to resolve the `D=SP+8` arg
|
|
conflict -- disproportionate blind-refactor risk for the gain.
|
|
- **Phase 5 -- DONE.** 5.1: `iigsSurfaceClearFastInner` -- post-init PHA stack-slam
|
|
(2 cyc/byte, ~1.5x) with AUXWRITE redirect to bank $01; SHR shadow already
|
|
inhibited by halInit; SP restored before CLI; the C macro dispatches here and
|
|
keeps the STA `iigsSurfaceClearInner` as the init-safe fallback. (Non-chunked
|
|
SEI ~23 ms; chunk like the present if audio glitches.) 5.2: `surfaceMarkDirtyAll`
|
|
is now two `memset`s. 5.3 subsumed by 5.1.
|
|
- **Phase 6 -- DONE (deep residual).** `jlSpriteSaveAndDraw` already shares
|
|
`shift`, dodges the MUL4 2D-index, and marks once (the C-level win). The deeper
|
|
dest-address sharing threaded through `spriteCompiledSaveUnder`/`Draw` is
|
|
codegen-level surgery for ~1.05x -- left as the **residual**.
|
|
- **Phase 7 -- DONE.** `docs/DESIGN.md` documents the deliberate double-buffer
|
|
cost and why single-buffering the IIgs was rejected.
|
|
|
|
### Highest hardware-validation priority (new asm, pixel-validate first)
|
|
|
|
1. `jlFillRect` at assorted x/w (NOT just 0/320): the computed-goto slam must
|
|
fill exactly `[x..x+w)` with no over/underrun. Check odd widths + odd x.
|
|
2. `iigsSurfaceClearFastInner`: the screen must clear to the right color with no
|
|
stray writes; watch for audio glitches (the ~23 ms SEI window).
|
|
3. `jlDrawCircle`: shape unchanged after the incremental-row-base rewrite.
|
|
4. `jlDrawPixel`: visually correct + ops/sec jump.
|
|
|
|
### Hardware-validation checklist (run on your MAME/GSplus + the 4 ports)
|
|
|
|
1. Build + run `examples/uber` on IIgs; confirm the showcase grid renders all
|
|
primitives legibly and the benchmark still completes (green exit).
|
|
2. Re-capture `joeylog.txt` on all four ports; regenerate `PERF.md` via
|
|
`tools/uber-perf-table`. Expect IIgs `fillRect 320x200` and `stagePresent`
|
|
to drop toward ~25-28 (SEI removal de-inflates fillRect).
|
|
3. `tools/diff-uber-hashes` IIgs-vs-{amiga,st,dos}: all ops must still match
|
|
(the C dirty-mark + memset changes are portable and must not alter pixels).
|
|
If any op diverges, the dirty-mark fold changed a present extent -- inspect.
|
|
4. Eyeball the circle (`jlDrawCircle`) output unchanged after the dcSavedCol
|
|
change; confirm `jlDrawPixel` ops/sec rose.
|
|
|
|
|
|
|
|
## Ground truth / framing
|
|
|
|
- **Hard floor:** any full-screen 32000-byte move = 16000 word-stores x 6 cyc
|
|
(`STA long,X $9F` or `PEI`, both 3 cyc/byte) / 2.8 MHz = ~34 ms = **~29 ops/sec
|
|
max**. Clear, full-width fillRect, and present are all pinned to this; you
|
|
cannot beat it without moving fewer bytes (the dirty-row system already does
|
|
that for sparse frames).
|
|
- **The benchmark lies for SEI ops.** `PERF.md` `fillRect 320x200`=60 and
|
|
`stagePresent`=42 exceed/approach the 29-ops/sec floor and are inflated; clear=28
|
|
is honest. Fix measurement FIRST (Phase 0) or you cannot tell wins from clock noise.
|
|
- **Two architectural realities that are NOT bugs:** (a) the back-buffer ->
|
|
`$E1` design pays the screen twice on full-redraw frames, but the single-buffer
|
|
"fix" is refuted (slow `$E1`, overdraw inverts the win, tearing) -- keep it; the
|
|
dirty-row system is the right mitigation. (b) The circle OUTLINE trails the ST
|
|
due to 65816 register pressure and may never fully win; bank the partial gain.
|
|
|
|
## Cross-cutting rules (apply to every asm change)
|
|
|
|
- Keep the assembler's register-width state correct after every `REP`/`SEP` so
|
|
immediates are sized correctly.
|
|
- Keep each asm routine in the SAME load segment as its C caller (DRAWPRIMS for
|
|
the draw primitives); a separate segment silently produces calls-to-nowhere.
|
|
- `gStageMinWord` / `gStageMaxWord` must stay `uint8_t[]` so the asm that walks
|
|
them per byte stays correctly aligned with the C-side element size.
|
|
- PHA/PEI stack-slam paths REQUIRE `SEI` and are unsafe during `joeyInit` -- gate
|
|
them post-init (`gModeSet`).
|
|
- 4bpp word stores need even byte index (`x/2` even, i.e. `x % 4 == 0`) for the
|
|
computed-jump slam entry.
|
|
- Build hygiene: `source toolchains/env.sh`; DELETE the example binary before
|
|
rebuild (make leaves IIgs binaries stale); pass `spriteEmitIigs.c` first in
|
|
LIB_SRCS (linker bank-packing order).
|
|
- **Validation after every change:** re-run `examples/uber`, confirm
|
|
`jlSurfaceHash` per op is UNCHANGED (pixel-identical to the IIgs golden) and
|
|
ops/sec moved the right way. Verify visually in MAME (checkpoint poll or
|
|
snapshot) -- joeyLog alone misses wrong-pixel-but-right-control-flow bugs.
|
|
|
|
---
|
|
|
|
## Phase 0 -- Make measurement trustworthy (PREREQUISITE)
|
|
|
|
Everything downstream is judged against the benchmark; fix it first.
|
|
|
|
- **0.1 Audit & likely remove `iigsFillRectInner`'s `SEI`** (joeyDraw.asm:315,
|
|
chunk-CLI machinery :585-614). The body does only long-mode writes to the
|
|
bank-$01 back buffer with no SP/shadow hijack (per its own comment :284-292),
|
|
so the SEI appears vestigial. Confirm nothing in the fill needs interrupt
|
|
atomicity (MVN is interruptible/resumable by design), then remove SEI + the
|
|
FRI_CHUNK_ROWS CLI logic. Expected: fillRect-full de-inflates 60 -> ~27
|
|
(matching clear), AND the audio-starvation risk goes away. If some atomicity IS
|
|
required, document exactly what and keep it.
|
|
- **0.2 Add a cycle-ceiling sanity check / annotation.** Document the ~29 ops/sec
|
|
full-screen floor in `PERF.md`; flag any full-screen op reading above it as
|
|
suspect. Optional: assert in the harness.
|
|
- **0.3 Reconcile `present`=42** (its SEI is REAL -- stack/shadow hijack -- so it
|
|
can't just be removed). Capture an honest present rate via an SEI-immune time
|
|
source or a diagnostic no-SEI measurement build; expect ~25 ops/sec. Update the
|
|
table and the cross-port comparison accordingly.
|
|
- **Re-baseline `PERF.md`** with honest numbers before Phase 2+.
|
|
|
|
Exit criterion: clear, fillRect-full, and present all read within ~10% of each
|
|
other (~25-28 ops/sec), consistent with the cycle model.
|
|
|
|
## Phase 1 -- Make the uber demo legible (show, don't smear)
|
|
|
|
Right now `examples/uber` shows the user a solid colored screen and teaches them
|
|
nothing. Cause, concretely:
|
|
- `runAllTests()` (uber.c:222-285) runs each op in a 16-frame tight loop and
|
|
NEVER calls `jlStagePresent()` during or between the timed ops. The screen is
|
|
frozen on the last present before the run (the red-bar-on-black at uber.c:379)
|
|
for the whole ~80 s run, then ends on a bare `jlSurfaceClear(gStage, 2)`
|
|
(uber.c:390) -> solid green.
|
|
- Even if it did present, the timed ops are full-screen or fixed-coordinate
|
|
(`fillRect 320x200`, `surfaceClear`, pixel@100,100, line 0,50->319,50, ...), so
|
|
they overwrite each other into one solid color. There is no spatial separation.
|
|
|
|
Goal: the demo should visibly demonstrate what each primitive does AND run the
|
|
benchmark without leaving a meaningless screen. Separate "show the user" from
|
|
"measure throughput" -- they are different jobs.
|
|
|
|
- **1.1 Visual showcase pass (new), runs once before the benchmark, UNTIMED.**
|
|
Lay out a labeled grid (e.g. 4 cols x 5 rows of ~72x36 cells with a gutter) and
|
|
render each primitive ONCE into its own cell so nothing overlaps: pixel, line
|
|
H/V/diag, rect outline, circle (small), fillRect (small), fillCircle, a tile,
|
|
tileCopyMasked result, the compiled sprite, a 16-entry palette ramp swatch, an
|
|
SCB per-line strip, a sample-pixel readout. Full-screen ops (surfaceClear,
|
|
fillRect 320x200, present) cannot be shown at native size in a cell -- represent
|
|
each with a small labeled swatch and say so. `jlStagePresent()` once, then hold
|
|
for a readable beat (frame-counted ~2-3 s, or `jlWaitForAnyKey` -> "press a key
|
|
to begin benchmark"). THIS is the part that makes sense to a user.
|
|
- **1.2 Benchmark pass: keep timing/hash fidelity, show a sane status screen.**
|
|
Keep the `timeOp` loop and every op's args + call order EXACTLY as-is -- the
|
|
cross-port `jlSurfaceHash` golden values depend on the precise op sequence, so
|
|
changing them invalidates the validation harness. Stop presenting the timing
|
|
garbage: draw a static "BENCHMARKING..." screen plus a progress indicator (op
|
|
counter / bar, optionally each op's just-measured ops/sec) updated BETWEEN
|
|
`timeOp` calls -- never inside the timed loop. End on an explicit "DONE -- press
|
|
any key" screen, not a bare solid clear.
|
|
- **1.3 Labels.** `jlDrawText` exists (tile.h:119) but needs a font surface +
|
|
asciiMap, which uber does not currently load. Either add a small font surface +
|
|
asciiMap and label each cell, or (fallback) use fixed spatial layout + a short
|
|
on-screen color-key strip and a legend printed via `jlLogF`.
|
|
|
|
Cautions:
|
|
- All showcase drawing AND the progress-UI updates must be OUTSIDE any timed
|
|
window -- zero effect on ops/sec.
|
|
- Do NOT change the timed ops' arguments or sequence (golden-hash stability).
|
|
- Keep it port-agnostic: uber is shared example code that also runs on
|
|
Amiga/ST/DOS -- no IIgs-only assumptions in the showcase.
|
|
- Do this AFTER Phase 0 so any per-op ops/sec shown on the status screen is the
|
|
corrected (honest) number, not the inflated one.
|
|
|
|
Exit criterion: a first-time user watching the demo can see and identify each
|
|
primitive's output, and the benchmark phase shows readable progress + a clear
|
|
finish -- never a bare solid color.
|
|
|
|
## Phase 2 -- Kill the per-primitive C+JSL tax (biggest real-world win)
|
|
|
|
`jlDrawPixel` costs **~1595 cyc/call** (2.8e6 / 1755), ~94% overhead; the actual
|
|
nibble RMW is ~30 cyc. Every marking primitive pays a second cross-segment JSL.
|
|
|
|
- **2.1 Fold dirty-marking into the asm inner loops.** `jlDrawPixel` (draw.c:95),
|
|
`jlFillRect` (draw.c:448-ish via fillRectOnSurface), line, and circle each call
|
|
`surfaceMarkDirtyRect` AFTER the plot JSL -- a second JSL into
|
|
`iigsMarkDirtyRowsInner` (~49 cyc plumbing x2 + the row loop). The inner loops
|
|
already know the row(s) and extent (fillRect has `midStart`/`trailingByte` per
|
|
row; pixel has `y` and the byte). Update `gStageMinWord[y]`/`gStageMaxWord[y]`
|
|
inline and make `surfaceMarkDirtyRect` a no-op on the IIgs fast paths.
|
|
- Confirmed win: removes ~120 cyc from EVERY marking primitive; **5-10x on
|
|
`jlDrawPixel`**.
|
|
- **2.2 Trim the asm prologue/epilogue for the tiny ops.** `iigsDrawPixelInner`
|
|
(joeyDraw.asm:1158-1227) keeps scratch in 6-cyc long-absolute (`dpxlTmp`,
|
|
`dpxlNibPart`) and pays full PHP/PHB/PHD ... PLD/PLB/PLP framing around a ~50-cyc
|
|
plot. Use registers/DP for the scratch; drop framing not needed when DBR/D are
|
|
already correct. (est ~1.2-1.4x on the inner once 2.1 lands.)
|
|
- **2.3 Batch/inline entry for pixel-heavy callers (optional).** A span/row-run
|
|
API (or inlining the in-bounds nibble RMW into the C via macro, the way fillRect
|
|
was de-wrappered) amortizes the JSL across many pixels for plotting-heavy code.
|
|
|
|
Exit criterion: `jlDrawPixel` >> 1755; line/rect call cost drops by the second-JSL
|
|
amount; hashes unchanged.
|
|
|
|
## Phase 3 -- fillRect partial-width slam gate
|
|
|
|
The 3-cyc/byte slam fires ONLY for exactly `x==0 && w==320` (`friFullWidth`,
|
|
joeyDraw.asm:406-420). Everything else falls to seed+`MVN` at **7 cyc/byte**
|
|
(2.33x slower), confirmed.
|
|
|
|
- **3.1 Computed jump into the unrolled `STA long,X` table.** Set `X =
|
|
curRow + midStart`, then `JMP (addr)` into the 80-store run at store
|
|
`#(80 - midBytes)`. Gives 3 cyc/byte for ANY width. Keep MVN only for tiny odd
|
|
remainders. Drop-in safe (bank-$01 long stores, no SEI/shadow). Respect the
|
|
4bpp even-byte-index rule for the entry. **1.8-2.3x on all partial-width fills.**
|
|
- **3.2 Cleanups (minor):** cache `curRow+midStart` instead of recomputing it
|
|
(joeyDraw.asm:542-565); hoist loop-invariant edge offsets (`leadingByte`/
|
|
`trailingByte` relative to curRow, += 160/row); one `sep`/`rep` window per row
|
|
for both edge RMWs.
|
|
|
|
## Phase 4 -- circle / line register pressure (~1.27x, partial)
|
|
|
|
- **4.1 Move Bresenham scratch to direct page.** `dcX/dcY/dcErr/dcRowYP/dcRowYN/
|
|
dcRowXP/dcRowXN` (circle, joeyDraw.asm:1500+) and the line's per-pixel state
|
|
(joeyDraw.asm:1253+) are 6-cyc long-absolute; DP is 5 cyc here (DL != 0 because
|
|
D=SP+8). Set up a DP scratch page for the hot state.
|
|
- **4.2 Remove the per-plot parity round-trip.** `sta dcSavedCol` / reload /
|
|
`and #1` (joeyDraw.asm:1603-1610) recovers a bit `lsr a` already left in carry
|
|
-- branch on carry instead. ~15 cyc/plot x 8 plots/step.
|
|
- **4.3 Maintain the 4 row bases incrementally** (+/-160 per Bresenham step)
|
|
instead of 4 LUT lookups/step.
|
|
- Net ~1.27x (0.65x -> ~0.83x vs ST). Stop there -- the architecture won't give a
|
|
full win on the outline. The fill-circle already beats the ST; leave it.
|
|
|
|
## Phase 5 -- faster clear (post-Phase-0; lower priority)
|
|
|
|
- **5.1 Post-init PHA-slam clear path.** `STA long,X` = 3 cyc/byte; SP-into-`$01`
|
|
PHA-slam = 2 cyc/byte. ~1.5x (37.6 -> ~24 ms). Gate on `gModeSet` (SEI unsafe at
|
|
init); keep the `STA long,X` path as the init-safe fallback. Chunk the SEI like
|
|
present to avoid audio starvation. NOTE: only meaningful to measure once Phase 0
|
|
fixes the clock.
|
|
- **5.2 Route `surfaceMarkDirtyAll` through `iigsMarkDirtyRowsInner`** (it's a
|
|
200-iter C loop today; surface.c:233-243) -- one source of truth with the
|
|
fillRect marker. ~7600 cyc/clear.
|
|
- **5.3 Hold `fillWord` in a register** across the clear loop instead of
|
|
`lda 9,s` per iteration (joeyDraw.asm:137). Subsumed if 5.1 lands.
|
|
|
|
## Phase 6 -- sprite dispatch overhead
|
|
|
|
Body store mode (`LDA #imm16`/`STA abs,Y`) is already near-optimal -- DO NOT touch
|
|
it. ~3000-4000 cyc/call of C dispatch wraps ~1000-1500 cyc of real blit.
|
|
|
|
- **6.1 Share address/shift/fnAddr across `SaveAndDraw`'s two halves** (sprite.c)
|
|
-- it recomputes dest address, shift, fnAddr, and re-patches the self-modifying
|
|
stub twice. (est ~1.05x SaveAndDraw.)
|
|
- **6.2 Trim per-call C dispatch:** cache fnAddr, dodge the 2D-index MUL4, do
|
|
`isFullyOnSurface` in 8-bit. (est ~1.10-1.15x per sprite op.)
|
|
|
|
## Phase 7 -- architecture (decision, no code)
|
|
|
|
Keep the back-buffer + dirty-row model. Document in `docs/DESIGN.md` that
|
|
full-redraw frames inherently pay the screen cost ~2x and that sparse-update
|
|
(sprite/HUD) games should lean on partial updates. Do NOT pursue single-buffer
|
|
direct-to-`$E1` (refuted on slow-bank/overdraw/tearing grounds).
|
|
|
|
---
|
|
|
|
## Suggested order of execution
|
|
|
|
1. Phase 0 (0.1 fillRect SEI, 0.2 annotate, 0.3 present) -> re-baseline PERF.md.
|
|
2. Phase 1 (make the demo legible) -- so every later change can be eyeballed, not
|
|
just hashed; cheap and unblocks visual validation.
|
|
3. Phase 2.1 (fold dirty-marking) -- highest real-world payoff, touches all primitives.
|
|
4. Phase 3.1 (fillRect computed-jump slam).
|
|
5. Phase 2.2, Phase 4 (circle/line DP scratch + parity).
|
|
6. Phase 5, Phase 6 (clear PHA-slam, sprite dispatch) as time permits.
|
|
7. Phase 7 doc.
|
|
|
|
Each step: re-run uber, assert hashes identical, record ops/sec delta in PERF.md.
|