78 KiB
JoeyLib cross-port performance (UBER benchmark)
Regenerated 2026-07-05 (PERF-AUDIT-PLAN.md Phase 8, #25b) from same-day captures of all four ports. This supersedes the 2026-07-04 Phase 0 baseline; every phase of the audit fix plan (0-8) has landed, so these are the numbers the library actually ships with.
2026-08-06 update -- Sharp X68000 moved to the 256-colour graphics
plane. The port previously drew native-planar into the TEXT plane, which has a
single 16-entry hardware palette and therefore could not honour the SCB. To make
SCB per-band palettes work it now presents into the 256-colour GRAPHICS plane
(_iocs_crtmod(13)), expanding the four 1bpp planes to one byte per pixel in
GVRAM with the DOS port's encoding, (scb[y] << 4) | nibble, against the
256-entry palette at $E82000. Storage is UNCHANGED -- still
JOEYLIB_NATIVE_PLANAR; only the present target moved. No raster interrupt is
involved. All 51 golden hashes are unchanged (jlSurfaceHash hashes the surface
planes, not GVRAM) and the port stays at 51/51. (The "total wall time is
identical at 15945 ms" originally written here was an artifact of the broken
frame clock -- see the clock note at the end of this entry.)
Measured 2026-08-06 on the passing 256-colour run with both clocks fixed. Honest UBER total wall time is 40890 ms (2249 frames). Two earlier figures in this entry's history were wrong and are corrected below: 15945 ms came from a broken frame clock, and 108309 ms was real but predated the present rewrite.
The "pre-LUT" column is the straightforward per-pixel expansion the graphics-plane switch first shipped with; "now" is after the word-store expansion tables.
| op | pre-LUT | now | IIgs | now / IIgs |
|---|---|---|---|---|
jlSurfaceClear |
30 | 106 | 33 | 3.21x |
jlPaletteSet |
16175 | 58516 | 1561 | 37.49x |
jlScbSetRange |
5946 | 21206 | 163 | 130.10x |
jlDrawPixel |
10225 | 36220 | 5576 | 6.50x |
jlDrawLine H |
2632 | 9370 | 1322 | 7.09x |
jlFillRect 16x16 |
687 | 2414 | 785 | 3.08x |
jlFillRect 80x80 |
84 | 306 | 133 | 2.30x |
jlFillRect 320x200 |
29 | 103 | 22 | 4.68x |
jlDrawCircle r=16 |
165 | 591 | 303 | 1.95x |
jlDrawCircle r=80 |
35 | 125 | 70 | 1.79x |
jlSpriteDraw |
2059 | 7397 | 1142 | 6.48x |
jlSpriteSaveUnder |
3570 | 12803 | 602 | 21.27x |
jlTileMapPaste 10x5 |
295 | 1055 | 140 | 7.54x |
jlStagePresent full |
392 | 1373 | 73 | 18.81x |
gameFrame composite |
2 | 12 | 12 | 1.00x |
gameFrameClean composite |
2 | 137 | 20 | 6.85x |
All 39 shared ops now meet or exceed the Apple IIgs reference; median 7.54x,
slower on none. The two composite rows are the ones that mattered -- they are
the closest thing UBER has to a real game frame, and they were the reason the
port missed the "within ~1.5x of the native ceiling" directive at 0.17x and
0.10x. gameFrame composite is now exactly at parity and gameFrameClean is
6.85x.
The fix was NOT a port of DOS's gExpandLut. DOS fuses two ADJACENT output bytes
into one 16-bit store (mode 13h pixels are neighbouring bytes); here each pixel
is a 16-bit WORD, so its bytes are two apart and that fusion is impossible.
Inverting it works: write whole words, and one 32-bit store covers exactly two
pixels. The high byte is free because the 256-colour index reads only
$000F | $00F0 (x68k_v.cpp:339-340). Expansion tables indexed by plane and
nibble are pre-interleaved into word positions, so four planes OR together with a
replicated palBase and the inner loop issues four long stores per eight pixels
with no shifting at all -- against roughly 64 operations and eight byte-stores
before. Verified three ways, because each layer misses what the others catch:
51/51 golden (surface correct), a GVRAM dump under PATTERN.X byte-identical to
the pre-rewrite run (pixel encoding correct), and an earlier framebuffer capture
(the display actually renders the bands).
Two clock bugs had to be fixed before any of these numbers meant anything, and
every X68000 figure published before 2026-08-06 was wrong. The port had no
JL_HAS_MILLIS_ELAPSED override, so milliseconds derived from the frame counter;
and that counter was POLLED (vdispPoll), so an op spanning more than a frame
lost the V-DISP edges in between. The undercount scaled with the op's cost and
cancelled it, making every op report ~10,000 ops/sec regardless of workload --
jlFillRect 320x200 (9735) matched jlFillRect 16x16 (10505) despite 250x the
work. The fix chains Human68k's MFP Timer C, reading its rate from TCDCR/TCDR
rather than assuming it, and derives the frame counter from that tick.
UBER-CLK now agrees frame-vs-millisecond on every row, which it could not
before because both figures came from the one broken source.
2026-07-06 update: the IIgs sprite column is now COMPILED. Finding #77 (the loader/segment-BSS + malloc-heap reservation that forced the codegen arena into the app's own BSS) was fixed toolchain-side and JOEY_IIGS_SPRITE_CODEGEN was re-enabled; the four IIgs sprite rows were re-captured (bench-iigs.sh) and jumped 5-25x (SaveUnder 104->543, Draw 48->907, RestoreUnder 110->568, SaveAndDraw 15->376 ops/sec), bit-identical to the interpreter goldens (diff-uber-hashes 45/45). The ST and DOS columns are unchanged from the 2026-07-05 capture.
Also 2026-07-06: the Amiga jlpFillRect partial-width path was moved
onto the blitter (BLTAFWM/BLTALWM hardware edge masks + masked minterm),
which lifted jlFillRect 16x16 567->1153, 80x80 71->269 (both from
below-floor to well above it), and -- because they build on fillRect --
jlDrawLine V 144->269 and jlDrawRect 100x100 122->203, all
hash-identical (Amiga 45/45). The one remaining Amiga below-floor cell
is jlFillCircle r=40 (74%); a whole-circle asm routine was tried and
did NOT help (per-span dispatch is not the cost -- the 8-way-symmetry
overlap fill dominates), so it stays on the C span walker for now.
2026-07-06, NATIVE-PERF Phase 0: four WORKLOAD rows were added (see
NATIVE-PERF-PLAN.md) and the goldens refrozen at 49 lines after a
clean regression gate (45 match + exactly 4 extras per port) and full
4-port cross-identity. The new rows measure what games actually do, and
they expose the real problem the per-op rows hid: jlSpriteDraw unaligned runs at 5% (ST) / 9% (Amiga) of the IIgs floor -- a
15x/12x cliff vs the aligned row, because only x%8==0 draws are
compiled on the 68k ports (see the plan's code-truth table). The
gameFrame composite row is the bottom line: ST 9 / Amiga 18 /
IIgs 10 / DOS 18 frames/sec for a modest game frame. Bonus finding:
cross-identity of the sweep16 row validated the ST/Amiga shifted c2p
walkers as pixel-perfect at all 14 previously-untested x phases.
(Re-capture note: IIgs input/audio rows wobbled vs the prior MAME run
-- e.g. jlAudioFrameTick 50223 -> 27612 -- run-to-run MAME variance on
cheap ops; hashes are unaffected and those rows gate nothing.)
2026-07-06, NATIVE-PERF Phase 1 (draw side) LANDED: all-shift
compiled sprite DRAW on both 68k ports. Amiga emits pre-shifted JIT
variants for all 8 bit phases; the ST emits per-phase thunk + data
tables driven by a shared masked-movem walker (spriteWalk.s) for all
16 intra-group phases, with the dispatcher's ST shift contract changed
to x & 15. jlSpriteCompile is now plan-based: it measures every
(shift, op) variant, greedily fits them under the uint16 offset cap
(draws first, save/restore as atomic pairs), and degrades to the
byte-aligned phases when the arena is too small instead of losing the
sprite. RESULT -- the alignment cliff is CLOSED: jlSpriteDraw unaligned ST 40 -> 500 ops/sec (12.5x), Amiga 75 -> 803 (10.7x);
sweep16 ST 2 -> 32 sweeps/sec (512 draws/sec), Amiga 5 -> 50 (800
draws/sec). Unaligned now performs like aligned on both ports (ST 82% /
Amiga 84% of their aligned rows), and every one of the 49 golden
hashes still matches -- the compiled paths are bit-identical to the
interpreted walkers they replace, on all four ports. (Found along the
way: the Amiga bench task stack was ~4 KB -- one deep call chain from
overflow -- now 32 KB in bench-amiga.sh; the emitters' accumulators are
file statics.)
Phase 1 step 5 (save/restore) ALSO LANDED, completing Phase 1: compiled save/restore on both 68k ports operate on the 16-px-group window (two width classes per sprite, like the IIgs MVN pair), the backup format is a pure function of the recorded geometry (identical for compiled and interpreted producers/consumers -- the ST HAL gained a raw-span interpreted branch, switched atomically with restore), and the gates route by SPRITE_SAVE_CLASS / SPRITE_RESTORE_CLASS with an odd- buffer-pointer demotion (68000 word ops need even addresses). The combined jlSpriteSaveAndDraw path now fires on 68k at EVERY x for the first time. RESULTS (floor-relative): Amiga SaveUnder 104% -> 276%, RestoreUnder 98% -> 236%, SaveAndDraw 94% -> 160%; ST SaveUnder 158% -> 203%, RestoreUnder 122% -> 163%, SaveAndDraw 87% -> 107%; Amiga gameFrame 18 -> 23 fps. Backup capacity contract: JOEY_SPRITE_BACKUP_BYTES in sprite.h (window worst case; all in-repo examples migrated + 2-aligned). DEFAULT_CODEGEN_BYTES is 64 KB on Amiga/ST (full variant sets). All 49 golden hashes held throughout.
2026-07-07, THE ST COLUMN DOUBLED -- idle-audio ISR tax found and fixed: every previous ST capture ran with ~49.5% of the CPU silently consumed by the 12.3 kHz Timer-A PWM ISR that jlpAudioInit started unconditionally -- it streamed buffer silence to the YM even with nothing playing, and UBER inits audio before its timed run. Found by a calibrated probe (known-cycle dbra spin + per-layer direct-call rows under the bench Hatari config): every probe row slowed by exactly 1.98x once audio was initialized, and the probe's un-taxed tilePaste cycles/call matched the instruction-count model that UBER's number mysteriously doubled. Fix: src/atarist/audio.c Timer A is now demand-driven -- installed on the first PlayMod/PlaySfx, parked again once the mix buffer drains back to silence with nothing live (drain-to-silence handshake in jlpAudioFrameTick; StopMod parks immediately when no SFX is riding the pump). Probe-verified: idle rows match the never-initialized rows exactly, the tax correctly returns while an SFX drains, and it vanishes after the drain. Full ST re-bench: 49/49 hashes bit-identical (pixels untouched) and EVERY row ~2x: gameFrame 9 -> 19 fps, tilePaste 1,574 -> 3,203/s, sprite draw 607 -> 1,188/s, unaligned draw 508 -> 992/s. The other three ports were audited for the same bug class and are clean: Amiga ~1.8% idle tax (PTPlayer rides Paula DMA), DOS ~0.5% (SB auto-init DMA, ~11 Hz half-buffer IRQ), IIgs 0% (no handler installed until PlayMod). ALL pre-2026-07-07 ST comparisons -- including the Phase-0 ceiling gaps and the Phase-1 relative results -- understated the ST by ~2x.
2026-07-07, NATIVE-PERF PHASE 4 (IIgs arm) LANDED -- whole-map asm walker: iigsTileMapPasteMark (joeyDraw.s) replaces the IIgs stage-arm per-tile loop of jlTileMapPaste with ONE JSL that walks the index map: dst offset advances incrementally (+4/column, +1280/tile row) from a single gRowOffsetLut read, the register-starved inner lives in a 22-byte stacked D-frame (X survives the 16-store body and carries the dst offset; the map index sits in the frame because Y is clobbered), the idx << 5 tileset add carries into the bank byte, all three wrapping branches are brl, and ONE band union closes the call. Gated 50/50 + make iigs-verify. RESULT: IIgs batch row 60 -> 140 maps/s = 7,000 tiles/s (+133%), ~400 cyc/tile -- BELOW the 480-cyc single-tile hand-coded ceiling, because the walker amortizes the JSL/derive/mark that even a hand-coded per-tile call pays. Batch row across all ports: ST 8,150 / Amiga 7,650 / IIgs 7,000 / DOS 22,500 tiles/sec (the morning's per-call loop numbers were 1,600 / 2,300 / 1,750 / 10,000). The Amiga blitter cookie-cut sprite was recon'd and DEFERRED with a full design on record: the synchronous WaitBlit the hash contract requires erases the parallelism win at <= 16x16 (every current sprite row), and the chip-RAM pre-shifted source data is a net-new subsystem -- the honest experiment is a >= 32x32 golden'd row measured with and without the fence.
2026-07-07, NATIVE-PERF PHASE 4 CLOSED -- clean-buffer experiment
PASSED, jlSurfaceCopyRect is public API: Scott signed off the
jlTileMapPaste shape and approved the clean-buffer experiment; both
landed same-day. jlSurfaceCopyRect(dst, src, x, y, w, h) is a
SAME-POSITION rect copy (clip, then snap outward to the 16-px group
grid on every port -- uniform snap keeps cross-port hashes
byte-identical): core wrapper + dirty mark in surface.c, generic
chunky row-memcpy engine (serves DOS and both IIgs surface flavors --
the IIgs stage pixels pointer reaches bank $01 through ordinary
large-model stores, recon-verified against the sprite-restore
precedent), ST override = strided interleaved row memcpys (the
raw-span restore recipe), Amiga override = per-plane CPU row copies
(CopyMemQuick is unusable at 16-px grain: word alignment only;
the modulo-driven blitter variant is the recorded production
upgrade). The STRATEGY row gameFrameClean composite (erase 3
sprite windows from a pristine clean surface + 3 draws + present; NO
per-frame background repaint, NO save-under -- the static-background
pattern, not an op-for-op gameFrame speedup) measures, vs gameFrame:
ST 48 vs 22 = 2.18x (acceptance bar was 1.4x), DOS 195 vs 19 =
10.3x, IIgs 18 vs 10 = 1.8x, Amiga 34 vs 25 = 1.36x (Amiga gains
least, exactly as the adversarial recon predicted: its fillRect was
already blitter-fast, and clean->stage is a chip-to-chip CPU copy
against display DMA -- the blitter rect variant is its lever).
GOLDEN CONTRACT now 51 lines: regression gate exactly 50 match +
1 extra per port, 4-port cross-identity of the new row (hash
C51FE171 x4), then freeze. Found along the way: make all does NOT
regenerate joey.2mg, so bench-iigs.sh silently benched the previous
UBER (clean 50/50, new row absent) until make iigs-disk -- recorded
in the stale-binary memory.
2026-07-07, NATIVE-PERF PHASE 4 (ST arm) LANDED -- movep.l map
engine: the ST jlTileMapPaste arm now dispatches to
stTileMapPasteMovep (src/atarist/tileMap.s): per tile row, one
move.l (a0)+ + movep.l d0,d16(a1) writes all 4 interleaved plane
bytes in 36 cycles vs ~80 for the four compiled move.b -- movep IS
the half-group byte-scatter pattern (the classic ST c2p trick), where
movem is useless (byte-granular dst) and byte-pair word fusion loses
(lsl.w #8 costs 22 cyc on a barrel-shifterless 68000, priced during
recon). Cells walk in even/odd pairs sharing one group pointer
(+8/pair, all row offsets compile-time d16); odd bx = head cell, odd
width = tail cell. The C pair walker (stTileMapPaste in stTile.h,
+17% alone -- adversarial recon correctly predicted the pure-C cap at
+15-23% because gcc already hoists the mulu and the 632-cyc copy body
dominates) stays as the odd-tileset-pointer fallback: move.l (a0)+
needs an even source, jlTileT is align-1, and the guard is one AND
per call. RESULT: ST batch row 163 -> 346 ops/sec = 17,300
tiles/s (2.12x, 462 cyc/tile) -- ABOVE the Phase-0 12,700-15,500
"hand-coded ceiling", which under-counted by assuming byte-move
bodies. Gate: ST 50/50 bit-identical (hash 52B9646B held through
both engines) + a 12-case on-emulator proof (TILETEST, session
scratchpad) covering odd-bx head, odd-width tail, w=1 columns, h=1,
checkerboard/end/all skips, and both odd-pointer fallback paths,
each with a border-ring containment check. Batch row across ports:
DOS 22,500 / ST 17,300 / Amiga 7,650 / IIgs 7,000 tiles/s. IIgs
PEI-slam map variant was EVALUATED and REJECTED: models ~+28%
(~312 cyc/tile) but requires SEI for the whole map (~5.6 ms
interrupts-off at 10x5, audio ISR jitter) AND the VBL-driven bench
clock cannot honestly measure SEI ops (PERF artifact: SEI inflation)
-- the walker already beats the 480-cyc per-tile ceiling without it.
2026-07-07, NATIVE-PERF PHASE 3 LANDED -- jlTileMapPaste batch API:
one public entry pastes a wTiles x hTiles rectangle from a tileset +
row-major uint8_t index map (TILE_MAP_SKIP = 0xFF leaves a cell
untouched): one validation, one dirty union (68k/DOS; the IIgs
stage-fused entries mark per pasted tile, and its loop hoists the
stage test + does tileset indexing as base + (idx << 5) -- the
natural tiles[idx] spelling cost a 32-bit indexed multiply per cell
on the 65816 and erased the batch win until specialized). GOLDEN
CONTRACT now 50 lines: refreeze followed the Phase-0 procedure
(regression exactly 49 match + 1 extra per port, then 4-port
cross-identity of the new jlTileMapPaste 10x5 row -- all four hash
52B9646B -- then freeze). RESULTS, tiles/sec batch vs per-call loop:
ST 8,150 vs 4,000 (+106%, gap-to-ceiling 1.6-1.9x); Amiga 7,650 vs
3,750 (+104%, gap 1.40x -- INSIDE the ~1.5x program target); DOS
22,500 vs 12,900 (+74%, 2.5x); IIgs 3,000 vs 2,950 (+2% -- its
Phase-2 fused entries already deleted the per-tile wrapper, so the
remaining cost is ~93% inner JSL; a whole-map IIgs asm walker is the
Phase-4 lever). Sprite batch APIs were DESCOPED from Phase 3: the
phase gate measures only the tilemap row, and post-R5 sprite dispatch
overhead is ~3-6% -- not worth new public surface without a measured
need.
2026-07-07, NATIVE-PERF PHASE 2 COMPLETE -- Steps D3 (R6 + R10) close the phase: (R6) the IIgs masked-copy inner's 4 x 8 jsr/rts round-trips to tmaskByte were inlined via the TMASK_BYTE assembler macro (identical body; the row loop-back became brl -- the expanded body exceeds 8-bit branch range): jlTileCopyMasked 625 -> 681 (+9%, exactly the 384-cycle model). (R10) the stage arm of jlDrawPixel is one fused JSL (iigsDrawPixelMark: single-row widen + nibble RMW plot against the pinned stage base): jlDrawPixel 3,179 -> 5,576 (+75%, 502 cyc/call incl. wrapper vs the measured 880 before); the batching plotters (lines/circles, plotPixelNoMark) keep the unfused inner and were flat, as were tiles/sprites. Final gate 49/49 x4 + make iigs-verify. PHASE-2 EXIT vs its gate ("small-op rows move measurably toward Phase-0 ceilings on all ports; goldens hold"): tilemap row tiles/sec ST 1,600 -> 4,050 / Amiga 2,300 -> 3,750 / IIgs 1,750 -> 2,950 / DOS 10,000 -> 12,900; gameFrame ST 9 -> 21 / Amiga 18 -> 25 / IIgs 10 / DOS 18 -> 19; IIgs tilePaste ceiling gap 3.5x -> 2.15x, ST 2.6x. The remaining per-call floor is the public wrapper + one JSL/call -- Phase 3 batching territory. Optional leftovers R7 (saveAndDraw consolidation) and R12 (IIgs C-residue reshapes) are recorded in the plan, not phase-gating.
2026-07-07, NATIVE-PERF Phase 2 Step D2 LANDED -- sprite dispatch
flattening (R5): the spriteCompiled* runtime dispatchers moved from
spriteCompile.c into spriteDispatch.h as always-inline functions
consumed only by sprite.c, deleting the cross-TU call per compiled
sprite op and letting the compiler share the slot/field loads the
caller already did. Gated 49/49 x4. Sprite rows: IIgs +6-13%
(SaveAndDraw 376 -> 423, Draw 906 -> 966), ST SaveUnder +15% (2,611),
Amiga +5-11% (SaveUnder 2,053), DOS SaveUnder +16% (10,743). Two
measurement footnotes: gameFrame wobbles +/-1 fps between runs
(single-digit iters per 16-frame window -- quantization, not signal;
ST has bounced 21/22 across three consecutive gates), and DOS
jlSpriteDraw unaligned is BIMODAL across runs of identical binaries
(~3,800 vs ~5,100; seen 07-06 and again today; the all-phase sweep16
row is stable at ~362, so per-phase behavior is fine -- likely DOSBox
dynamic-core state; do not chase single-run deltas on that row).
2026-07-07, NATIVE-PERF Phase 2 Step D1 LANDED -- IIgs fused tile entries (R2): jlTileFill/Copy/Paste on the stage now run ONE JSL each: new joeyDraw.s entries (iigsTile*Mark, section .text.iigsFusedTile) derive the destination in asm from the pinned $01:2000 stage base + gRowOffsetLut, write via long,X with the stage literal folded into each store, and widen the dirty band via a shared stageTileMark helper -- replacing C-side pointer math plus two JSLs (op + marker). Gated 49/49 (pixels) AND make iigs-verify (framebuffer -- the mark half is hash-blind, so the display check is load-bearing). RESULTS: IIgs tileFill 1,981 -> 3,251 (+64%), tilePaste 1,687 -> 2,710 (+61%), tileCopy 1,441 -> 1,988 (+38%), map row 35 -> 59 (+69% = 2,950 tiles/s), gameFrame 10 -> 11 fps; unfused controls (copyMasked, snap, sprites) flat. IIgs tilePaste is now ~1,030 cyc/call vs the 480-cyc hand-coded ceiling (gap 2.15x, was 3.5x). Toolchain landmine found en route: a 16-bit jsr between .section blocks can cross OMF segment banks (no link error, wild jump into $00CFxx) -- helpers must share their callers' section or be jsl/rtl; see the cross-section-jsr feedback memory.
2026-07-07, NATIVE-PERF Phase 2 Steps A-E LANDED (per-op tax): five hash-gated steps, each 49/49 x4. (A) plane-pointer access inlined into the 68k compiled-sprite dispatchers (amigaPlanar.h / stPlanar.h; the jlpSurfacePlanePtr cross-TU call -- ~190 cyc x4 per Amiga dispatch -- deleted): Amiga sprite ops +10-23%. (B) whole-tile dirty marks specialized (band = [bx2, bx2+1] with compile-time proof, unrolled 8-row widen): tiles +9-28% on 68k/DOS. (C) jlDrawRect raw-edge fills + ONE band-identical mark (+22-40%); the multi-row widen unroll is Amiga-only by measurement (gcc-mint compiled it WORSE, ST -5% -- see the per-port #if in surface.c). (D=R9) presents consume the core gStageScbDirty/gStagePaletteDirty flags instead of memcmp'ing 512-712-byte cached mirrors: gameFrame +1-2 fps on ST/Amiga/DOS. (E) ST/Amiga tile ops moved into always-inline per-port headers (stTile.h / amigaTile.h), deleting the second cdecl layer (ST tiles +9-14%, Amiga +5-11%), then the Amiga row loops unrolled to d16(An) form (tilePaste +36%, tileSnap +52%). NET SESSION MOVEMENT on the tile gate row (tiles/sec, map 10x5): ST 1,600 -> 4,000; Amiga 2,300 -> 3,750; DOS 10,000 -> 12,900. gameFrame: ST 9 -> 21, Amiga 23 -> 25, DOS 18 -> 19 fps (IIgs untouched at 10 -- its fused-asm glue is Phase 2 Step D, still open). The present-row jumps (ST 467, DOS 1,137) are the dirty-union idle early-out running flag-gated; gameFrame is the honest present signal.
Correctness statement backing this table: all FIVE ports produce BYTE-IDENTICAL pixels on every measured op and every correctness check -- 51/51 surface hashes match the frozen goldens in tests/goldens/uber/ (39 timed ops + 12 hashed checks; the X68000 has no golden of its own and is diffed straight against iigs.txt; the PASS/FAIL checks -- palette round-trip, PRNG golden sequence, sprite clip round-trip, arena churn, sprite-from-surface -- pass everywhere). The Phase 0-era divergences (#75 tilePasteMono three-way, #78 ST fillCircle/tileFill, #87 IIgs circle outline) are all fixed.
2026-07-09, Phase 7 P7-3 -- Amiga present sped up (src/amiga/hal.c only). The Amiga present was Phase 6's single largest out-of-bar item (2.8-5.3x). A probe forensic split it: a fixed per-present scan (P8 idle = 1.51 ms) plus, for large presents, a ~11 ms CPU copy (P1 present96 copied 12,288 B). Two bisectable changes, each P11-gated (presentPageSum = exhaustive displayed-page-vs-shadow byte parity, the gate for the hash-invisible present -- PASS on both): (A) the per-present snapshot pass was folded into the scan's COLD union-dirty path (a row clean in both frames already holds the clean sentinel in gPrev, so the hot path stays 2 loads -- a first attempt that touched every row REGRESSED and re-confirmed the recorded gcc-amigaos register-pressure trap), dropping the present from two 200-row passes to one (scanned 400 -> 200); (B) the large-run copy now runs on the blitter via the W3 amigaBlitterCopyPlanes (A->D, offset+modulo) above a 128-plane-word threshold, so present96 / gameFrame's 96-row bands move off the CPU while present3win's scattered windows stay on the CPU THIN path. RESULTS: present96 46 -> 71 (+54%), present3win 93 -> 103 (+11%), gameFrame 25 -> 31 (+28%), gameFrameClean 62 -> 71, jlStagePresent full 584 -> 609; Amiga 51/51 golden hashes hold (present changes are hash-invisible -- pixels unchanged). Remaining Amiga-present lever: a fused-scan present asm (ST presentSlam analogue over 4 plane buffers) for the scan-bound rows -- C fusion is now empirically register-pressure-capped, and the async-fence blitter contract stays parked behind a >=32x32 row.
2026-07-09 -- fused-scan present asm + a component probe that located the real cost. src/amiga/presentScan.s (hand-rolled 68000, the 4 band arrays walked in address registers) replaces jlpPresent's C union scan. The plain asm (no fast path) barely moved present3win (103->106), which first read as "the scan isn't the bottleneck" -- but a present-component skip-mask probe corrected that. It split present3win's 9.4 ms: SCAN ~4.9 ms, FLUSH ~3.3 ms (chip->chip CPU copy of the 3 tiny windows), FLIP ~0.4 ms, residual (copper-check/reset/WaitBlit/marks/harness) ~1.2 ms; the 3 marks alone are 0.64 ms. So the scan IS ~half the present -- the plain asm just wasn't faster than gcc because the scan is INSTRUCTION-COUNT-bound (~16 instr x 200 rows), not the register spill it was assumed to be. The ST presentSlam 4-row fast path (one move.l compares 4 clean-sentinel min bytes, cur AND prev, and skips the group; needs gPrev arrays aligned(2)) LANDED: present3win 106->128 (+21%), present96 ->78, gameFrame 32->34, gameFrameClean 62->82, P11 PASS, 51/51. A first cut passed all PIXEL gates but broke the gameFrameClean golden row -- a bisect proved the fast path's 'bra .rowLoop' skipped the loop-exit check, so a clean group ending at row 199 landed the row index on 200 and processed a PHANTOM out-of-bounds row 200 (writing gPrevStageMin/MaxWord[200] one byte past the arrays). That corrupts a neighbouring global invisibly to P11/hashes/counters yet breaks a later jlSurfaceCreate AllocMem -- gameFrameClean (the only op allocating a 2nd surface) was an accidental memory-integrity canary. FIX: the exit check now lives at the TOP of .rowLoop, guarding both re-entry paths. The FLUSH lever also LANDED: the ~3.3 ms chip->chip CPU copy of the small windows was contention-bound, so lowering AMIGA_PRESENT_BLIT_MIN_WORDS 128 -> 16 routes them to the blitter (parallel chip->chip DMA) -- present3win 128 -> 153 (+20%), gameFrameClean 82 -> 90, 51/51 + P11. An OWN-ONCE flush polish (the present OwnBlitters once instead of per window) also landed correct (51/51 + P11) but PERF-FLAT (present3win 153 -> 154, within noise): after the threshold drop the flush is blitter-DMA-bound, so the per-window Own/Disown it removes was already negligible -- kept for the right blitter discipline, not a speedup. Session present3win arc 106 -> 128 -> 153 -> 154 = +45%. No known remaining flush lever.
2026-08-05, FIFTH PORT: Sharp X68000 -- the hashes PASS, the clock does NOT. src/x68000 + make/x68000.mk (JOEYLIB_PLATFORM_X68000) is on the golden gate: tools/diff-uber-hashes reports OK: 51/51 ops match. The port carries NO frozen golden of its own -- it is diffed straight against tests/goldens/uber/iigs.txt, which it matches exactly. Shape: 10 MHz 68000, NATIVE PLANAR storage (4 x 1bpp planes, 40 bytes/row, X68K_PLANE_SIZE 8,000), byte-for-byte the Amiga surface layout -- but the DISPLAY is not the surface. It draws into the TEXT plane (four 1bpp planes at $E00000/$E20000/$E40000/$E60000, 128-byte stride), NOT graphics VRAM, so jlpPresent scatters the compact 40-byte rows into that stride: 32,000 bytes/frame against 128,000 bytes of GVRAM word writes on the chunky path it replaced (src/x68000/x68kPlanar.h:20-33). The 320x200 window is CENTRED in a 512x512 mode (_iocs_crtmod 12) at x=96 / y=156, byte-aligned by construction (96/8 = 12). Palette is the 16-entry PCG/text palette at $E82200 -- NOT the $E82000 graphic palette -- in GGGGGRRRRRBBBBBI order, green and red SWAPPED versus RGB555 plus ONE SHARED intensity LSB; and because the text plane has exactly one 16-entry palette, per-scanline SCB palettes cannot be honoured (palette 0 always wins) -- unlike the ST, which really does emulate per-band SCB palettes from an MFP Timer B ISR, so this port stands alone in ignoring the SCB on screen. [SUPERSEDED 2026-08-06: the port moved to the 256-colour graphics plane (_iocs_crtmod 13, palette $E82000) precisely to fix this -- SCB per-band palettes now work, DOS-style. See the 2026-08-06 entry at the top. The rest of this paragraph still describes the port accurately.] Sprites COMPILE (shared 68k planar emitter, see the Reading notes). Toolchain is elf2x68k at toolchains/x68000/m68k-xelf (gcc 13.4.0, ELF -- so no leading-underscore asm symbols, and src/m68k/.s does not link there); disks are built by tools/xdftool.py (Human68k FAT12, 1024-byte sectors, .X executables); the emulator is a patched MAME 0.264 single-driver build (patches/mame-0.264-x68k-.patch). The only recorded rate-card input so far is from those patches: TVRAM costs ~1.92 extra cycles/access and GVRAM ~1.03 versus main memory, applied only at a 10 MHz CPU clock (docs/x68000_mame_crtc_timing.md:307-309, 328) -- which is why the 4x traffic cut to the text plane wins despite TVRAM's higher per-access cost. NO ceilings are derived for this port anywhere in this file, and none should be until the clock is fixed.
THE COLUMN BELOW IS NOT COMPARABLE TO THE OTHER FOUR. It is recorded raw
because the capture is expensive and the log is not kept in-tree, not
because the numbers mean what they say -- see measurement artifact 5.
Capture route: UBER built with the serial log tee
(EXTRA_CFLAGS=-DJOEY_LOG_SERIAL_TEE, which now reaches every port via
make/common.mk:26; src/core/debug.c calls jlpLogTee, implemented at
src/x68000/hal.c:353) against scripts/x68kSerialPeer.py, which MUST be
run as python3 -u or a killed peer loses the whole capture to stdio
buffering. This is the only route that gets a full log off the machine:
the disk route (scripts/verify-x68000-golden.sh) forces -DUBER_FRAMES=1u
and cannot produce timings at all -- and in practice cannot produce the
log either, because joeylog.txt never commits its size to the Human68k
image. Run header as reported: total wall time 15,945 ms, 877 frames
@ 55 Hz (both from the suspect clock).
| Op | X68000 (reported ops/sec) |
|---|---|
| jlSurfaceClear | 8775 |
| jlPaletteSet | 14353 |
| jlScbSetRange | 8305 |
| jlDrawPixel | 8580 |
| jlDrawLine H | 8855 |
| jlDrawLine V | 8800 |
| jlDrawLine diag | 8195 |
| jlDrawRect 100x100 | 8470 |
| jlDrawCircle r=16 | 8305 |
| jlDrawCircle r=80 | 7316 |
| jlFillRect 16x16 | 9295 |
| jlFillRect 80x80 | 22182 |
| jlFillRect 320x200 | 6986 |
| jlFillCircle r=40 | 9240 |
| jlSamplePixel | 10395 |
| jlTileFill | 8195 |
| jlTileCopy | 8525 |
| jlTileCopyMasked | 8745 |
| jlTilePaste | 9405 |
| jlTileSnap | 8690 |
| jlSpriteSaveUnder | 8580 |
| jlSpriteDraw | 8250 |
| jlSpriteRestoreUnder | 8886 |
| jlSpriteSaveAndDraw | 8305 |
| jlStagePresent full | 572 |
| jlInputPoll | 8250 |
| jlKeyDown | 53006 |
| jlKeyPressed | 53611 |
| jlMouseX | 125529 |
| joeyJoyConnected | 53006 |
| jlAudioFrameTick | 23756 |
| jlAudioIsPlayingMod | 84472 |
| surfaceMarkDirtyRect (via jlFillRect 32x32) | 8965 |
| jlSpriteDraw unaligned | 8030 |
| jlSpriteDraw sweep16 | 8470 |
| jlTilePaste map 10x5 | 8305 |
| gameFrame composite | 938 |
| jlTileMapPaste 10x5 | 8030 |
| gameFrameClean composite | 1245 |
The tell is in the table itself, no model needed: 26 of the 39 rows land
inside a single 7,316-9,405 band -- a 1.29x spread across ops whose real
work differs by orders of magnitude. jlDrawPixel (8,580), a 100x100
rect outline (8,470) and a 16x16 sprite draw (8,250) cannot all cost the
same; jlFillRect 80x80 (22,182) cannot be 2.4x FASTER than jlFillRect 16x16 (9,295); jlSurfaceClear (8,775) cannot land within 1.2x of
jlSamplePixel (10,395). The physics check, in the style of artifact 2:
the stage is 4 x 8,000 = 32,000 bytes, so 8,775 clears/sec is 114
us/clear = ~1,140 cycles at 10 MHz = 28 bytes PER CYCLE. Even at the ST
rate card's write-only fill rate of ~2.2 cyc/B, 32,000 bytes is ~70k
cycles = ~7.0 ms = ~140 ops/sec, so that row reads ~60x high. Same for
the composite: gameFrame at 938/s is 1.07 ms = ~10,700 cycles for a
256x96 fill + 16 tile pastes + 3 saveAndDraw + 3 restores + a 96-row
band present, when the fill ALONE is 12,288 B x 2.2 = 27k cycles (the
very figure this file uses for the ST gameFrame ceiling) -- 2.5x the
whole measured frame. Only the cheap rows survive: at ~8 us/call
jlMouseX (125,529), jlKeyDown (53,006) and jlAudioIsPlayingMod
(84,472) poll the clock far more often than once per frame and read
honestly.
For calibration once the clock is fixed, this port sits between the two 68k ports already in the table: a 10 MHz 68000 with NO blitter -- the ST's situation with 25% more clock -- writing a compact 4-plane surface in plain malloc'd main RAM (no Chip RAM constraint, nothing but the CPU ever reads it) and then scattering 32,000 bytes/frame into TVRAM at its ~1.92-cycle access penalty. So CPU-bound rows should be expected ABOVE the 8.0 MHz ST roughly on clock ratio, while the present should land between the ST (bus-floor CPU copy, 8 MHz) and the Amiga (7.09 MHz but a blitter and display-DMA contention). None of that is measured yet, and none of it can be until artifact 5 is closed.
Measurement honesty (read before trusting any number)
Five measurement artifacts have shaped this table's history. Three are fixed, one is inherent, and one -- the X68000 frame clock -- is OPEN and invalidates that port's entire ops/sec column: know them before comparing rows:
-
Per-iteration poll tax -- FIXED (Phase 0). The pre-audit bench loop called jlFrameCount() between every op; on the IIgs that was a GetTick toolbox chain (~600-800 cycles), so every small-op number carried the harness cost, not the op cost. Ops now run in adaptive batches per poll. Small-op numbers jumped 2-10x on every port when this landed; never compare against pre-Phase-0 captures.
-
SEI tick loss -- INHERENT on IIgs, rows tagged below. The benchmark clock is jlFrameCount = the VBL-interrupt-driven GetTick. An op that holds SEI while it runs suppresses those interrupts, the clock loses ticks, and the op over-reports. The three SEI-holding rows --
jlSurfaceClear,jlFillRect 320x200, andjlStagePresent full-- are CLOCK-UNRELIABLE in the IIgs column. The physics check: a full-screen op moves 32000 bytes at ~6 cyc/word, so ~29 ops/sec is the hard ceiling at 2.8 MHz. The present row's 68 and the clear row's 35 sit above it -- inflated; the honest values are at or below the ceiling. -
ST frame-rate miscalibration -- FOUND and FIXED 2026-07-06; the ST column below is the corrected (millis-clock) capture. Before the fix, ops/sec was
iters * jlFrameHz() / actualFramesand the ST'sjlpFrameHz()hardcoded 50 -- but Hatari runs the emulated ST at ~60 Hz VBL, so every ST number was scaled by 50/60 ~= 0.83 (~17% pessimistic). Caught by cross-checking against UBER's independent Timer-C 200 Hz millis clock (refresh-independent), which ran ~16-22% faster; the Amiga's two clocks agreed exactly (FS-UAE = PAL 50 Hz), confirming it was ST-only. Two fixes landed: (a) UBER now computes ops/sec fromjlMillisElapsed(its own hardware clock on the ST and DOS; derived from the frame counter on the IIgs -- identical to the old formula -- and on the Amiga and X68000, neither of which overrides jlpMillisElapsed; on the X68000 that counter is POLLED, so both UBER clocks carry the same error and the frame-vs-millis cross-check cannot see it -- artifact 5), and (b) STjlpFrameHz()now reads the shifter resolution + sync-mode registers to return the true 50/60/71 Hz. The per-opUBER-CLK:log line reports frame-vs-millis as a standing cross-check (they now agree). ST cells rose ~17% (e.g. jlSpriteDraw 503->585, jlFillRect 80x80 97->115); DOS/IIgs/Amiga unchanged. -
ST idle-audio ISR tax -- FOUND and FIXED 2026-07-07; the ST column below is the post-fix capture. jlpAudioInit started the 12.3 kHz Timer-A PWM ISR at init and left it running forever, even silent (~320 cycles/tick = ~49.5% of the CPU), and UBER inits audio before timing anything. Every ST capture before 2026-07-07 is ~2x pessimistic. Timer A is demand-driven now (runs only while a mod or SFX is audible), so the benches measure the op and an idle game gets the whole CPU. Amiga/DOS/IIgs audited clean for the same bug class (~1.8% / ~0.5% / 0% idle cost).
-
X68000 polled frame clock -- OPEN; that port's whole ops/sec column is invalid. _iocs_vdispst HANGS the machine when handed a plain C handler, so the V-DISP interrupt is never installed (src/x68000/hal.c:85-92; gVdispInstalled stays false) and jlpFrameCount instead edge-detects the MFP GPIP V-DISP bit each time something calls it (hal.c:310-323). A polled edge detector loses ticks exactly when the work is heavy, and the HAL says so in its own comment: "Measured: 400 audio refills across ~19 frames reported ZERO elapsed frames, because nothing polled in between" (hal.c:275-279). UBER polls once per 16-op batch (uber.c:201-219), so any op whose batch stretches past a V-DISP period makes the counter under-count, and ops/sec = iters * 1000 / millis over-reports by that same factor. The frame-vs-millis cross-check built for artifact 3 cannot catch it: the X68000 has no jlpMillisElapsed override, so millis comes from jlpGenericMillisElapsed, which is computed FROM jlpFrameCount (src/core/port.h:753-758, src/generic/genericPort.c:52-97) -- every UBER-CLK line on that port compares a number with itself, which is why they all agree to within 0.5%. Cheap rows (jlMouseX, jlKeyDown, jlAudioIsPlayingMod) poll often enough to catch every edge and read honestly; everything else is inflated, ~60x on the full-screen rows. THE FIX is recorded in the HAL: an asm thunk for _iocs_vdispst (save d0-d1/a0-a1, call, restore, correct return) so the interrupt owns the counter (hal.c:87-91, 298-300). Until that lands the X68000 gates on HASHES ONLY -- it is a correctness column, not a performance column.
Baseline (2026-07-08): absolute ops/sec per port
The IIgs is the reference and the perf floor (every other port must
meet or beat it -- project directive). Non-IIgs cells show absolute
ops/sec plus percentage of the IIgs number; sub-100% cells are bolded
as below-floor. tools/uber-perf-table emits this table directly;
regenerate after any re-capture (Amiga needs TIMEOUT=600 -- the
default cuts off the final checks). The tool still knows only the
original four ports (its PORTS list, tools/uber-perf-table:40-45) and
there is no bench-x68000.sh, so an X68000 column has to be added by
hand from a serial-tee capture -- see the 2026-08-05 entry. Capture provenance: all four
ports 2026-07-08 from a full make clean rebuild (Phase 6 W1+W2+W3
landed: fused IIgs sprite/fillRect entries, sprite-mark
specialization, per-row/fused presents on Amiga/ST/IIgs, ST movep
tile singles, DOS map walker, Amiga blitter copyRect + map hoist);
51/51 golden hashes x4 on this exact build; Amiga solo TIMEOUT=600;
IIgs from a fresh make iigs-disk image.
| Op | IIgs (ops/sec) | Amiga (ops/sec, % of IIGS) | Atari ST (ops/sec, % of IIGS) | DOS (ops/sec, % of IIGS) |
|---|---|---|---|---|
| jlSurfaceClear | 33 | 75 (227%) | 88 (267%) | 66 (200%) |
| jlPaletteSet | 1561 | 11503 (737%) | 10388 (665%) | 29828 (1911%) |
| jlScbSetRange | 163 | 4003 (2456%) | 4833 (2965%) | 8491 (5209%) |
| jlDrawPixel | 5576 | 5903 (106%) | 6522 (117%) | 9009 (162%) |
| jlDrawLine H | 1322 | 1753 (133%) | 2670 (202%) | 5682 (430%) |
| jlDrawLine V | 104 | 303 (291%) | 261 (251%) | 456 (438%) |
| jlDrawLine diag | 42 | 40 (95%) | 88 (210%) | 109 (260%) |
| jlDrawRect 100x100 | 110 | 285 (259%) | 279 (254%) | 470 (427%) |
| jlDrawCircle r=16 | 303 | 285 (94%) | 359 (118%) | 461 (152%) |
| jlDrawCircle r=80 | 70 | 63 (90%) | 83 (119%) | 98 (140%) |
| jlFillRect 16x16 | 785 | 1153 (147%) | 1225 (156%) | 3017 (384%) |
| jlFillRect 80x80 | 133 | 285 (214%) | 240 (180%) | 412 (310%) |
| jlFillRect 320x200 | 22 | 66 (300%) | 77 (350%) | 59 (268%) |
| jlFillCircle r=40 | 42 | 31 (74%) | 155 (369%) | 156 (371%) |
| jlSamplePixel | 4797 | 7253 (151%) | 9784 (204%) | 46266 (964%) |
| jlTileFill | 3239 | 3753 (116%) | 3855 (119%) | 11824 (365%) |
| jlTileCopy | 1981 | 2903 (147%) | 3626 (183%) | 10910 (551%) |
| jlTileCopyMasked | 681 | 1003 (147%) | 1749 (257%) | 3391 (498%) |
| jlTilePaste | 2700 | 3753 (139%) | 4685 (174%) | 12575 (466%) |
| jlTileSnap | 2520 | 6103 (242%) | 8516 (338%) | 19218 (763%) |
| jlSpriteSaveUnder | 602 | 2053 (341%) | 2611 (434%) | 10718 (1780%) |
| jlSpriteDraw | 1142 | 1153 (101%) | 1341 (117%) | 8146 (713%) |
| jlSpriteRestoreUnder | 662 | 1603 (242%) | 1865 (282%) | 7826 (1182%) |
| jlSpriteSaveAndDraw | 423 | 803 (190%) | 934 (221%) | 5478 (1295%) |
| jlStagePresent full | 73 | 609 (834%) | 407 (558%) | 1132 (1551%) |
| jlInputPoll | 229 | 4003 (1748%) | 2788 (1217%) | 3148 (1375%) |
| jlKeyDown | 10191 | 39753 (390%) | 44448 (436%) | 77688 (762%) |
| jlKeyPressed | 11432 | 38003 (332%) | 44448 (389%) | 74903 (655%) |
| jlMouseX | 20138 | 97703 (485%) | 117920 (586%) | 155583 (773%) |
| joeyJoyConnected | 10191 | 38903 (382%) | 44260 (434%) | 77697 (762%) |
| jlAudioFrameTick | 27612 | 56503 (205%) | 39885 (144%) | 11490 (42%) |
| jlAudioIsPlayingMod | 11029 | 56453 (512%) | 60441 (548%) | 127856 (1159%) |
| surfaceMarkDirtyRect (via jlFillRect 32x32) | 397 | 853 (215%) | 774 (195%) | 1479 (373%) |
| jlSpriteDraw unaligned | 1022 | 953 (93%) | 1089 (107%) | 6316 (618%) |
| jlSpriteDraw sweep16 | 67 | 58 (87%) | 70 (104%) | 463 (691%) |
| jlTilePaste map 10x5 | 59 | 75 (127%) | 94 (159%) | 257 (436%) |
| gameFrame composite | 12 | 31 (258%) | 32 (267%) | 19 (158%) |
| jlTileMapPaste 10x5 | 140 | 191 (136%) | 346 (247%) | 1224 (874%) |
| gameFrameClean composite | 20 | 71 (355%) | 90 (450%) | 192 (960%) |
IIgs full-screen ceiling (~29 ops/sec). A full-screen op moves 32000 bytes = 16000 word-stores x 6 cyc (STA long,X / PEI, both 3 cyc/byte) / 2.8 MHz ~= 34 ms ~= 29 ops/sec maximum. So
jlSurfaceClear,jlFillRect 320x200, andjlStagePresent fullcannot legitimately exceed ~29; a higher reading means the benchmark clock was perturbed. The clock isjlFrameCount= the VBL-interrupt- driven GetTick, so an op that holdsSEIwhile it runs can over- report its ops/sec. Treat full-screen IIgs numbers above ~29 as inflated and cross-check against the cycle model.
IIgs full-screen writes: the paper ceiling is refuted in MAME (Phase 6 W2). The old note here derived ~34 ms / ~29 ops/sec for any full-screen op from 3 cyc/byte stores. The W2 fitted model -- which closes P1, P2, and the idle-present row simultaneously -- measured a shadowed push write at ~1.73 us (~4.8 fast cycles), so the honest MAME full-screen write floor is ~55 ms of writes alone; on real hardware (datasheet ~0.98 us/write) it is ~31 ms. The
jlStagePresent fullrow is NOT a full-screen copy either way: it is the all-clean idle present (early-out scan only), which is why it reads far above any copy ceiling on every port. The GetTick clock itself is proven honest post-SEI-chunking (Lua cross-clock, tick-loss ~1.00);jlSurfaceClearremains the one mildly inflated row (1.21x).
Re-derived native-technique ceilings (NATIVE-PERF Phase 6 close-out, 2026-07-08)
The Phase-0 ceilings were paper models. Three measurement waves (W0 probes + ledgers, W1 IIgs sprite autopsy + fusion, W2 present fixes + fitted IIgs write model, W3 tile engines) invalidated or refined several of them; this section re-derives every headline ceiling from the MEASURED models and gates each row against the directive bar (measured within ~1.5x of hand-coded native technique for the row's exact work). "Ceiling" = the ops/sec a hand-coded native routine would achieve doing the row's exact work -- pixel/byte traffic at the measured machine rate, plus a minimal call and the dirty-band bookkeeping the present contract requires. Gap = ceiling / measured. Measured values: golden rows from the 2026-07-08 PERF.md capture; probe rows (P1/P2/P8) from the Phase-6 W0/W2 tracker entries. Gaps > 1.5x are bolded. Cells marked UNKNOWN have no recorded model and no honest derivation -- they are bounds, not ceilings.
Superseded by this section: (a) the "IIgs full-screen ceiling
(~29 ops/sec)" callout above -- the 34 ms full-screen "PEI floor" is
REFUTED in MAME: under the fitted write model (below), 32,000 shadowed
bytes cost ~55.4 ms of writes alone, so the honest MAME full-present
ceiling is ~14.5/s (~22/s real-HW variant), and the golden
jlStagePresent full reading (73/s) is the idle-early-out artifact,
not a copy; (b) the Phase-0 ST tile ceilings (12,700-15,500/s --
assumed byte-move bodies; the landed movep engine beats them); (c) the
Phase-0 IIgs sprite-draw ceiling (4,300/s -- assumed no
derive/mark/wrapper; W1 autopsy honest ceiling is 2,500-2,650/s);
(d) the Phase-0 Amiga/ST/DOS "present full copy" ceilings (the timed
row never re-dirtied; the P1/P2 probe models below are the real ones).
Row shapes (exact UBER geometry the ceilings are computed for): tile ops are one 8x8 tile; the map row is 10x5 = 50 tiles; sprite per-op rows are a 16x16 sprite at (40,30) -- x&15 = 8, so the 68k save/restore window is TWO 16-px groups = 256 B (IIgs/DOS chunky window is 128 B); gameFrame = 256x96 fill + 16 tile pastes + 3 saveAndDraw + 3 restores (sprites mod-16 aligned = 1-group windows) + a 96-row x 256-px dirty present; gameFrameClean = 3 16x16 copyRect erases + 3 draws + present; P1 = the 96-row x 256-px band alone (12,288 shadow bytes; 24,576 VGA bytes on DOS); P2 = three 16x16 windows (384 B; 768 VGA bytes). Clocks: ST 8.0 MHz, Amiga 7.09 MHz, IIgs 2.8 MHz; DOS has no cycle model -- only recorded rates are used.
Measured rate cards (replace the Phase-0 paper inputs)
- ST (68000): bus floor 4.0 cyc/B copied; move.l inline 5.0 cyc/B; movem streams ~4.6-4.8 cyc/B (write-only fill ~2.2 cyc/B); landed movep tile bodies 36 cyc/tile-row (paste/snap) and 48 (copy); specialized whole-tile mark ~480 cyc; measured per-present idle base 0.87 ms (hunt-3 refresh: jlStagePresent full 1150/s post the group-skip landings; the old 2.2 ms / P8 452/s figure predates them -- every ceiling row below that adds "+ 2.2 ms base" is pessimistic by ~1.3 ms and ST P2's honest re-derived ceiling is ~775/s, which RE-OPENS that row). P1 machine ceiling ~100 ops/s (session model) carries the same staleness.
- Amiga: blitter A->D ~2 chip cyc/word + ~230 cyc/plane setup + ~240 cyc Own/Disown; display-DMA contention ~2x on chip traffic (Phase-0 full-blit 110/s precedent); measured contended CPU tile body ~700 cyc/cell (backed out of the landed map walker, 191 maps/s = 742 cyc/tile incl. walk); P8 idle base 0.87 ms (hunt-3 refresh: jlStagePresent full 1143/s -- the old NEVER RECORDED note is resolved by the golden idle-dominated row post group-skips).
- DOS: the W0 "~1.7 us/VGA byte -- pure hardware rate" model is REFUTED (hunt-3): the dword-fused expand measured probe present96 at 29 ops/s = ~1.40 us/B over 24,576 B, EXCEEDING the old ceiling (~24) -- under DOSBox's instruction-counting core the write loop was software-bound, not hardware-bound. DOS present ceilings should be re-fit from ~1.40 us/B; gameFrameClean is PASS-at-floor (245/s vs the probe-split-derived ~244 floor); CPU-side recorded models only: single-tile paste ~57,000/s (Phase-0, unrevised) and map walker ~76,000 tiles/s (W3: ~26 measured vs 18-20 hand instructions/cell).
- IIgs: MAME-fitted write model (closed the P1/P2/P8 ledgers simultaneously): shadowed SHR push write wSh ~1.73 us/B, slow-I/O softswitch access ~1.0 us; measured per-present fixed base 2.75 ms (hunt-3 probe: presentIdle 363/s; the original 13.7 ms figure died with the band-memset fix -- any ceiling below still adding "+13.7" is pessimistic by ~11 ms). Real-HW datasheet variant: wSh ~0.977 us/B (fixed base assumed unchanged -- unverified on hardware). Tile hand-call floor 480 cyc; W1 sprite autopsy: honest full-op draw ~1,060-1,120 cyc (body ~670 + minimal per-row mark ~390 + 3-far-load derive + JSL; the fnAddr cache was built, measured a LOSS, and removed).
Atari ST (8.0 MHz 68000)
| Row | Ceiling (ops/s) | Derivation | Measured | Gap | Bar |
|---|---|---|---|---|---|
| jlTilePaste | ~9,200 | movep body 8x36=288 + mark ~480 + call ~100 = 868 cyc | 4,685 | 2.0x | FAIL |
| jlTileCopy | ~8,300 | movep body 8x48=384 + 480 + 100 = 964 cyc | 3,626 | 2.3x | FAIL |
| jlTileSnap | ~20,600 | body 288 + call ~100 = 388 cyc (no mark; measured overhead is ~651) | 8,516 | 2.4x | FAIL |
| jlTileMapPaste 10x5 | ~440 | 50 x ~350 cyc/tile (movep-honest) + one mark 480 + one validation ~200 = 18.2k cyc | 346 | 1.27x | PASS |
| jlSpriteDraw | 2,235-2,750 | Phase-0 pre-shifted movem-RMW model, 2,909-3,580 cyc -- UNREVISED paper | 1,341 | 1.7-2.1x | FAIL |
| jlSpriteSaveUnder | ~5,300 | 2-group window 256 B x 5.0 cyc/B + row stride ~230 = 1,510 cyc. P7-1 probe split (2026-07-09): full call 3,064 cyc = wrapper+validation 1,088 + dispatch/metadata ~660 + body 1,310; the window delta (g16 row) measured the body at 5.1 cyc/B = AT the move.l floor -- the whole gap is C glue | 2,611 | 2.0x | FAIL (lever: 68k wrapper trim/fusion) |
| jlSpriteRestoreUnder | ~4,000 | save copy 1,510 + sprite mark ~480 = 1,990 cyc. P7-1 probe: full 4,211 cyc, validation alone 422 (odd-bx row) | 1,900 | 2.0x | FAIL (same lever) |
| gameFrame | ~58 fps | fill 12,288 B x 2.2 = 27k + 16 pastes 13.9k + 3 saves 2.6k + 3 draws 9.3k (mid-range of the paper draw model) + 3 restores 4.1k + present96 80k = 136.9k cyc = 17.1 ms | 32 | 1.8x | FAIL |
| gameFrameClean | ~245 | 3 copyRect x ~750 cyc + 3 draws 9.3k + P2-class present 20.8k = 32.4k cyc = 4.0 ms | 90 | 2.7x | FAIL |
| P1 present96 | ~100 | model: 12,288 B x 4.0 cyc/B = 6.1 ms + P8 base 2.2 ms + snapshot traffic in one fused 200-row scan/snapshot pass ~= 10 ms | 73 | 1.37x | PASS |
| P2 present3win | ~385 | measured idle base 2.2 ms (P8 452/s) + 384 B copy+snapshot x ~8 cyc/B ~= 2.6 ms | 245 | ~1.6x | FAIL (marginal) |
Amiga (7.09 MHz 68000 + blitter)
| Row | Ceiling (ops/s) | Derivation | Measured | Gap | Bar |
|---|---|---|---|---|---|
| jlTilePaste | ~5,500 | measured contended body ~700 + mark 480 + call ~100 = 1,280 cyc (uncontended instruction model gives ~6,100-6,400 -> gap 1.6-1.7; the Phase-0 18% pair discrepancy persists as this band) | 3,753 | ~1.48x | PASS (at bar) |
| jlTileCopy | ~4,800-5,300 | chip-to-chip body ~750-900 + 480 + 100 cyc | 2,903 | 1.7-1.8x | FAIL (marginal) |
| jlTileSnap | ~8,900 | body ~700 + call ~100 = 800 cyc (no mark; writes go to fast RAM) | 6,103 | ~1.45x | PASS (at bar) |
| jlTileMapPaste 10x5 | ~200 | 50 x ~700 + mark 480 + validation ~200 = 35.7k cyc; the landed walker already beat its own design model (176-186 -> 191) | 191 | ~1.05x | PASS |
| jlSpriteDraw | ~1,350 | Phase-0 compiled-CPU ceiling; the 3,600-4,300/s blitter cookie-cut is DEFERRED by recon -- sync WaitBlit forfeits the win at <= 16x16 (this row's size) | 1,153 | 1.17x | PASS |
| jlSpriteSaveUnder | ~4,400-6,100 | SUPERSEDES the blitter paper (6,300-7,400): P7-1 probe (2026-07-09) measured the CPU body at 3.8 cyc/B via the window delta (482 cyc/128 B), so body ~960-1,900 cyc (contention band) + call ~200. Full call 3,372 cyc = wrapper+validation 1,524 + dispatch/metadata ~880 + body -- the gap is C glue, not the copy | 2,103 | ~2.1-2.9x | FAIL (lever: 68k wrapper trim/fusion) |
| jlSpriteRestoreUnder | ~2,900-3,700 | body + mark ~480 + call (same probe-derived band); full 4,423 cyc, validation alone 672 (odd-bx row) | 1,603 | ~1.8-2.3x | FAIL (same lever) |
| gameFrame | ~70 fps | fill blit 13.4k x2 contention = 27k + 16 pastes 20.5k + 3 saves 3k + 3 draws 15.8k + 3 restores 4.4k (sprite components interpolated from the paper models) + present96 29k = 99.7k cyc = 14.1 ms | 31 | 2.3x | FAIL |
| gameFrameClean | ~230 | 3 blit copyRects ~6.8k contended + 3 draws 15.8k + P2 present ~8k = 30.6k cyc = 4.3 ms | 71 | 3.2x | FAIL |
| P1 present96 | ~245 | 4-plane modulo blit 4x(230 + 1,536 w x 2) + 240 = 13.4k cyc, x2 contention + union scan ~2k = 29k cyc = 4.1 ms (the landed P7-3 blitter flush now takes this path; CPU-copy variant was 12,288 B x 4 cyc/B + scan ~= 7.5 ms -> ~133/s) | 78 | ~3.1x | FAIL |
| P2 present3win | UNKNOWN (write-only bound ~890-1,300) | 12 plane-blit setups 2.8k + 192 w x2 contention + Own/Disown + scan ~2k ~= 5.4-8k cyc = 890-1,310 ops/s; honest ceiling needs the Amiga P8 idle base = 661/s = 1.51 ms. Component split (2026-07-09 skip-mask probe): scan ~4.9 ms, flush ~3.3 ms (chip->chip CPU copy, below the 128-word blitter threshold), flip ~0.4 ms, residual ~1.2 ms. TWO levers landed: (1) presentScan.s 4-row fast path cut the scan (106 -> 128, +21%, after fixing a phantom-row-200 OOB); (2) the flush moved to the blitter by lowering AMIGA_PRESENT_BLIT_MIN_WORDS 128 -> 16 (chip->chip DMA vs contention-bound CPU copy): 128 -> 153 (+20%). Both 51/51 + P11 | 153 | <= ~6-8x vs bound | FAIL/UNKNOWN |
DOS (AT-class VGA; no cycle model -- recorded rates only)
| Row | Ceiling (ops/s) | Derivation | Measured | Gap | Bar |
|---|---|---|---|---|---|
| jlTilePaste | ~57,000 | Phase-0 single-tile model -- UNREVISED (W3 walker data is consistent with it) | 12,575 | 4.5x | FAIL |
| jlTileCopy | ~57,000 class | same 64-B traffic as paste; no separate model recorded | 10,910 | ~5.2x | FAIL |
| jlTileSnap | >= ~57,000 class | paste class minus the mark; no separate model recorded | 19,218 | ~3.0x | FAIL |
| jlTileMapPaste 10x5 | ~1,520 | recorded W3 model: ~76k tiles/s (walker at 61,200 x 26/[18-20] hand instructions/cell) / 50 | 1,224 | 1.24x | PASS (residue structural: map indirection + mark) |
| jlSpriteDraw | ~11,000 | Phase-0 model -- UNREVISED | 8,146 | 1.35x | PASS |
| jlSpriteSaveUnder | 14,500-18,000 | Phase-0 model -- UNREVISED | 10,718 | 1.35-1.68x | AT BAR (range straddles) |
| jlSpriteRestoreUnder | 14,500-18,000 | Phase-0 model -- UNREVISED (add mark: real ceiling somewhat lower) | 7,826 | 1.9-2.3x | FAIL |
| gameFrame | ~22 fps | present writes 24,576 B x 1.7 us = 41.8 ms + components ~2.5 ms ~= 44.3 ms (component figure not derivable from recorded rates -- no DOS fill rate; verdict insensitive: even at the measured P4 9.2 ms the ceiling is ~19.6 fps, gap ~1.03); W0 verdict: present-bound AT hardware | 19 | 1.2x | PASS |
| gameFrameClean | UNKNOWN (upper bound ~500) | erases+draws ~0.7 ms + present writes 768 B x 1.7 us = 1.3 ms; per-present fixed scan (~2.0 ms observed) never modeled | 192 | <= ~2.6x vs bound | FAIL/UNKNOWN |
| P1 present96 | ~24 | 24,576 VGA B x 1.7 us = 41.8 ms -- pure hardware rate | 23 | 1.04x | PASS (AT hardware) |
| P2 present3win | UNKNOWN (upper bound ~766) | 768 B x 1.7 us = 1.3 ms writes only; fixed base unmeasured (W0 closed DOS present work) | 305 | <= ~2.5x vs bound | FAIL/UNKNOWN |
Apple IIgs (2.8 MHz 65816; MAME-fitted write model)
| Row | Ceiling (ops/s) | Derivation | Measured | Gap | Bar |
|---|---|---|---|---|---|
| jlTilePaste | ~5,830 | recorded 480-cyc hand per-tile call | 2,700 | 2.2x | FAIL |
| jlTileCopy | ~4,860 | 480 + ~96 cyc extra far-src reads = 576 cyc | 1,981 | 2.5x | FAIL |
| jlTileSnap | ~5,800 | 480-cyc class (mark savings ~= far-read tax) | 2,520 | 2.3x | FAIL |
| jlTileMapPaste 10x5 | ~140 | 50 x ~390 cyc/tile floor (no tracker source; within noise of the landed ~400-cyc/tile walker) + one mark ~= 19.9k cyc; the landed walker already runs BELOW the 480-cyc per-call ceiling | 140 | ~1.0x | PASS |
| jlSpriteDraw | 2,500-2,650 | W1 autopsy honest full-op model: ~1,060-1,120 cyc = body ~670 + per-row mark ~390 + 3-far-load derive + JSL (fnAddr cache rejected by measurement) | 1,142 | ~2.2-2.3x | FAIL |
| jlSpriteSaveUnder | ~2,050 | MVN window body ~1,111 cyc (128 B x ~7 cyc/B + row stepping; = Phase-0 2,520/s) + derive/JSL ~250 = 1,360 cyc | 602 | 3.4x | FAIL |
| jlSpriteRestoreUnder | ~1,600 | save body 1,111 + fused mark ~390 + derive/JSL ~250 = 1,750 cyc (matches the W1 "restore 1,600+" honest target) | 662 | 2.4x | FAIL |
| gameFrame | ~16.5 fps (real HW ~20) | P4 components ~20.4 ms (fill 6,144 w x 6 cyc = 13.2 + pastes 2.7 + saves 1.5 + draws 1.2 + restores 1.9) + present96 40.3 ms = 60.7 ms | 12 | 1.37x | PASS |
| gameFrameClean | ~60 | 3 stage-copy erases ~1.2 ms + 3 draws ~1.2 ms + P2 present 14.4 ms ~= 16.7 ms | 20 | 3.0x | FAIL |
| P1 present96 | ~25 MAME (real HW ~35) | 15,360 B x 1.73 us = 26.6 ms writes + 13.7 ms measured base = 40.3 ms (landed Stage-B copies 13,824 B -> 37.6 ms -> ~27/s, gap 1.33x); minimal-copy variant (12,288 truly dirty B) 35.0 ms -> ~29/s, gap 1.43x -- PASS every way | 20 | 1.24x | PASS |
| P2 present3win | ~70 (real HW ~71) | 384 B x 1.73 us = 0.7 ms + 13.7 ms base = 14.4 ms | 52 | 1.34x | PASS |
IIgs real-HW-datasheet variants (wSh ~0.977 us/B, MAME-measured 13.7 ms base assumed unchanged -- unverified on hardware): P1 15.0 + 13.7 = 28.7 ms -> ~35/s; P2 ~14.1 ms -> ~71/s; full 200-row present 31.3 + 13.7 = 45.0 ms -> ~22/s (vs ~14.5/s MAME-fitted); gameFrame ~49 ms -> ~20 fps. The MAME-fitted column is what the bench harness can verify; the real-HW column is the datasheet claim.
Rows still outside the 1.5x bar -- reasons and recorded levers
- Tile singles, all four ports (1.7x-5.2x). Bodies are at machine floor (probe-validated splits, restated against the 2026-07-08 capture: ST paste overhead = 1,708 - 288 = 1,420 cyc, snap overhead 651 cyc with no mark; the W3-landing probe read 1,388/669); the gap is the per-call public wrapper + validation + whole-tile mark (Phase-2 exit finding: "remaining per-call floor = public wrapper + one JSL/call"). Recorded answer is the batch API, which PASSES on every port (1.0x-1.27x); no further singles lever is recorded -- the singles gap is accepted per-call API tax (absolute cost <= ~0.3 ms/call everywhere; IIgs tileCopy ~0.3 ms and tileSnap ~0.22 ms are the worst cases, all other rows <= ~0.2 ms).
- IIgs sprite rows (2.2-2.3x/3.4x/2.4x). Autopsy buckets: validation ~22% + wrapper residue in the register-starved public wrappers; fusion already landed (+18-22%), the fnAddr cache measured a loss and was removed. VALIDATION LEVER LANDED 2026-07-12 via the trusted fast-path API (jlSprite{SaveUnder,RestoreUnder,SaveAndDraw} Trusted, extended from the 68k P7-1 work): the trusted entries skip the ~22% validation and route into the existing fused MVN entries. IIgs save 602->665 (+10.5%), restore 662->722 (+9%), saveAndDraw 423->455 (+7.6%); shrParity badRows=0, 51/51 golden (safe entries unchanged). Smaller than the 68k wins because 65816 validation is a smaller fraction. Do not resurrect the fnAddr cache without a compiler-level win (LLVM816-ASKS candidate).
- ST sprite rows (~2.0x) / Amiga save-restore (~1.8-2.9x) / DOS restore (1.9x-2.3x). The P7-1 layer-split probe (2026-07-09) RESOLVED where the 68k gap lives: the copy bodies are AT the machine floor (window-delta-measured 5.1 cyc/B ST / 3.8 cyc/B Amiga), and ~55-70% of every call is the C wrapper+validation (1,088 cyc ST / 1,524 Amiga for save; 422/672 for restore's validation alone) plus dispatch/backup-metadata glue (~660-880 cyc). The old paper ceilings (Amiga blitter 6,300-7,400) are superseded by probe-derived CPU ceilings in the tables above; true gaps ~2x. LEVER ATTEMPT 1 (macro tightening, 2026-07-11): REVERTED, net loss. Both class macros were rewritten to their algebraically-minimal form, proven value-identical (102k save + 1.1M restore cases, 0 mismatches) and gated 51/51 -- benched save bit-identical on both ports, restore neutral on ST but -3% on Amiga (1603->1553; gcc-amigaos lowered the "simpler" C to +3 instructions). The wrapper is gated by register allocation (movem of 10-11 regs + interleaved validation), not arithmetic op count -- no C micro-opt moves these rows, same compiler-bound wall as the IIgs fnAddr-cache autopsy. LEVER ATTEMPT 2 (restore+draw mark fusion, 2026-07-11): LANDED, +5-6% on restore. Isolated the mark at ~26-29% of a 68k restore (probe: restore to a non-stage surface skips it). Three builds: a per-port body+mark asm shim REGRESSED both ports (broke the always-inline body's shared loads + added a call layer); a lean rolled-loop asm widen call REGRESSED (rolled dbra lost to gcc's unrolled indexed mark); INLINING the unrolled 16-row mark into jlSpriteRestoreUnder for the h==16 stage window (C only, no asm) WON -- ST restore 1900->1996, Amiga 1603->1703, both 51/51 + P11/P10/P5 parity. The recoverable win was the CALL BOUNDARY (only inlining removes it; asm always keeps a jsr). Draw bundled via the same inline (jlSpriteDraw + SaveAndDraw): Amiga draw +4% / saveAndDraw +6% / unaligned +5%; ST draw neutral (masked body dwarfs the mark); DOS untouched (guarded). gameFrame flat (present-bound). LEVER ATTEMPT 3 (TRUSTED fast-path entry, 2026-07-11): LANDED, the biggest 68k sprite win yet. New opt-in public API (jlSpriteSaveUnderTrusted / RestoreUnderTrusted / SaveAndDrawTrusted) that skips the per-call geometry-validation chain for a caller that pre-compiles its sprites and clamps positions on-screen (UB-if-violated contract; falls back to the safe sibling for anything unrecognised; Amiga/ST fast path, DOS just calls the safe entry). WINS ON ALL THREE COMPILED PORTS: save ST 705->817 (+16%) / Amiga 657->753 (+15%) / IIgs 602->665 (+10.5%); restore ST 529->609 (+15%) / Amiga 545->609 (+12%) / IIgs 662->722 (+9%); saveAndDraw ST 257->289 (+12.5%) / Amiga 273->289 (+6%) / IIgs 423->455 (+7.6%). On 68k the win EXCEEDS the raw ~422/672 cyc validation because dropping the checks also shrank gcc's frame/spills; IIgs routes into the fused MVN entries minus validation. Parity gates pass (P10 ST, P11 Amiga, shrParity IIgs); DOS 51/51 + IIgs 51/51 (falls back / safe entries unchanged). Batch APIs stay descoped (dispatch alone ~3-6%); the blitter cookie-cut stays deferred until a >= 32x32 golden row exists.
- Amiga presents P1/P2 (+ both Amiga composites, 2.8x-5.3x; P2 bound-only). W2 fixed the bounding-box over-copy (P2 copied bytes 4,992 -> 384 exact; W0's conviction figure was 3.5x vs the ~1,536 truly-dirty estimate) but the ~11 ms per-present fixed scan survives (P2 = 11.1 ms to move 384 B; P1 even regressed 8% on added scan cost -- accepted deviation). Recorded levers: fused-scan present asm on the ST presentSlam pattern (W2 follow-up) and a modulo-blitter present flush (the landed W3 copyRect blitter arm is the template). This is the single largest remaining item in the program.
- ST P2 (~1.6x, marginal) and ST composites (1.8x/2.7x). P2 residue is ~1.5 ms absolute (per-run move.l-chain entry + snapshot maintenance); W2 closed ST present work and no lever is recorded. gameFrame's gap is spread across sprite/fill wrapper taxes (no single lever); gameFrameClean's was dominated by jlSurfaceCopyRect's strided per-row memcpy dispatch -- the same overhead class the Amiga W3 blitter item deleted. LANDED 2026-07-12 (#3): an ST move.l-chain copyRect (src/atarist/copyRect.s, stCopyRectSpans68k) replaces the per-row memcpy CALL for short spans (the 48 8-byte-row window copies where mintlib memcpy's frame dwarfs the two move.l; long rows keep memcpy past the 64-byte crossover). gameFrameClean 90 -> 120 ops/sec (+33%), 51/51 + P10 displayMismatch=0; gameFrame unchanged (uses save/restore, not copyRect). Bigger than expected -- the small-window memcpy call overhead dominated the clean-buffer op.
- IIgs gameFrameClean (was 3.0x). Recorded forensic: ~30 ms/frame
unattributed, llvm-mos far-memcpy byte loops in the generic copyRect
erase engine (finding #79 class). LEVER LANDED 2026-07-12: an
IIgs-specific jlpSurfaceCopyRect override (src/iigs/hal.c, gated
JL_HAS_SURFACE_COPY_RECT) does an inline word copy for short spans --
compiles to
lda [dp],y / sta, NO far-memcpy call (verified in the -S output) -- for the 8-byte-row clean-buffer windows; long rows keep memcpy. gameFrameClean 18 -> 24 ops/sec (+33%), 51/51 + shrParity. NOTE: this was FIRST tried in the shared generic copyRect but that REGRESSED DOS -12% (x86 rep-movs memcpy already beats an inline word loop for 8-byte rows), so it is IIgs-only; DOS stays on generic memcpy. Remaining IIgs gameFrameClean gap (if any) would want the flagged llvm816 memcpy fix or an MVN block-copy for the long rows. - DOS gameFrameClean / P2 -- RESOLVED 2026-07-13 (perf-hunt-2 ranks 4/5/6). The missing display gate now exists: P12 vgaMismatch (an UBER_PROBE row + dosProbeVgaMismatch in src/dos/hal.c) re-derives every VGA byte independently of the expand LUT and reads 0/0/0 across the battery. The per-present scan was cut by the 4-row dword group skip (P8 idle 1137 -> 3881 ops/s, ~0.26 ms base) and the expand loop dword-fused (two LUT words per 32-bit store; DOSBox's instruction-counting core rewarded it directly): gameFrameClean 196 -> 243 ops/s (+24%), gameFrame 19 -> 23 (+21%), both 51/51. The probe component split answers the old unknown: gameFrameClean at 244/s = 4.10 ms/frame = present 3.06 ms + window erases 0.70 ms + sprite draws 0.34 ms -- the row is present-dominated and the remaining present cost is expand + per-row glue at the measured rate, not an unexplained residual. No further DOS present lever is recorded. (-fomit-frame-pointer also landed for flag parity with the 68k ports: input rows +28-29%, map paste +7%, zero regressions.)
Assumptions
- Clocks: ST 8.0 MHz, Amiga 7.09 MHz (PAL), IIgs 2.8 MHz; DOS has no cycle model, so DOS ceilings use only the recorded rates (1.7 us/VGA byte, 57k paste, 76k tiles/s walker).
- Irreducible per-call floors assumed: ~60-100 cyc call/shim on 68k, ~250 cyc JSL+derive on IIgs; the landed specialized dirty marks (~480 cyc 68k whole-tile, ~390 cyc IIgs fused) are treated as irreducible for row shapes that mark per call.
- Amiga display-DMA contention factor ~2x on chip-RAM traffic, taken from the Phase-0 full-blit 110/s precedent; uncontended variants are shown wherever they would change a verdict.
- The IIgs fixed present base was 13.7 ms MAME-measured (P8) when the ceilings below were derived. 2026-07-13: ~11 ms of that base was the dirty-band memset pair (see History) and is now ~2.75 ms -- ceilings that add "+ 13.7 ms base" are stale and pessimistic; re-derive with ~2.75 ms before gating new work on them. The real-HW column assumption (base carries over unchanged) is still unverified on hardware.
- Composite ceilings are sums of component ceilings at the rows' exact geometry (per-op sprite rows at x=40 use the 2-group 256-B 68k window; gameFrame's mod-16-aligned sprites use the 1-group 128-B window). Composite sprite components not separately recorded (ST draws, Amiga saves/restores/copyRects) are interpolations inside the quoted Phase-0 ranges.
- Ceilings labeled UNREVISED (ST/Amiga/DOS sprite models, DOS tile singles) are Phase-0 paper carried forward because no wave re-measured them; their gaps are quoted as ranges and should be re-derived before gating any new work on those rows. The ST P1 ~100 ops/s machine ceiling is a session model consistent with the P8 base + bus-floor copy, not a tracker-recorded number.
- Where no model or measurement exists (Amiga/DOS P8 idle bases, DOS copy/snap singles), the cell is an explicit upper bound or UNKNOWN -- no ceiling was invented.
Reading notes
- Sprite rows measure COMPILED code on all five ports since NATIVE-PERF Phase 1 (2026-07-06): the 68k ports JIT all-shift draw variants plus group-window save/restore (previously stubs that fell to interpreters), IIgs codegen re-enabled via finding #77, DOS via finding #74. The X68000 joined on the shared 68k PLANAR emitter with no emitter changes -- it takes the AMIGA arm of src/core/spriteDispatch.h (4-plane cdecl routines from src/m68k/spriteEmitPlanar68k.c, one pre-shifted variant per x % 8 phase, all 8 compiled), because its surface stride is the Amiga's 40 bytes and the display's 128-byte text-plane stride never reaches the emitter. The interpreted walkers remain only as arena-full degrade paths -- and are proven bit-identical to the compiled code by the 51-hash gate.
- The below-floor cells are the standing punch list, and after the 2026-07-07 audio-tax fix the ST has NONE -- every remaining below-floor graphics cell is the Amiga's: the circle family (diag line 95%, circle r=16 94%, r=80 87%, fillCircle 74% -- planar 4-plane RMW tax) and the two sprite workload rows a hair under floor (unaligned 95%, sweep16 89%; chip-RAM display-DMA contention that the fast-RAM-code ST does not pay). DOS jlAudioFrameTick (42%) is the SB half-buffer silence refill on an otherwise-trivial call -- it gates nothing. Floor parity is effectively DONE; the open work is ceiling-relative (NATIVE-PERF-PLAN.md Phase 2+), where every port still runs 2-5x below hand-coded technique on tile/sprite glue. Amiga jlFillCircle r=40 (74%) was probed 2026-07-06 with a full whole-circle asm routine (Bresenham + inline 4-plane long-fill span): it held 45/45 hashes but did NOT move the number (34 vs 31, within the coarse 11-iter sample), so the per-span C dispatch is NOT the bottleneck -- the asm was reverted. The real cost is the midpoint fill's 8-way symmetry filling each row 2-4x (idempotent overlap): the true lever is a scanline rewrite (one span per row, no overlap), which would roughly halve the byte writes and stay hash-identical. NB: this op is timed at batch=1 (~29 ms each) by the 50 Hz frame counter, a coarse measurement. (The old note here blaming a "4-5x model-vs-measured gap" partly on measurement was vindicated 2026-07-07: on the ST that gap was the idle-audio tax, artifact #4.)
- jlRandom has no timed row. Phase 8 (#43) rewrote its xorshift32 update in 16-bit halves on the IIgs only (~2-3x fewer cycles per call on the 65816; 32-bit shifts there are per-bit helper loops). Output is bit-identical -- host-proven over the golden sequence and a 50M-state sweep, and gated by UBER's random-golden check on all five ports.
- Input predicates stay functions (Phase 8 #40 decision): the post-clock-fix numbers (12K-77K ops/sec) put call overhead nowhere near a real frame budget; macro forms in the public header were rejected as not worth the API exposure.
- IIgs libc memset/memcpy remain ~4x slow (finding #79, toolchain-owned): llvm-mos runtime byte loops vs ORCA's MVN. The jlScbSetRange and jlPaletteSet IIgs rows carry that tax; library-side mitigations landed (#16 unrolled word copy, #21), the libc fix is flagged to the llvm816 session.
History
-
2026-07-20 (later still): music layer v2 = the complete feature set. jlAudioNoise percussion voice implemented natively on all four ports (ST YM noise generator on channel C with shared-volume guards both directions + a Supexec'd PSG all-off in shutdown; DOS OPL2 rhythm-mode snare on ch 7 with pitch via the hi-hat modulator and a shared adlibFreqToFnum factored out of jlpAudioVoice; Amiga 1024-byte LFSR chip sample looped on Paula ch 3, mt_musicmask re-applied after every mt_end; IIgs resident LFSR sample on SFX slot 3 with retrigger phase rotation -- and a real guard fix: jlpAudioVoice bounded by SFX_SLOTS could stamp tone squares over slots 3-4). Sequencer v2: noise as uniform voice 3, tempo-change opcode, jlMusicSetAtten fades, jlMusicPause/Resume. songbake v2: noise column, repeat blocks, mid-song tickms, ST-overlap warning. AGI's always-parsed-never-sounded SN76489 noise channel is now audible via psgNoiseApply. Design recon = 10-agent workflow (5 design + 5 adversarial verify) that caught 3 hardware-level bugs pre-implementation. Gates: builds x4, AGI host suite 1/1, sequencer host harness all-checks, golden 51/51 x4 cross-diff.
-
2026-07-20 (later): portable chip-tracker music layer landed (the cross-platform answer to the ST MOD tax -- one authored song, four native chip backends, no CPU mixing anywhere). New: src/core/music.c jlMusicPlay/Stop/IsPlaying sequencer over the JYM1 event stream (tools/songbake bakes text notation; pacing = jlMillisElapsed so tempo holds at 50/60/70 Hz; pumped by jlAudioFrameTick above the gAudioReady gate). Hunt-3 rank 4 landed as the IIgs backend: jlpAudioVoice rewritten from synthesize-per-call (~120-150 ms: C toneFill + forced 2 KB XOR recopy, plus it burned that cost even with audio uninitialized) to a resident DOC-unsigned square with freqWord/length/volume struct pokes + one NTPstreamsound JSL per note (~10 far stores). Amiga jlpAudioVoice implemented for the first time (was a stub): Paula ch 0-2 looping a 2-byte chip-RAM square, DMA auto-reload sustains for free. Voices vs MOD stay mutually exclusive on ST/Amiga (same chip). jlAudioFrameTick gains a no-song boolean check (musicPump) -- DOS row 57969 ops/s, unchanged within noise. Gates: 4-port builds clean, DOS/ST/Amiga/ IIgs golden 51/51 cross-diff OK.
-
2026-07-20: hunt-3 top ranks landed. Rank 1 probe re-baseline (IIgs+ST -DUBER_PROBE battery): the IIgs gameFrameClean mystery SOLVED -- the split rows (newly wired into the IIgs timeOp arm) attributed the 28.6 ms as erases 15.9 + draws 3.3 + present 9.4, exposing jlSurfaceCopyRect at ~5.3 ms PER 16x16 call: the Phase-7 C word-copy override had regressed into the dead-lever-#12 shape (llvm-mos re-deriving both far pointers per word, ~30+ instr/word, visible in the current -S). Fix = iigsCopyRectRows asm ([dp],y pairs, ~21 cyc/word): gameFrameClean 35 -> 56 ops/s (+60%), hash C51FE171 unchanged, 51/51. Rank 5: jlpInputPoll idle clear memset -> iigsByteFill: jlInputPoll 228 -> 662 ops/s (+190%). Rank 3: AGI restoreRectFromBackdrop -> jlSurfaceCopyRect (the stale "planar = rare (Amiga only)" fallback full-copied 32KB per object erase on BOTH 68k ports and forced full-dirty presents); builds clean x4, host AGI suite 1/1 (harness stub rot also repaired). Rank 2: rate cards refreshed in place (IIgs base 2.75 ms, ST idle 0.87 ms, Amiga idle 0.87 ms recorded for the first time, DOS 1.7 us/B model REFUTED -- measured 29/s EXCEEDS the old ceiling, refit ~1.40 us/B; ST P2 RE-OPENED at ~775/s re-derived ceiling). IIgs probe battery also refreshed: present96 26/s, present3win 133/s, trusted save 795/s, drawTrusted 1262/s, shrParity OK, band fingerprints cross-port identical.
-
2026-07-19: perf-hunt-2 close-out. Ranks 8/9/10/12/13/14/17 landed (ST draw-JIT group rewrite: jlSpriteDraw 1366 -> 1807 +32%, saveAndDraw +21%, gameFrameClean +14%; jlSpriteDrawTrusted on all ports; IIgs fused save: saveUnder 602 -> 722 +20%; ST PS_ROW inline arm; DOS inline-mark + full trusted family: safe restore +5%, trusted restore +21%; Amiga fillCircle exact-disk walk 31 -> 40 +29%; cleanup batch incl. 115 lines of unreachable ST walkers). Rank 11 v2 landed after a re-probe (presentScan.s 50-fixed-group restructure; Amiga gameFrameClean 91 -> 97 +6.6%; the re-measured split: scan 3.24 / flush 1.61 / flip 0.36 / residual 1.33 ms of a 6.54 ms present3win). Ranks 15/16 killed by their own gates. P7-2 stage 2 (metadata fills into a 10-arg asm entry) was built, measured 722 -> 602, and REVERTED -- llvm-mos argument marshalling costs more than the far stores it replaces (the fnAddr-cache wall; P7-2 is closed, both directions measured). All landings 51/51 cross-port + P5/P11/P12 parity.
-
2026-07-13 (later): perf-hunt-2 ranks 2/4/5/6/7 landed. ST+Amiga idle-present group scan (4 rows per longword vs the 0xFF clean sentinel): jlStagePresent full ST 407 -> 1150/s, Amiga 609 -> 1131/s; composites unchanged. DOS: -fomit-frame-pointer (flag parity; input rows +28-29%), 4-row group skip in jlpPresent (P8 idle 1137 -> 3881/s), dword-fused VGA expand (gameFrameClean 196 -> 243, gameFrame 19 -> 23), NEW P12 vgaMismatch display gate (dosProbeVgaMismatch + UBER_PROBE battery, 0/0/0), and the gameFrameClean component split (present 3.06 / erases 0.70 / draws 0.34 ms). All four ports 51/51 cross-diffed after each landing.
-
2026-07-13: IIgs dirty-band memset fix (perf-hunt-2 rank 1+3). The "13.7 ms irreducible present base" was mostly two libc memset far calls: stageDirtyClearAll (runs after EVERY present) memset both 200-byte band arrays, and llvm-mos lowers memset to
jsl memset-- libc'swhile (n--) *d++ = cbyte loop (the old "MVN-seeded" comments were wrong; finding-#79 class). Replaced with asm word fills in joeyDraw.s (iigsStageBandsClearAll / MarkAll, abs,x descending; a C word-fill emitted ~35 instr/store and was rejected) plus iigsByteFill ([dp],y) for jlScbSetRange. Measured (MAME): jlStagePresent full 67 -> 363/s (fixed base ~13.7 -> ~2.75 ms), jlScbSetRange 163 -> 1322/s, jlSurfaceClear 33 -> 56/s, gameFrameClean 24 -> 35/s, gameFrame 12 -> 14/s. 51/51 golden vs Amiga reference; verify-iigs UBER framebuffer checksum byte- identical to the memset build. NOTE: every IIgs ceiling that bakes in the 13.7 ms base (P1/P2 rows, assumption 4, the rate-card fixed base) is now stale on the optimistic side for the library and needs re-derivation with the ~2.75 ms base; the 2026-07-08 baseline table rows for the five ops above are superseded by the numbers here. -
2026-07-08: Phase 6 W1+W2+W3 landed (fused IIgs sprite/fillRect entries + sprite-mark specialization; per-row/fused presents on Amiga/ST/IIgs with display-parity probe gates P8-P11; ST movep tile singles, DOS map walker, Amiga blitter copyRect + map hoist). Table regenerated from a clean rebuild, 51/51 x4. Ceilings re-derived from the measured models (fitted IIgs write cost, ST bus floor, movep/blitter rates) -- the old ~29 ops/s IIgs full-screen note is superseded in place.
Pre-2026-07-04 tables (ORCA-era IIgs column, artifact-inflated DOS numbers, un-batched loops) live in git history only -- every conclusion drawn from them was re-derived from Phase 0 onward. The Phase 0 baseline (2026-07-04) and per-phase deltas are recorded in PERF-AUDIT-PLAN.md's progress tracker; headline gains since Phase 0: IIgs interpreted spriteDraw 18 -> 48 ops/sec (Phase 4), IIgs tilePasteMono asm 3.35x (Phase 7 #3), ST flood plane hooks 5.6x (Phase 7 #4), DOS paletteSet +92% / tileSnap +42% (Phases 2/5), DOS compiled sprites 284 -> ~6100 ops/sec (#74).