diff --git a/NATIVE-PERF-PLAN.md b/NATIVE-PERF-PLAN.md index 71c53e4..38bae11 100644 --- a/NATIVE-PERF-PLAN.md +++ b/NATIVE-PERF-PLAN.md @@ -241,7 +241,17 @@ perf-directive memory with the outcome. ZERO below-floor cells for the first time; remaining below-floor is Amiga-only (circle family + unaligned 95% / sweep16 89%) + DOS audioFrameTick (gates nothing). -- [~] Phase 2 IN PROGRESS 2026-07-07. Recon: 6-agent workflow + +- [x] Phase 2 COMPLETE 2026-07-07 (final gate 49/49 x4 + iigs-verify; + detail below). Steps D3 closed it: R6 TMASK_BYTE macro inline + (copyMasked +9%, loop-back needed brl past 8-bit branch range) and + R10 iigsDrawPixelMark fused plot+mark (jlDrawPixel 3,179 -> 5,576, + +75%; batching plotters keep the unfused inner). EXIT NUMBERS: + tilemap tiles/s ST 4,050 / Amiga 3,750 / IIgs 2,950 / DOS 12,900 + (from 1,600 / 2,300 / 1,750 / 10,000); gameFrame ST 21 / Amiga 25 / + IIgs 10 / DOS 19; IIgs tilePaste gap-to-ceiling 2.15x. Remaining + per-call floor = public wrapper + one JSL/call -> Phase 3's job. + Optional non-gating leftovers: R7 saveAndDraw consolidation, R12 + IIgs C-residue reshapes. Recon: 6-agent workflow + merged synthesis (12 ranked items R1-R12) + calibrated-probe layer split (see ERRATUM above). Step A LANDED (gate 49/49 x4): plane-pointer inlining via new src/amiga/amigaPlanar.h + @@ -263,10 +273,50 @@ perf-directive memory with the outcome. pointers either way; simple loop re-measured exactly at ST baseline; DOS flat) -- so Amiga unrolls, everyone else keeps the simple loop (documented at the #if in surface.c). - REMAINING: R4a 68k tile call-chain flatten + R8 Amiga unroll check; - Step D IIgs glue (R2/R5/R6/R10 fused asm entries -- biggest IIgs - upside, ~2x its tile rows; own session recommended); R9 palette/SCB - memcmp->flags on 68k/DOS presents (hygiene). + Step E1 LANDED (gate 49/49 x4): R4a = ST/Amiga tile ops moved to + always-inline per-port headers (stTile.h / amigaTile.h; jlpTile* + macro-aliased in port.h before the dispatch tail, headers included + at its END so the portData==NULL fallbacks see the jlpGenericTile* + prototypes; ST_BYTES_PER_ROW/GROUP moved to stPlanar.h; + surface68kStTileFill8x8 extern moved to stTile.h) -- ST tiles + +9-14%, Amiga +5-11%, DOS flat (already single-layer). R9 = + ST/Amiga/DOS presents consume gStageScbDirty/gStagePaletteDirty + instead of memcmp'ing 512-712-byte cached mirrors (mirrors + Amiga + memcmpLongs deleted; stageAlloc re-arms both flags for + jlShutdown/jlInit cycles) -- gameFrame ST 19->21 / Amiga 23->25 / + DOS 18->19 fps. Step E2 LANDED (gate 49/49 x4): R8 Amiga 8-row + d16(An) unrolls in amigaTile.h (fill/copy/paste/snap; + copyMasked/pasteMono stay rolled) -- tilePaste +36%, tileSnap +52%, + tileFill +32%, tileCopy +28%, map row 56->75. SESSION NET (from + honest post-audio-fix baselines): map row tiles/s ST 3,100->4,000 / + Amiga 2,300->3,750 / DOS 10,000->12,900; gameFrame ST 19->21 / + Amiga 23->25 / DOS 18->19 (and vs the pre-audio-fix world: ST + gameFrame 9->21). PERF.md table regenerated from the E2-gate logs. + Step D1 LANDED (IIgs gate 49/49 + make iigs-verify framebuffer + check; other ports covered by the D1-gate C-side pass): R2 fused + stage-tile entries iigsTileFillMark/PasteMark/CopyMark in + joeyDraw.s (.text.iigsFusedTile -- ONE section for entries + the + stageTileMark helper; a cross-section 16-bit jsr wild-jumped into + $00CFxx twice before the section-atomicity rule was learned, see + the cross-section-jsr feedback memory). One JSL now does + derive($01:2000 + gRowOffsetLut[by*8] + bx*4) + op via long,X + + band widen. IIgs: tileFill +64% (3,251), tilePaste +61% (2,710), + tileCopy +38%, map row 35->59 (+69%), gameFrame 10->11 fps; + tilePaste gap to ceiling 3.5x -> 2.15x. port.h stage-gates the + three macros; tile.c's tileMarkDirtyFused knob keeps + copyMasked/pasteMono on the separate marker. Step D2 LANDED (gate + 49/49 x4): R5 spriteCompiled* dispatchers moved to + src/core/spriteDispatch.h as always-inline, consumed only by + sprite.c (extern decls removed from spriteInternal.h; compile-time + machinery stays in spriteCompile.c). Sprite rows: IIgs +6-13% + (SaveAndDraw 423), ST SaveUnder +15%, Amiga +5-11%, DOS +2-16%. + Measurement notes recorded in PERF.md: gameFrame quantizes +/-1 + fps run-to-run; DOS unaligned-draw row is bimodal (~3,800/~5,100) + across identical binaries -- don't chase it. REMAINING: R6 + copyMasked tmaskByte de-layering (32 jsr/call = 384 cyc + BSS + scratch, 13-20% of the op; + optional CopyMaskedMark fusion), R10 + fused drawPixel+mark asm entry (~820 -> ~300 cyc), optional R7 + saveAndDraw consolidation, optional R12 IIgs C-residue reshapes. - [ ] Phase 3 - [ ] Phase 4 - [ ] Phase 5 diff --git a/PERF.md b/PERF.md index bc58b20..ee9269e 100644 --- a/PERF.md +++ b/PERF.md @@ -102,6 +102,85 @@ half-buffer IRQ), IIgs 0% (no handler installed until PlayMod). ALL pre-2026-07-07 ST comparisons -- including the Phase-0 ceiling gaps and the Phase-1 relative results -- understated the ST by ~2x. +**2026-07-07, NATIVE-PERF PHASE 2 COMPLETE -- Steps D3 (R6 + R10) +close the phase:** (R6) the IIgs masked-copy inner's 4 x 8 jsr/rts +round-trips to tmaskByte were inlined via the TMASK_BYTE assembler +macro (identical body; the row loop-back became brl -- the expanded +body exceeds 8-bit branch range): jlTileCopyMasked 625 -> 681 (+9%, +exactly the 384-cycle model). (R10) the stage arm of jlDrawPixel is +one fused JSL (iigsDrawPixelMark: single-row widen + nibble RMW plot +against the pinned stage base): jlDrawPixel 3,179 -> 5,576 (+75%, 502 +cyc/call incl. wrapper vs the measured 880 before); the batching +plotters (lines/circles, plotPixelNoMark) keep the unfused inner and +were flat, as were tiles/sprites. Final gate 49/49 x4 + make +iigs-verify. PHASE-2 EXIT vs its gate ("small-op rows move measurably +toward Phase-0 ceilings on all ports; goldens hold"): tilemap row +tiles/sec ST 1,600 -> 4,050 / Amiga 2,300 -> 3,750 / IIgs 1,750 -> +2,950 / DOS 10,000 -> 12,900; gameFrame ST 9 -> 21 / Amiga 18 -> 25 / +IIgs 10 / DOS 18 -> 19; IIgs tilePaste ceiling gap 3.5x -> 2.15x, ST +2.6x. The remaining per-call floor is the public wrapper + one +JSL/call -- Phase 3 batching territory. Optional leftovers R7 +(saveAndDraw consolidation) and R12 (IIgs C-residue reshapes) are +recorded in the plan, not phase-gating. + +**2026-07-07, NATIVE-PERF Phase 2 Step D2 LANDED -- sprite dispatch +flattening (R5):** the spriteCompiled* runtime dispatchers moved from +spriteCompile.c into spriteDispatch.h as always-inline functions +consumed only by sprite.c, deleting the cross-TU call per compiled +sprite op and letting the compiler share the slot/field loads the +caller already did. Gated 49/49 x4. Sprite rows: IIgs +6-13% +(SaveAndDraw 376 -> 423, Draw 906 -> 966), ST SaveUnder +15% (2,611), +Amiga +5-11% (SaveUnder 2,053), DOS SaveUnder +16% (10,743). Two +measurement footnotes: gameFrame wobbles +/-1 fps between runs +(single-digit iters per 16-frame window -- quantization, not signal; +ST has bounced 21/22 across three consecutive gates), and DOS +`jlSpriteDraw unaligned` is BIMODAL across runs of identical binaries +(~3,800 vs ~5,100; seen 07-06 and again today; the all-phase sweep16 +row is stable at ~362, so per-phase behavior is fine -- likely DOSBox +dynamic-core state; do not chase single-run deltas on that row). + +**2026-07-07, NATIVE-PERF Phase 2 Step D1 LANDED -- IIgs fused tile +entries (R2):** jlTileFill/Copy/Paste on the stage now run ONE JSL +each: new joeyDraw.s entries (iigsTile*Mark, section +.text.iigsFusedTile) derive the destination in asm from the pinned +$01:2000 stage base + gRowOffsetLut, write via long,X with the stage +literal folded into each store, and widen the dirty band via a shared +stageTileMark helper -- replacing C-side pointer math plus two JSLs +(op + marker). Gated 49/49 (pixels) AND make iigs-verify (framebuffer +-- the mark half is hash-blind, so the display check is load-bearing). +RESULTS: IIgs tileFill 1,981 -> 3,251 (+64%), tilePaste 1,687 -> 2,710 +(+61%), tileCopy 1,441 -> 1,988 (+38%), map row 35 -> 59 (+69% = 2,950 +tiles/s), gameFrame 10 -> 11 fps; unfused controls (copyMasked, snap, +sprites) flat. IIgs tilePaste is now ~1,030 cyc/call vs the 480-cyc +hand-coded ceiling (gap 2.15x, was 3.5x). Toolchain landmine found en +route: a 16-bit jsr between .section blocks can cross OMF segment +banks (no link error, wild jump into $00CFxx) -- helpers must share +their callers' section or be jsl/rtl; see the cross-section-jsr +feedback memory. + +**2026-07-07, NATIVE-PERF Phase 2 Steps A-E LANDED (per-op tax):** +five hash-gated steps, each 49/49 x4. (A) plane-pointer access +inlined into the 68k compiled-sprite dispatchers (amigaPlanar.h / +stPlanar.h; the jlpSurfacePlanePtr cross-TU call -- ~190 cyc x4 per +Amiga dispatch -- deleted): Amiga sprite ops +10-23%. (B) whole-tile +dirty marks specialized (band = [bx*2, bx*2+1] with compile-time +proof, unrolled 8-row widen): tiles +9-28% on 68k/DOS. (C) jlDrawRect +raw-edge fills + ONE band-identical mark (+22-40%); the multi-row +widen unroll is Amiga-only by measurement (gcc-mint compiled it WORSE, +ST -5% -- see the per-port #if in surface.c). (D=R9) presents consume +the core gStageScbDirty/gStagePaletteDirty flags instead of +memcmp'ing 512-712-byte cached mirrors: gameFrame +1-2 fps on +ST/Amiga/DOS. (E) ST/Amiga tile ops moved into always-inline per-port +headers (stTile.h / amigaTile.h), deleting the second cdecl layer +(ST tiles +9-14%, Amiga +5-11%), then the Amiga row loops unrolled to +d16(An) form (tilePaste +36%, tileSnap +52%). NET SESSION MOVEMENT on +the tile gate row (tiles/sec, map 10x5): ST 1,600 -> 4,000; Amiga +2,300 -> 3,750; DOS 10,000 -> 12,900. gameFrame: ST 9 -> 21, Amiga +23 -> 25, DOS 18 -> 19 fps (IIgs untouched at 10 -- its fused-asm +glue is Phase 2 Step D, still open). The present-row jumps (ST 467, +DOS 1,137) are the dirty-union idle early-out running flag-gated; +gameFrame is the honest present signal. + Correctness statement backing this table: all four ports produce BYTE-IDENTICAL pixels on every measured op and every correctness check -- **49/49** surface hashes match the frozen goldens in @@ -171,43 +250,43 @@ unaffected by the ST fix). | Op | IIgs (ops/sec) | Amiga (ops/sec, % of IIGS) | Atari ST (ops/sec, % of IIGS) | DOS (ops/sec, % of IIGS) | | --- | --- | --- | --- | --- | -| jlSurfaceClear | 33 | 75 (227%) | 89 (270%) | 67 (203%) | -| jlPaletteSet | 1561 | 11453 (734%) | 10255 (657%) | 29828 (1911%) | -| jlScbSetRange | 163 | 4003 (2456%) | 4744 (2910%) | 8491 (5209%) | -| jlDrawPixel | 3179 | 5853 (184%) | 6645 (209%) | 23689 (745%) | -| jlDrawLine H | 1142 | 1703 (149%) | 2660 (233%) | 5708 (500%) | -| jlDrawLine V | 99 | 269 (272%) | 265 (268%) | 456 (461%) | -| jlDrawLine diag | 42 | **40 (95%)** | 86 (205%) | 108 (257%) | -| jlDrawRect 100x100 | 89 | 203 (228%) | 216 (243%) | 378 (425%) | -| jlDrawCircle r=16 | 303 | **285 (94%)** | 359 (118%) | 461 (152%) | -| jlDrawCircle r=80 | 70 | **61 (87%)** | 83 (119%) | 98 (140%) | -| jlFillRect 16x16 | 605 | 1153 (191%) | 1211 (200%) | 3017 (499%) | -| jlFillRect 80x80 | 110 | 269 (245%) | 236 (215%) | 412 (375%) | -| jlFillRect 320x200 | 21 | 66 (314%) | 78 (371%) | 59 (281%) | +| jlSurfaceClear | 33 | 75 (227%) | 88 (267%) | 66 (200%) | +| jlPaletteSet | 1561 | 11403 (730%) | 10388 (665%) | 29809 (1910%) | +| jlScbSetRange | 163 | 3953 (2425%) | 4833 (2965%) | 8495 (5212%) | +| jlDrawPixel | 5576 | 5853 (105%) | 6522 (117%) | 23687 (425%) | +| jlDrawLine H | 1142 | 1753 (154%) | 2670 (234%) | 5708 (500%) | +| jlDrawLine V | 99 | 303 (306%) | 261 (264%) | 468 (473%) | +| jlDrawLine diag | 42 | **40 (95%)** | 88 (210%) | 109 (260%) | +| jlDrawRect 100x100 | 73 | 285 (390%) | 279 (382%) | 454 (622%) | +| jlDrawCircle r=16 | 304 | **285 (94%)** | 366 (120%) | 447 (147%) | +| jlDrawCircle r=80 | 70 | **63 (90%)** | 83 (119%) | 98 (140%) | +| jlFillRect 16x16 | 602 | 1153 (192%) | 1203 (200%) | 3013 (500%) | +| jlFillRect 80x80 | 110 | 285 (259%) | 240 (218%) | 414 (376%) | +| jlFillRect 320x200 | 21 | 68 (324%) | 77 (367%) | 58 (276%) | | jlFillCircle r=40 | 42 | **31 (74%)** | 155 (369%) | 156 (371%) | -| jlSamplePixel | 4815 | 7253 (151%) | 9603 (199%) | 46266 (961%) | -| jlTileFill | 1981 | 2253 (114%) | 3144 (159%) | 9301 (470%) | -| jlTileCopy | 1441 | 1853 (129%) | 2720 (189%) | 8764 (608%) | -| jlTileCopyMasked | 623 | 953 (153%) | 1366 (219%) | 3147 (505%) | -| jlTilePaste | 1741 | 2253 (129%) | 3203 (184%) | 9834 (565%) | -| jlTileSnap | 2520 | 3503 (139%) | 5751 (228%) | 19290 (765%) | -| jlSpriteSaveUnder | 545 | 1503 (276%) | 2196 (403%) | 9248 (1697%) | -| jlSpriteDraw | 902 | 953 (106%) | 1188 (132%) | 6234 (691%) | -| jlSpriteRestoreUnder | 510 | 1203 (236%) | 1662 (326%) | 5925 (1162%) | -| jlSpriteSaveAndDraw | 376 | 603 (160%) | 803 (214%) | 4271 (1136%) | -| jlStagePresent full | 67 | 359 (536%) | 125 (187%) | 360 (537%) | -| jlInputPoll | 228 | 4003 (1756%) | 2788 (1223%) | 3148 (1381%) | -| jlKeyDown | 10191 | 39303 (386%) | 45286 (444%) | 77548 (761%) | -| jlKeyPressed | 11389 | 38053 (334%) | 45286 (398%) | 75092 (659%) | -| jlMouseX | 20214 | 96703 (478%) | 117981 (584%) | 155372 (769%) | -| joeyJoyConnected | 10191 | 39303 (386%) | 43440 (426%) | 77548 (761%) | -| jlAudioFrameTick | 27509 | 55853 (203%) | 40637 (148%) | **11527 (42%)** | -| jlAudioIsPlayingMod | 17627 | 56453 (320%) | 60441 (343%) | 127708 (725%) | -| surfaceMarkDirtyRect (via jlFillRect 32x32) | 323 | 853 (264%) | 774 (240%) | 1472 (456%) | -| jlSpriteDraw unaligned | 845 | **803 (95%)** | 992 (117%) | 5118 (606%) | -| jlSpriteDraw sweep16 | 56 | **50 (89%)** | 64 (114%) | 363 (648%) | -| jlTilePaste map 10x5 | 35 | 46 (131%) | 62 (177%) | 200 (571%) | -| gameFrame composite | 10 | 23 (230%) | 19 (190%) | 18 (180%) | +| jlSamplePixel | 4797 | 7253 (151%) | 9603 (200%) | 46397 (967%) | +| jlTileFill | 3251 | 3753 (115%) | 3855 (119%) | 11840 (364%) | +| jlTileCopy | 1988 | 2903 (146%) | 3445 (173%) | 10913 (549%) | +| jlTileCopyMasked | 681 | 1003 (147%) | 1632 (240%) | 3391 (498%) | +| jlTilePaste | 2710 | 3803 (140%) | 4033 (149%) | 12545 (463%) | +| jlTileSnap | 2530 | 6003 (237%) | 6403 (253%) | 19247 (761%) | +| jlSpriteSaveUnder | 602 | 2103 (349%) | 2611 (434%) | 10743 (1785%) | +| jlSpriteDraw | 966 | 1103 (114%) | 1283 (133%) | 6390 (661%) | +| jlSpriteRestoreUnder | 545 | 1453 (267%) | 1781 (327%) | 6246 (1146%) | +| jlSpriteSaveAndDraw | 423 | 753 (178%) | 892 (211%) | 4656 (1101%) | +| jlStagePresent full | 67 | 540 (806%) | 467 (697%) | 1132 (1690%) | +| jlInputPoll | 228 | 4003 (1756%) | 2841 (1246%) | 3148 (1381%) | +| jlKeyDown | 10191 | 39753 (390%) | 44448 (436%) | 77738 (763%) | +| jlKeyPressed | 11389 | 38003 (334%) | 44448 (390%) | 74907 (658%) | +| jlMouseX | 20214 | 97703 (483%) | 117920 (583%) | 155685 (770%) | +| joeyJoyConnected | 10191 | 38903 (382%) | 44260 (434%) | 77748 (763%) | +| jlAudioFrameTick | 27509 | 57003 (207%) | 39885 (145%) | 59372 (216%) | +| jlAudioIsPlayingMod | 11071 | 56403 (509%) | 60441 (546%) | 128214 (1158%) | +| surfaceMarkDirtyRect (via jlFillRect 32x32) | 323 | 853 (264%) | 788 (244%) | 1476 (457%) | +| jlSpriteDraw unaligned | 906 | **903 (100%)** | 1070 (118%) | 3837 (424%) | +| jlSpriteDraw sweep16 | 60 | **56 (93%)** | 68 (113%) | 362 (603%) | +| jlTilePaste map 10x5 | 59 | 78 (132%) | 80 (136%) | 259 (439%) | +| gameFrame composite | 10 | 25 (250%) | 21 (210%) | 19 (190%) | > **IIgs full-screen ceiling (~29 ops/sec).** A full-screen op moves > 32000 bytes = 16000 word-stores x 6 cyc (STA long,X / PEI, both diff --git a/include/joey/platform.h b/include/joey/platform.h index b93cd25..6c49247 100644 --- a/include/joey/platform.h +++ b/include/joey/platform.h @@ -140,12 +140,12 @@ #define JL_HAS_DRAW_CIRCLE // amiga planar circle, function #define JL_HAS_FILL_CIRCLE // amiga planar fill circle, function -- returns bool #define JL_HAS_FILL_RECT // amiga planar fill rect, function - #define JL_HAS_TILE_FILL // amiga planar tile fill, function - #define JL_HAS_TILE_COPY // amiga planar tile copy, function - #define JL_HAS_TILE_COPY_MASKED // amiga planar masked copy, function - #define JL_HAS_TILE_PASTE // amiga planar tile paste, function - #define JL_HAS_TILE_SNAP // amiga planar tile snap, function - #define JL_HAS_TILE_PASTE_MONO // amiga planar mono colorize+paste, function + #define JL_HAS_TILE_FILL // amiga planar tile fill, inline (amigaTile.h) + #define JL_HAS_TILE_COPY // amiga planar tile copy, inline (amigaTile.h) + #define JL_HAS_TILE_COPY_MASKED // amiga planar masked copy, inline (amigaTile.h) + #define JL_HAS_TILE_PASTE // amiga planar tile paste, inline (amigaTile.h) + #define JL_HAS_TILE_SNAP // amiga planar tile snap, inline (amigaTile.h) + #define JL_HAS_TILE_PASTE_MONO // amiga planar mono colorize+paste, inline (amigaTile.h) #define JL_HAS_SPRITE_DRAW // amiga planar sprite draw, function #define JL_HAS_SPRITE_SAVE // amiga planar sprite save, function #define JL_HAS_SPRITE_RESTORE // amiga planar sprite restore, function @@ -194,12 +194,12 @@ #define JL_HAS_DRAW_CIRCLE // st planar circle, function #define JL_HAS_FILL_CIRCLE // st planar fill circle, function -- returns bool #define JL_HAS_FILL_RECT // st planar fill rect, function - #define JL_HAS_TILE_FILL // st planar tile fill, function - #define JL_HAS_TILE_COPY // st planar tile copy, function - #define JL_HAS_TILE_COPY_MASKED // st planar masked copy, function - #define JL_HAS_TILE_PASTE // st planar tile paste, function - #define JL_HAS_TILE_SNAP // st planar tile snap, function - #define JL_HAS_TILE_PASTE_MONO // st planar mono colorize+paste, function + #define JL_HAS_TILE_FILL // st planar tile fill, inline (stTile.h) + #define JL_HAS_TILE_COPY // st planar tile copy, inline (stTile.h) + #define JL_HAS_TILE_COPY_MASKED // st planar masked copy, inline (stTile.h) + #define JL_HAS_TILE_PASTE // st planar tile paste, inline (stTile.h) + #define JL_HAS_TILE_SNAP // st planar tile snap, inline (stTile.h) + #define JL_HAS_TILE_PASTE_MONO // st planar mono colorize+paste, inline (stTile.h) #define JL_HAS_SPRITE_DRAW // st planar sprite draw, function #define JL_HAS_SPRITE_SAVE // st planar sprite save, function #define JL_HAS_SPRITE_RESTORE // st planar sprite restore, function diff --git a/src/amiga/amigaTile.h b/src/amiga/amigaTile.h new file mode 100644 index 0000000..e330c04 --- /dev/null +++ b/src/amiga/amigaTile.h @@ -0,0 +1,261 @@ +// Amiga tile ops as always-inline functions (NATIVE-PERF Phase 2 +// R4a). Bodies moved verbatim from src/amiga/hal.c so the public +// wrappers in src/core/tile.c inline them, deleting the second full +// cdecl call layer (~130-190 cycles/op on the 68000). The jlpTile* +// dispatch names are macro-aliased to these in port.h, which also +// #includes this header at its tail -- AFTER the jlpGenericTile* +// prototypes the portData == NULL fallbacks call. Do not include this +// header directly; it comes with port.h on the Amiga. +#ifndef JOEY_AMIGA_TILE_H +#define JOEY_AMIGA_TILE_H + +#include +#include + +#include "joey/tile.h" +#include "amigaPlanar.h" + +// 8-row unrolled tile access (NATIVE-PERF Phase 2 R8): row offsets +// are compile-time constants row * 40, so each row is one d16(An) +// access with no pointer stepping and no loop control. gcc-amigaos +// -O2 keeps the equivalent row loops rolled as dbf (objdump-verified, +// ~18 cycles/row of pure loop overhead). Only the row dimension +// unrolls -- the 4-plane loop keeps real per-plane setup and stays +// register-light per the m68k-gcc-unroll-backfire lesson. +#define AMIGA_TILE_ROW_OFF(_r) ((uint16_t)((_r) * AMIGA_BYTES_PER_ROW)) + +// Phase 5 planar dual-write for tile ops, fully planar after Phase 9 +// dropped the chunky shadow. All tiles are 8-pixel aligned (8x8 blocks +// at multiples of 8), so plane writes are byte-aligned -- one plane +// byte per row, 8 rows per tile, no edge masks. Stride between rows +// in a plane is AMIGA_BYTES_PER_ROW (40). + +static inline __attribute__((always_inline)) void amigaTileFill(jlSurfaceT *s, uint8_t bx, uint8_t by, uint8_t colorIndex) { + AmigaPlanarT *pd; + uint16_t plane; + uint8_t fillByte; + uint8_t *p; + + pd = (AmigaPlanarT *)s->portData; + if (pd == NULL) { + jlpGenericTileFill(s, bx, by, colorIndex); + return; + } + for (plane = 0; plane < AMIGA_BITPLANES; plane++) { + fillByte = ((colorIndex >> plane) & 1u) ? 0xFFu : 0x00u; + p = pd->planes[plane] + (uint16_t)by * 8u * AMIGA_BYTES_PER_ROW + bx; + p[AMIGA_TILE_ROW_OFF(0)] = fillByte; + p[AMIGA_TILE_ROW_OFF(1)] = fillByte; + p[AMIGA_TILE_ROW_OFF(2)] = fillByte; + p[AMIGA_TILE_ROW_OFF(3)] = fillByte; + p[AMIGA_TILE_ROW_OFF(4)] = fillByte; + p[AMIGA_TILE_ROW_OFF(5)] = fillByte; + p[AMIGA_TILE_ROW_OFF(6)] = fillByte; + p[AMIGA_TILE_ROW_OFF(7)] = fillByte; + } +} + + +static inline __attribute__((always_inline)) void amigaTileCopy(jlSurfaceT *dst, uint8_t dstBx, uint8_t dstBy, const jlSurfaceT *src, uint8_t srcBx, uint8_t srcBy) { + AmigaPlanarT *dstPd; + AmigaPlanarT *srcPd; + uint16_t plane; + uint8_t *dp; + const uint8_t *sp; + + dstPd = (AmigaPlanarT *)dst->portData; + srcPd = (AmigaPlanarT *)src->portData; + if (dstPd == NULL || srcPd == NULL) { + jlpGenericTileCopy(dst, dstBx, dstBy, src, srcBx, srcBy); + return; + } + for (plane = 0; plane < AMIGA_BITPLANES; plane++) { + dp = dstPd->planes[plane] + (uint16_t)dstBy * 8u * AMIGA_BYTES_PER_ROW + dstBx; + sp = srcPd->planes[plane] + (uint16_t)srcBy * 8u * AMIGA_BYTES_PER_ROW + srcBx; + dp[AMIGA_TILE_ROW_OFF(0)] = sp[AMIGA_TILE_ROW_OFF(0)]; + dp[AMIGA_TILE_ROW_OFF(1)] = sp[AMIGA_TILE_ROW_OFF(1)]; + dp[AMIGA_TILE_ROW_OFF(2)] = sp[AMIGA_TILE_ROW_OFF(2)]; + dp[AMIGA_TILE_ROW_OFF(3)] = sp[AMIGA_TILE_ROW_OFF(3)]; + dp[AMIGA_TILE_ROW_OFF(4)] = sp[AMIGA_TILE_ROW_OFF(4)]; + dp[AMIGA_TILE_ROW_OFF(5)] = sp[AMIGA_TILE_ROW_OFF(5)]; + dp[AMIGA_TILE_ROW_OFF(6)] = sp[AMIGA_TILE_ROW_OFF(6)]; + dp[AMIGA_TILE_ROW_OFF(7)] = sp[AMIGA_TILE_ROW_OFF(7)]; + } +} + + +// Pure-planar masked copy. For each row of the 8x8 tile, read 4 src +// plane bytes; compute a per-pixel "non-transparent" mask via XOR +// against the transparent index's per-plane bit pattern (a pixel +// matches transparent IFF all 4 plane bits match transparent's 4 +// bits = OR of 4 XOR'd bytes is 0 in that bit). Then for each plane, +// dst = (dst & ~mask) | (src & mask) -- copy src bits at mask-set +// positions, preserve dst bits elsewhere. +static inline __attribute__((always_inline)) void amigaTileCopyMasked(jlSurfaceT *dst, uint8_t dstBx, uint8_t dstBy, const jlSurfaceT *src, uint8_t srcBx, uint8_t srcBy, uint8_t transparent) { + AmigaPlanarT *dstPd; + AmigaPlanarT *srcPd; + uint8_t transBitByte[AMIGA_BITPLANES]; + uint16_t i; + uint8_t row; + uint16_t srcByteOff; + uint16_t dstByteOff; + uint8_t srcPlaneBytes[AMIGA_BITPLANES]; + uint8_t maskByte; + + dstPd = (AmigaPlanarT *)dst->portData; + srcPd = (AmigaPlanarT *)src->portData; + if (dstPd == NULL || srcPd == NULL) { + jlpGenericTileCopyMasked(dst, dstBx, dstBy, src, srcBx, srcBy, transparent); + return; + } + transparent = (uint8_t)(transparent & 0x0Fu); + /* Per-plane "all bits set if transparent's bit at this plane is 1 + * else all 0" -- so XOR gives bit set where pixel differs from + * transparent in that plane. */ + for (i = 0; i < AMIGA_BITPLANES; i++) { + transBitByte[i] = ((transparent >> i) & 1u) ? 0xFFu : 0x00u; + } + + for (row = 0; row < 8u; row++) { + srcByteOff = (uint16_t)((uint16_t)srcBy * 8u + row) * AMIGA_BYTES_PER_ROW + srcBx; + dstByteOff = (uint16_t)((uint16_t)dstBy * 8u + row) * AMIGA_BYTES_PER_ROW + dstBx; + srcPlaneBytes[0] = srcPd->planes[0][srcByteOff]; + srcPlaneBytes[1] = srcPd->planes[1][srcByteOff]; + srcPlaneBytes[2] = srcPd->planes[2][srcByteOff]; + srcPlaneBytes[3] = srcPd->planes[3][srcByteOff]; + /* maskByte: bit set where pixel differs from transparent in + * ANY plane -- i.e., where the pixel is non-transparent. */ + maskByte = (uint8_t)((srcPlaneBytes[0] ^ transBitByte[0]) + | (srcPlaneBytes[1] ^ transBitByte[1]) + | (srcPlaneBytes[2] ^ transBitByte[2]) + | (srcPlaneBytes[3] ^ transBitByte[3])); + for (i = 0; i < AMIGA_BITPLANES; i++) { + uint8_t existing = dstPd->planes[i][dstByteOff]; + dstPd->planes[i][dstByteOff] = (uint8_t)((existing & (uint8_t)~maskByte) + | (srcPlaneBytes[i] & maskByte)); + } + } +} + + +// Phase 9 plane-to-chunky derivation for jlTileSnap. Reads 8 plane +// bytes (1 byte per row x 8 rows) from each of 4 planes for the +// 8-pixel-aligned tile column at bx, then assembles 32 chunky bytes +// (4 per row x 8 rows, packed 2 px/byte high-then-low nibble) into +// On Amiga, jlTileT.pixels is opaque port-specific storage (cross- +// platform tile.c never reads it directly when s->pixels is NULL). +// We use those 32 bytes as 4 planes x 8 rows, plane-major: +// bytes [0..7] = plane 0, rows 0..7 +// bytes [8..15] = plane 1, rows 0..7 +// bytes [16..23] = plane 2, rows 0..7 +// bytes [24..31] = plane 3, rows 0..7 +// snap/paste then become 32 plain byte loads + stores -- no chunky +// <-> planar conversion at all. The previous c2p-based path paid +// 4 KB LUT lookups + bit shuffling per pixel; this is ~50x cheaper. +#define AMIGA_TILE_PLANE_STRIDE 8 +static inline __attribute__((always_inline)) void amigaTileSnap(const jlSurfaceT *src, uint8_t bx, uint8_t by, uint8_t *chunkyOut) { + AmigaPlanarT *pd; + uint16_t rowBase; + uint8_t plane; + + pd = (AmigaPlanarT *)src->portData; + if (pd == NULL) { + jlpGenericTileSnap(src, bx, by, chunkyOut); + return; + } + rowBase = (uint16_t)((uint16_t)by * 8u) * AMIGA_BYTES_PER_ROW + bx; + for (plane = 0; plane < AMIGA_BITPLANES; plane++) { + const uint8_t *p = pd->planes[plane] + rowBase; + uint8_t *q = chunkyOut + plane * AMIGA_TILE_PLANE_STRIDE; + q[0] = p[AMIGA_TILE_ROW_OFF(0)]; + q[1] = p[AMIGA_TILE_ROW_OFF(1)]; + q[2] = p[AMIGA_TILE_ROW_OFF(2)]; + q[3] = p[AMIGA_TILE_ROW_OFF(3)]; + q[4] = p[AMIGA_TILE_ROW_OFF(4)]; + q[5] = p[AMIGA_TILE_ROW_OFF(5)]; + q[6] = p[AMIGA_TILE_ROW_OFF(6)]; + q[7] = p[AMIGA_TILE_ROW_OFF(7)]; + } +} + + +static inline __attribute__((always_inline)) void amigaTilePaste(jlSurfaceT *dst, uint8_t bx, uint8_t by, const uint8_t *chunkyTile) { + AmigaPlanarT *pd; + uint8_t plane; + uint16_t rowBase; + + pd = (AmigaPlanarT *)dst->portData; + if (pd == NULL) { + jlpGenericTilePaste(dst, bx, by, chunkyTile); + return; + } + /* jlTileT.pixels[] holds plane-major planar bytes (see + * jlpTileSnap header for layout). Paste = 32 byte stores + * with no chunky -> planar conversion. */ + rowBase = (uint16_t)((uint16_t)by * 8u) * AMIGA_BYTES_PER_ROW + bx; + for (plane = 0; plane < AMIGA_BITPLANES; plane++) { + const uint8_t *q = chunkyTile + plane * AMIGA_TILE_PLANE_STRIDE; + uint8_t *p = pd->planes[plane] + rowBase; + p[AMIGA_TILE_ROW_OFF(0)] = q[0]; + p[AMIGA_TILE_ROW_OFF(1)] = q[1]; + p[AMIGA_TILE_ROW_OFF(2)] = q[2]; + p[AMIGA_TILE_ROW_OFF(3)] = q[3]; + p[AMIGA_TILE_ROW_OFF(4)] = q[4]; + p[AMIGA_TILE_ROW_OFF(5)] = q[5]; + p[AMIGA_TILE_ROW_OFF(6)] = q[6]; + p[AMIGA_TILE_ROW_OFF(7)] = q[7]; + } +} + + +// Planar monochrome paste. monoTile follows the cross-port mono +// contract (include/joey/tile.h + jlpGenericTilePasteMono): 32 chunky +// nibble-pair bytes, row-major, TILE_BYTES_PER_ROW bytes per row, two +// pixels per byte. A pixel renders fgColor when its source nibble is +// nonzero, bgColor when zero; the HIGH nibble is the LEFT pixel. Each +// row's 4 source bytes fold into an 8-bit shape mask (bit 7 = +// leftmost pixel, matching the planar bit order), then each plane k +// writes +// outPlaneK = (shape & maskFgK) | (~shape & maskBgK) +// where maskXK = $FF if (X & (1 << k)) else $00. +static inline __attribute__((always_inline)) void amigaTilePasteMono(jlSurfaceT *dst, uint8_t bx, uint8_t by, const uint8_t *monoTile, uint8_t fgColor, uint8_t bgColor) { + AmigaPlanarT *pd; + uint8_t shape[TILE_PIXELS_PER_SIDE]; + uint8_t plane; + uint8_t row; + uint8_t col; + uint16_t rowBase; + + pd = (AmigaPlanarT *)dst->portData; + if (pd == NULL) { + jlpGenericTilePasteMono(dst, bx, by, monoTile, fgColor, bgColor); + return; + } + // Fold each chunky nibble-pair row into one 1bpp shape byte. + for (row = 0; row < TILE_PIXELS_PER_SIDE; row++) { + uint8_t bits = 0u; + for (col = 0; col < TILE_BYTES_PER_ROW; col++) { + uint8_t srcByte = *monoTile++; + bits = (uint8_t)(bits << 2); + if (srcByte & 0xF0u) { + bits = (uint8_t)(bits | 0x02u); + } + if (srcByte & 0x0Fu) { + bits = (uint8_t)(bits | 0x01u); + } + } + shape[row] = bits; + } + rowBase = (uint16_t)((uint16_t)by * 8u) * AMIGA_BYTES_PER_ROW + bx; + for (plane = 0; plane < AMIGA_BITPLANES; plane++) { + uint8_t maskFg = (uint8_t)((fgColor & (1u << plane)) ? 0xFFu : 0x00u); + uint8_t maskBg = (uint8_t)((bgColor & (1u << plane)) ? 0xFFu : 0x00u); + uint8_t *p = pd->planes[plane] + rowBase; + for (row = 0; row < 8u; row++) { + uint8_t s = shape[row]; + p[row * AMIGA_BYTES_PER_ROW] = (uint8_t)((uint8_t)(s & maskFg) | (uint8_t)((uint8_t)(~s) & maskBg)); + } + } +} + +#endif diff --git a/src/amiga/hal.c b/src/amiga/hal.c index 7be3626..889ea09 100644 --- a/src/amiga/hal.c +++ b/src/amiga/hal.c @@ -128,15 +128,11 @@ static bool gPrevAnyDirty = true; // demo after the initial paint) leave both alone. MrgCop + LoadView + // WaitTOF is hundreds of milliseconds on a 7 MHz 68000, so skipping // them on clean frames is a major win. -/* Long-aligned so memcmpLongs (uint32_t pointer compare) won't - * address-error on 68000. jlSurfaceT.scb/palette are at long-aligned - * offsets (8 and 208), so the source side is safe; these cached - * mirrors need the same property. */ -static uint8_t gCachedScb [SURFACE_HEIGHT] __attribute__((aligned(4))); -static uint16_t gCachedPalette[SURFACE_PALETTE_COUNT][SURFACE_COLORS_PER_PALETTE] __attribute__((aligned(4))); -static bool gCacheValid = false; - -static bool paletteOrScbChanged(const jlSurfaceT *src); +/* Change detection is the core gStageScbDirty / gStagePaletteDirty + * flags (surfaceInternal.h): the mutation APIs set them and + * updateCopperIfNeeded clears them after rebuilding, replacing a + * 712-byte cached-mirror long-compare per present (NATIVE-PERF + * Phase 2 R9; the IIgs consumed the flags first). */ // ----- Internal helpers (alphabetical) ----- @@ -453,63 +449,25 @@ static void patchBplCopperPokes(UBYTE *planes[AMIGA_BITPLANES]) { } -/* Long-aligned compare. SCB is 200 bytes, palette is 16*16*2 = 512 - * bytes; both gCached* are statically aligned and src->scb/palette - * sit at long-aligned offsets in jlSurfaceT. libnix memcmp walks - * byte-by-byte (~10 cyc/byte = ~7 ms for 712 bytes); a long-pointer - * inline compare drops that to ~2 ms, which dominates per-call - * overhead for tight stagePresentRect loops where there's nothing - * to actually present. Returns nonzero on first mismatch. */ -static bool memcmpLongs(const void *a, const void *b, uint16_t bytes) { - const uint32_t *pa = (const uint32_t *)a; - const uint32_t *pb = (const uint32_t *)b; - uint16_t n = (uint16_t)(bytes >> 2); - while (n > 0u) { - if (*pa != *pb) { - return true; - } - pa++; - pb++; - n--; - } - return false; -} - - -// Returns true if the SCB table or palette RGB values differ from the -// last presented frame, or if no frame has been presented yet. -static bool paletteOrScbChanged(const jlSurfaceT *src) { - if (!gCacheValid) { - return true; - } - if (memcmpLongs(gCachedScb, src->scb, (uint16_t)sizeof(gCachedScb))) { - return true; - } - if (memcmpLongs(gCachedPalette, src->palette, (uint16_t)sizeof(gCachedPalette))) { - return true; - } - return false; -} - - // Rebuild and install the user copper list only if the palette/SCB // state visible to the display differs from what the surface carries // now. On clean frames we skip the AllocMem + MrgCop + LoadView + -// WaitTOF chain entirely. +// WaitTOF chain entirely. `src` is always gStage (jlpPresent's +// contract), so the core dirty flags are the change signal; clear +// them only after rebuilding. static uint32_t gCopperRebuildCount = 0; static uint32_t gPresentCallCount = 0; static void updateCopperIfNeeded(const jlSurfaceT *src) { - if (!paletteOrScbChanged(src)) { + if (!gStageScbDirty && !gStagePaletteDirty) { return; } gCopperRebuildCount++; uploadFirstBandPalette(src); buildCopperList(src); installCopperList(); - memcpy(gCachedScb, src->scb, sizeof(gCachedScb)); - memcpy(gCachedPalette, src->palette, sizeof(gCachedPalette)); - gCacheValid = true; + gStageScbDirty = false; + gStagePaletteDirty = false; } @@ -1143,115 +1101,6 @@ void jlpFillRect(jlSurfaceT *s, int16_t x, int16_t y, int16_t w, int16_t h, uint } -// Phase 5 planar dual-write for tile ops, fully planar after Phase 9 -// dropped the chunky shadow. All tiles are 8-pixel aligned (8x8 blocks -// at multiples of 8), so plane writes are byte-aligned -- one plane -// byte per row, 8 rows per tile, no edge masks. Stride between rows -// in a plane is AMIGA_BYTES_PER_ROW (40). - -void jlpTileFill(jlSurfaceT *s, uint8_t bx, uint8_t by, uint8_t colorIndex) { - AmigaPlanarT *pd; - uint16_t plane; - uint8_t fillByte; - uint8_t *p; - uint8_t row; - - pd = (AmigaPlanarT *)s->portData; - if (pd == NULL) { - jlpGenericTileFill(s, bx, by, colorIndex); - return; - } - for (plane = 0; plane < AMIGA_BITPLANES; plane++) { - fillByte = ((colorIndex >> plane) & 1u) ? 0xFFu : 0x00u; - p = pd->planes[plane] + (uint16_t)by * 8u * AMIGA_BYTES_PER_ROW + bx; - for (row = 0; row < 8u; row++) { - *p = fillByte; - p += AMIGA_BYTES_PER_ROW; - } - } -} - - -void jlpTileCopy(jlSurfaceT *dst, uint8_t dstBx, uint8_t dstBy, const jlSurfaceT *src, uint8_t srcBx, uint8_t srcBy) { - AmigaPlanarT *dstPd; - AmigaPlanarT *srcPd; - uint16_t plane; - uint8_t *dp; - const uint8_t *sp; - uint8_t row; - - dstPd = (AmigaPlanarT *)dst->portData; - srcPd = (AmigaPlanarT *)src->portData; - if (dstPd == NULL || srcPd == NULL) { - jlpGenericTileCopy(dst, dstBx, dstBy, src, srcBx, srcBy); - return; - } - for (plane = 0; plane < AMIGA_BITPLANES; plane++) { - dp = dstPd->planes[plane] + (uint16_t)dstBy * 8u * AMIGA_BYTES_PER_ROW + dstBx; - sp = srcPd->planes[plane] + (uint16_t)srcBy * 8u * AMIGA_BYTES_PER_ROW + srcBx; - for (row = 0; row < 8u; row++) { - *dp = *sp; - dp += AMIGA_BYTES_PER_ROW; - sp += AMIGA_BYTES_PER_ROW; - } - } -} - - -// Pure-planar masked copy. For each row of the 8x8 tile, read 4 src -// plane bytes; compute a per-pixel "non-transparent" mask via XOR -// against the transparent index's per-plane bit pattern (a pixel -// matches transparent IFF all 4 plane bits match transparent's 4 -// bits = OR of 4 XOR'd bytes is 0 in that bit). Then for each plane, -// dst = (dst & ~mask) | (src & mask) -- copy src bits at mask-set -// positions, preserve dst bits elsewhere. -void jlpTileCopyMasked(jlSurfaceT *dst, uint8_t dstBx, uint8_t dstBy, const jlSurfaceT *src, uint8_t srcBx, uint8_t srcBy, uint8_t transparent) { - AmigaPlanarT *dstPd; - AmigaPlanarT *srcPd; - uint8_t transBitByte[AMIGA_BITPLANES]; - uint16_t i; - uint8_t row; - uint16_t srcByteOff; - uint16_t dstByteOff; - uint8_t srcPlaneBytes[AMIGA_BITPLANES]; - uint8_t maskByte; - - dstPd = (AmigaPlanarT *)dst->portData; - srcPd = (AmigaPlanarT *)src->portData; - if (dstPd == NULL || srcPd == NULL) { - jlpGenericTileCopyMasked(dst, dstBx, dstBy, src, srcBx, srcBy, transparent); - return; - } - transparent = (uint8_t)(transparent & 0x0Fu); - /* Per-plane "all bits set if transparent's bit at this plane is 1 - * else all 0" -- so XOR gives bit set where pixel differs from - * transparent in that plane. */ - for (i = 0; i < AMIGA_BITPLANES; i++) { - transBitByte[i] = ((transparent >> i) & 1u) ? 0xFFu : 0x00u; - } - - for (row = 0; row < 8u; row++) { - srcByteOff = (uint16_t)((uint16_t)srcBy * 8u + row) * AMIGA_BYTES_PER_ROW + srcBx; - dstByteOff = (uint16_t)((uint16_t)dstBy * 8u + row) * AMIGA_BYTES_PER_ROW + dstBx; - srcPlaneBytes[0] = srcPd->planes[0][srcByteOff]; - srcPlaneBytes[1] = srcPd->planes[1][srcByteOff]; - srcPlaneBytes[2] = srcPd->planes[2][srcByteOff]; - srcPlaneBytes[3] = srcPd->planes[3][srcByteOff]; - /* maskByte: bit set where pixel differs from transparent in - * ANY plane -- i.e., where the pixel is non-transparent. */ - maskByte = (uint8_t)((srcPlaneBytes[0] ^ transBitByte[0]) - | (srcPlaneBytes[1] ^ transBitByte[1]) - | (srcPlaneBytes[2] ^ transBitByte[2]) - | (srcPlaneBytes[3] ^ transBitByte[3])); - for (i = 0; i < AMIGA_BITPLANES; i++) { - uint8_t existing = dstPd->planes[i][dstByteOff]; - dstPd->planes[i][dstByteOff] = (uint8_t)((existing & (uint8_t)~maskByte) - | (srcPlaneBytes[i] & maskByte)); - } - } -} - - // Phase 9 sprite save/restore plane-backup hooks. // Sprite save at (x, y, w, h) records the destination region into the // caller's backup buffer; restore writes it back. x and w are 2-pixel @@ -1581,119 +1430,6 @@ void jlpSpriteDrawPlanes(jlSurfaceT *s, const jlSpriteT *sp, int16_t x, int16_t } -// Phase 9 plane-to-chunky derivation for jlTileSnap. Reads 8 plane -// bytes (1 byte per row x 8 rows) from each of 4 planes for the -// 8-pixel-aligned tile column at bx, then assembles 32 chunky bytes -// (4 per row x 8 rows, packed 2 px/byte high-then-low nibble) into -// On Amiga, jlTileT.pixels is opaque port-specific storage (cross- -// platform tile.c never reads it directly when s->pixels is NULL). -// We use those 32 bytes as 4 planes x 8 rows, plane-major: -// bytes [0..7] = plane 0, rows 0..7 -// bytes [8..15] = plane 1, rows 0..7 -// bytes [16..23] = plane 2, rows 0..7 -// bytes [24..31] = plane 3, rows 0..7 -// snap/paste then become 32 plain byte loads + stores -- no chunky -// <-> planar conversion at all. The previous c2p-based path paid -// 4 KB LUT lookups + bit shuffling per pixel; this is ~50x cheaper. -#define AMIGA_TILE_PLANE_STRIDE 8 -void jlpTileSnap(const jlSurfaceT *src, uint8_t bx, uint8_t by, uint8_t *chunkyOut) { - AmigaPlanarT *pd; - uint16_t row; - uint16_t rowBase; - uint8_t plane; - - pd = (AmigaPlanarT *)src->portData; - if (pd == NULL) { - jlpGenericTileSnap(src, bx, by, chunkyOut); - return; - } - rowBase = (uint16_t)((uint16_t)by * 8u) * AMIGA_BYTES_PER_ROW + bx; - for (plane = 0; plane < AMIGA_BITPLANES; plane++) { - const uint8_t *p = pd->planes[plane] + rowBase; - uint8_t *q = chunkyOut + plane * AMIGA_TILE_PLANE_STRIDE; - for (row = 0; row < 8u; row++) { - q[row] = p[row * AMIGA_BYTES_PER_ROW]; - } - } -} - - -void jlpTilePaste(jlSurfaceT *dst, uint8_t bx, uint8_t by, const uint8_t *chunkyTile) { - AmigaPlanarT *pd; - uint8_t row; - uint8_t plane; - uint16_t rowBase; - - pd = (AmigaPlanarT *)dst->portData; - if (pd == NULL) { - jlpGenericTilePaste(dst, bx, by, chunkyTile); - return; - } - /* jlTileT.pixels[] holds plane-major planar bytes (see - * jlpTileSnap header for layout). Paste = 32 byte stores - * with no chunky -> planar conversion. */ - rowBase = (uint16_t)((uint16_t)by * 8u) * AMIGA_BYTES_PER_ROW + bx; - for (plane = 0; plane < AMIGA_BITPLANES; plane++) { - const uint8_t *q = chunkyTile + plane * AMIGA_TILE_PLANE_STRIDE; - uint8_t *p = pd->planes[plane] + rowBase; - for (row = 0; row < 8u; row++) { - p[row * AMIGA_BYTES_PER_ROW] = q[row]; - } - } -} - - -// Planar monochrome paste. monoTile follows the cross-port mono -// contract (include/joey/tile.h + jlpGenericTilePasteMono): 32 chunky -// nibble-pair bytes, row-major, TILE_BYTES_PER_ROW bytes per row, two -// pixels per byte. A pixel renders fgColor when its source nibble is -// nonzero, bgColor when zero; the HIGH nibble is the LEFT pixel. Each -// row's 4 source bytes fold into an 8-bit shape mask (bit 7 = -// leftmost pixel, matching the planar bit order), then each plane k -// writes -// outPlaneK = (shape & maskFgK) | (~shape & maskBgK) -// where maskXK = $FF if (X & (1 << k)) else $00. -void jlpTilePasteMono(jlSurfaceT *dst, uint8_t bx, uint8_t by, const uint8_t *monoTile, uint8_t fgColor, uint8_t bgColor) { - AmigaPlanarT *pd; - uint8_t shape[TILE_PIXELS_PER_SIDE]; - uint8_t plane; - uint8_t row; - uint8_t col; - uint16_t rowBase; - - pd = (AmigaPlanarT *)dst->portData; - if (pd == NULL) { - jlpGenericTilePasteMono(dst, bx, by, monoTile, fgColor, bgColor); - return; - } - // Fold each chunky nibble-pair row into one 1bpp shape byte. - for (row = 0; row < TILE_PIXELS_PER_SIDE; row++) { - uint8_t bits = 0u; - for (col = 0; col < TILE_BYTES_PER_ROW; col++) { - uint8_t srcByte = *monoTile++; - bits = (uint8_t)(bits << 2); - if (srcByte & 0xF0u) { - bits = (uint8_t)(bits | 0x02u); - } - if (srcByte & 0x0Fu) { - bits = (uint8_t)(bits | 0x01u); - } - } - shape[row] = bits; - } - rowBase = (uint16_t)((uint16_t)by * 8u) * AMIGA_BYTES_PER_ROW + bx; - for (plane = 0; plane < AMIGA_BITPLANES; plane++) { - uint8_t maskFg = (uint8_t)((fgColor & (1u << plane)) ? 0xFFu : 0x00u); - uint8_t maskBg = (uint8_t)((bgColor & (1u << plane)) ? 0xFFu : 0x00u); - uint8_t *p = pd->planes[plane] + rowBase; - for (row = 0; row < 8u; row++) { - uint8_t s = shape[row]; - p[row * AMIGA_BYTES_PER_ROW] = (uint8_t)((uint8_t)(s & maskFg) | (uint8_t)((uint8_t)(~s) & maskBg)); - } - } -} - - // Phase 3 planar dual-write for jlSurfaceCopy: 4 plane memcpys after // the cross-platform chunky pixel memcpy. Both src and dst planes // are off-screen shadow buffers; the displayed gPlanes[] is updated diff --git a/src/atarist/hal.c b/src/atarist/hal.c index 380e427..2f0c963 100644 --- a/src/atarist/hal.c +++ b/src/atarist/hal.c @@ -44,9 +44,9 @@ // Word-interleaved ST planar uses the same 160 bytes/scanline as our // chunky source, but organized as 20 groups of 4 words per scanline, // with each word holding the 16 one-bit samples for one bitplane. -#define ST_BYTES_PER_ROW 160 +// (ST_BYTES_PER_ROW and ST_BYTES_PER_GROUP live in stPlanar.h so the +// stTile.h inlines can see them; the rest are hal-private.) #define ST_GROUPS_PER_ROW 20 -#define ST_BYTES_PER_GROUP 8 // 4 plane words back-to-back #define ST_PLANE_OFF_BYTES 2 // step between adjacent plane words within a group #define ST_BITPLANES 4 #define ST_PLANAR_SIZE (ST_BYTES_PER_ROW * SURFACE_HEIGHT) @@ -184,7 +184,6 @@ static long writePrevPaletteRegs(void); static __attribute__((interrupt_handler)) void timerBIsr(void); static __attribute__((interrupt_handler)) void vblIsr(void); static void buildTransitions(const jlSurfaceT *src); -static bool paletteOrScbChanged(const jlSurfaceT *src); static void refreshPaletteStateIfNeeded(const jlSurfaceT *src); // ----- Module state ----- @@ -271,15 +270,14 @@ static volatile uint16_t gLastBandCount = 0; static void (*gOldVblVec)(void) = NULL; static void (*gOldTimerBVec)(void) = NULL; -// Cached SCB + palette from the last present. flattenScbPalettes runs -// 200 * 16 quantize conversions and buildTransitions rescans the full -// SCB; neither is cheap on a 7 MHz 68000. In the typical game loop -// (and every frame of the keys demo after the initial paint) SCB and -// palette never change, so caching and skipping those passes keeps -// rect presents down to just the screen blit. -static uint8_t gCachedScb [SURFACE_HEIGHT]; -static uint16_t gCachedPalette[SURFACE_PALETTE_COUNT][SURFACE_COLORS_PER_PALETTE]; -static bool gCacheValid = false; +// Palette/SCB change detection is the core gStageScbDirty / +// gStagePaletteDirty flags (surfaceInternal.h) -- the mutation APIs +// set them and refreshPaletteStateIfNeeded clears them after +// rebuilding, replacing a 712-byte cached-mirror memcmp per present +// (NATIVE-PERF Phase 2 R9; the IIgs consumed the flags first). +// flattenScbPalettes runs 200 * 16 quantize conversions and +// buildTransitions rescans the full SCB; neither is cheap on a 7 MHz +// 68000, and in the typical game loop SCB and palette never change. // ----- Internal helpers (alphabetical) ----- @@ -333,34 +331,20 @@ static void flattenScbPalettes(const jlSurfaceT *src) { } -// Returns true if SCB or palette values differ from the last present. -static bool paletteOrScbChanged(const jlSurfaceT *src) { - if (!gCacheValid) { - return true; - } - if (memcmp(gCachedScb, src->scb, sizeof(gCachedScb)) != 0) { - return true; - } - if (memcmp(gCachedPalette, src->palette, sizeof(gCachedPalette)) != 0) { - return true; - } - return false; -} - - // Rebuild the per-line palette table and band-transition table only // when the SCB/palette state has actually changed. Both are hot -- the // flatten pass runs 3200 palette entries through quantization -- so -// skipping them on clean frames dominates rect-present timing. +// skipping them on clean frames dominates rect-present timing. `src` +// is always gStage (jlpPresent's contract), so the core dirty flags +// are the change signal; clear them only after rebuilding. static void refreshPaletteStateIfNeeded(const jlSurfaceT *src) { - if (!paletteOrScbChanged(src)) { + if (!gStageScbDirty && !gStagePaletteDirty) { return; } flattenScbPalettes(src); buildTransitions(src); - memcpy(gCachedScb, src->scb, sizeof(gCachedScb)); - memcpy(gCachedPalette, src->palette, sizeof(gCachedPalette)); - gCacheValid = true; + gStageScbDirty = false; + gStagePaletteDirty = false; } @@ -961,7 +945,6 @@ extern void surface68kStFillCircle(uint8_t *base, uint16_t cx, uint16_t cy, uint extern void surface68kStFillRectSingleGroup(uint8_t *firstGroupPtr, uint16_t mask, uint16_t h, uint8_t color); extern void surface68kStFillRectMulti(uint8_t *base, int16_t x, int16_t y, uint16_t w, uint16_t h, uint8_t color); extern void surface68kStLongFill(uint8_t *dst, uint16_t numGroups, uint32_t loLong, uint32_t hiLong); -extern void surface68kStTileFill8x8(uint8_t *firstGroupPtr, uint16_t mask, uint8_t color); extern void surface68kStSprite16x16Save(uint8_t *base, uint16_t x, uint16_t y, uint8_t *dstBuf); extern void surface68kStSprite16x16Restore(uint8_t *base, uint16_t x, uint16_t y, const uint8_t *srcBuf); extern void surface68kStSpriteSaveByteAligned(uint8_t *base, uint16_t x, uint16_t y, uint16_t w, uint16_t h, uint8_t *dstPlaneBytes); @@ -1344,300 +1327,6 @@ void jlpSurfaceCopyPlanes(jlSurfaceT *dst, const jlSurfaceT *src) { // every callsite folds the 4-plane RMW into its body. -// Phase 5 tile ops. 8x8 tiles at byte position (bx, by) start at -// pixel (bx*8, by*8). 8 pixels wide always covers exactly half a -// 16-pixel group: high half (bits 15..8) when bx is even, low half -// (bits 7..0) when bx is odd. Per-row work is 4 plane half-word RMWs. -void jlpTileFill(jlSurfaceT *s, uint8_t bx, uint8_t by, uint8_t colorIndex) { - StPlanarT *pd; - uint16_t group; - uint16_t halfMask; - uint8_t *gp; - - // Core jlTileFill already proved s != NULL and clipped bx/by; only - // the non-planar (offscreen, no portData) fallback needs handling. - pd = (StPlanarT *)s->portData; - if (pd == NULL) { - jlpGenericTileFill(s, bx, by, colorIndex); - return; - } - group = (uint16_t)((uint16_t)bx >> 1); - halfMask = ((bx & 1u) == 0u) ? 0xFF00u : 0x00FFu; - gp = pd->base + (uint16_t)by * 8u * ST_BYTES_PER_ROW + group * ST_BYTES_PER_GROUP; - /* Phase 10 final: specialized 8x8 unrolled tile-fill skips the - * generic FRG_LOOP's per-row subq+bne overhead. */ - surface68kStTileFill8x8(gp, halfMask, colorIndex); -} - - -// Phase 10: tile paste/snap reuse the asm sprite save/restore -// helpers -- identical per-row work patterns at byte-aligned -// positions. Width 8 = single tile column = single half-group -// write per plane. The asm walker handles 8 rows just as well -// as a sprite's variable height. -void jlpTilePaste(jlSurfaceT *dst, uint8_t bx, uint8_t by, const uint8_t *chunkyTile) { - StPlanarT *pd; - uint16_t group; - uint8_t *dstAddr; - int16_t row; - - if (dst->portData == NULL) { - jlpGenericTilePaste(dst, bx, by, chunkyTile); - return; - } - pd = (StPlanarT *)dst->portData; - /* Phase 10.5: jlTileT.pixels holds plane-major bytes (4 plane bytes - * per row * 8 rows = 32 bytes). Direct byte copy to the planar - * buffer; no chunky <-> planar conversion. Mirrors the sibling - * jlpTileCopy pattern but reads from the contiguous tile - * buffer. Drops the asm-walker entry/exit overhead. */ - group = (uint16_t)((uint16_t)bx >> 1); - dstAddr = pd->base - + (uint16_t)by * 8u * ST_BYTES_PER_ROW - + group * ST_BYTES_PER_GROUP - + (uint16_t)(bx & 1u); - (void)row; -#define ST_TILE_PASTE_ROW \ - do { \ - dstAddr[0] = chunkyTile[0]; \ - dstAddr[2] = chunkyTile[1]; \ - dstAddr[4] = chunkyTile[2]; \ - dstAddr[6] = chunkyTile[3]; \ - dstAddr += ST_BYTES_PER_ROW; \ - chunkyTile += TILE_BYTES_PER_ROW; \ - } while (0) - ST_TILE_PASTE_ROW; - ST_TILE_PASTE_ROW; - ST_TILE_PASTE_ROW; - ST_TILE_PASTE_ROW; - ST_TILE_PASTE_ROW; - ST_TILE_PASTE_ROW; - ST_TILE_PASTE_ROW; - ST_TILE_PASTE_ROW; -#undef ST_TILE_PASTE_ROW -} - - -// Planar monochrome paste. monoTile follows the cross-port mono -// contract (include/joey/tile.h + jlpGenericTilePasteMono): 32 chunky -// nibble-pair bytes, row-major, TILE_BYTES_PER_ROW bytes per row, two -// pixels per byte. A pixel renders fgColor when its source nibble is -// nonzero, bgColor when zero; the HIGH nibble is the LEFT pixel. Each -// row's 4 source bytes fold into an 8-bit shape mask (bit 7 = leftmost -// pixel, matching the planar bit order), then each plane k writes -// outPlaneK = (shape & maskFgK) | (~shape & maskBgK) -// at dstAddr + k*2 inside the 16-pixel group (4 plane words = 8 bytes -// per group). bx odd selects the second byte of each plane word (the -// low / right 8 pixels of the group). -void jlpTilePasteMono(jlSurfaceT *dst, uint8_t bx, uint8_t by, const uint8_t *monoTile, uint8_t fgColor, uint8_t bgColor) { - StPlanarT *pd; - uint16_t group; - uint8_t *dstAddr; - uint8_t row; - uint8_t col; - uint8_t plane; - uint8_t masksFg[4]; - uint8_t masksBg[4]; - - if (dst->portData == NULL) { - jlpGenericTilePasteMono(dst, bx, by, monoTile, fgColor, bgColor); - return; - } - pd = (StPlanarT *)dst->portData; - for (plane = 0u; plane < 4u; plane++) { - masksFg[plane] = (uint8_t)((fgColor & (1u << plane)) ? 0xFFu : 0x00u); - masksBg[plane] = (uint8_t)((bgColor & (1u << plane)) ? 0xFFu : 0x00u); - } - group = (uint16_t)((uint16_t)bx >> 1); - dstAddr = pd->base - + (uint16_t)by * 8u * ST_BYTES_PER_ROW - + group * ST_BYTES_PER_GROUP - + (uint16_t)(bx & 1u); - for (row = 0u; row < TILE_PIXELS_PER_SIDE; row++) { - uint8_t shape = 0u; - - for (col = 0u; col < TILE_BYTES_PER_ROW; col++) { - uint8_t srcByte = *monoTile++; - - shape = (uint8_t)(shape << 2); - if (srcByte & 0xF0u) { - shape = (uint8_t)(shape | 0x02u); - } - if (srcByte & 0x0Fu) { - shape = (uint8_t)(shape | 0x01u); - } - } - dstAddr[0] = (uint8_t)((uint8_t)(shape & masksFg[0]) | - (uint8_t)((uint8_t)(~shape) & masksBg[0])); - dstAddr[2] = (uint8_t)((uint8_t)(shape & masksFg[1]) | - (uint8_t)((uint8_t)(~shape) & masksBg[1])); - dstAddr[4] = (uint8_t)((uint8_t)(shape & masksFg[2]) | - (uint8_t)((uint8_t)(~shape) & masksBg[2])); - dstAddr[6] = (uint8_t)((uint8_t)(shape & masksFg[3]) | - (uint8_t)((uint8_t)(~shape) & masksBg[3])); - dstAddr += ST_BYTES_PER_ROW; - } -} - - -void jlpTileSnap(const jlSurfaceT *src, uint8_t bx, uint8_t by, uint8_t *chunkyOut) { - const StPlanarT *pd; - uint16_t group; - const uint8_t *srcAddr; - int16_t row; - - if (src->portData == NULL) { - jlpGenericTileSnap(src, bx, by, chunkyOut); - return; - } - pd = (const StPlanarT *)src->portData; - /* Phase 10.5: write plane-major bytes to jlTileT (4 per row * 8 rows). */ - group = (uint16_t)((uint16_t)bx >> 1); - srcAddr = pd->base - + (uint16_t)by * 8u * ST_BYTES_PER_ROW - + group * ST_BYTES_PER_GROUP - + (uint16_t)(bx & 1u); - (void)row; -#define ST_TILE_SNAP_ROW \ - do { \ - chunkyOut[0] = srcAddr[0]; \ - chunkyOut[1] = srcAddr[2]; \ - chunkyOut[2] = srcAddr[4]; \ - chunkyOut[3] = srcAddr[6]; \ - srcAddr += ST_BYTES_PER_ROW; \ - chunkyOut += TILE_BYTES_PER_ROW; \ - } while (0) - ST_TILE_SNAP_ROW; - ST_TILE_SNAP_ROW; - ST_TILE_SNAP_ROW; - ST_TILE_SNAP_ROW; - ST_TILE_SNAP_ROW; - ST_TILE_SNAP_ROW; - ST_TILE_SNAP_ROW; - ST_TILE_SNAP_ROW; -#undef ST_TILE_SNAP_ROW -} - - -// Phase 10: direct planar->planar tile copy. Each tile occupies one -// half-byte of one plane word per plane per row (8 rows total). -// We just byte-copy 4 plane bytes per row -- no chunky scratch, no -// bit transpose, no LUT. ~640 cyc per tile vs ~5000 cyc for the -// snap+paste path. -void jlpTileCopy(jlSurfaceT *dst, uint8_t dstBx, uint8_t dstBy, const jlSurfaceT *src, uint8_t srcBx, uint8_t srcBy) { - StPlanarT *dstPd; - const StPlanarT *srcPd; - uint8_t *dstAddr; - const uint8_t *srcAddr; - uint16_t srcGroup; - uint16_t dstGroup; - int16_t row; - - if (dst->portData == NULL || src->portData == NULL) { - jlpGenericTileCopy(dst, dstBx, dstBy, src, srcBx, srcBy); - return; - } - dstPd = (StPlanarT *)dst->portData; - srcPd = (const StPlanarT *)src->portData; - srcGroup = (uint16_t)((uint16_t)srcBx >> 1); - dstGroup = (uint16_t)((uint16_t)dstBx >> 1); - srcAddr = srcPd->base - + (uint16_t)srcBy * 8u * ST_BYTES_PER_ROW - + srcGroup * ST_BYTES_PER_GROUP - + (uint16_t)(srcBx & 1u); - dstAddr = dstPd->base - + (uint16_t)dstBy * 8u * ST_BYTES_PER_ROW - + dstGroup * ST_BYTES_PER_GROUP - + (uint16_t)(dstBx & 1u); - /* gcc-mint -O2 does NOT unroll the 8-iter byte-copy loop, - * leaving cmpl + bnes loop overhead per row. Manual unroll - * drops ~150 cyc/call. (void)row keeps the unused decl quiet. */ - (void)row; -#define ST_TILE_COPY_ROW \ - do { \ - dstAddr[0] = srcAddr[0]; \ - dstAddr[2] = srcAddr[2]; \ - dstAddr[4] = srcAddr[4]; \ - dstAddr[6] = srcAddr[6]; \ - srcAddr += ST_BYTES_PER_ROW; \ - dstAddr += ST_BYTES_PER_ROW; \ - } while (0) - ST_TILE_COPY_ROW; /* row 0 */ - ST_TILE_COPY_ROW; /* row 1 */ - ST_TILE_COPY_ROW; /* row 2 */ - ST_TILE_COPY_ROW; /* row 3 */ - ST_TILE_COPY_ROW; /* row 4 */ - ST_TILE_COPY_ROW; /* row 5 */ - ST_TILE_COPY_ROW; /* row 6 */ - ST_TILE_COPY_ROW; /* row 7 */ -#undef ST_TILE_COPY_ROW -} - - -void jlpTileCopyMasked(jlSurfaceT *dst, uint8_t dstBx, uint8_t dstBy, const jlSurfaceT *src, uint8_t srcBx, uint8_t srcBy, uint8_t transparent) { - StPlanarT *dstPd; - uint8_t scratch[TILE_BYTES]; - int16_t row; - uint16_t dstX0; - uint16_t dstY0; - uint16_t group; - uint16_t halfOff; - uint8_t *dstByte; - uint8_t p0; - uint8_t p1; - uint8_t p2; - uint8_t p3; - uint8_t xK0; - uint8_t xK1; - uint8_t xK2; - uint8_t xK3; - uint8_t mask; - uint8_t notMask; - - if (dst->portData == NULL) { - jlpGenericTileCopyMasked(dst, dstBx, dstBy, src, srcBx, srcBy, transparent); - return; - } - dstPd = (StPlanarT *)dst->portData; - /* Phase 10.5: bulk-plane fast path. scratch holds plane-major bytes - * (4 plane bytes per row * 8 rows). For each row, build a "non- - * transparent" mask = OR of (plane_byte XOR replicated transparent - * bit) -- 1s where the source pixel != transparent. Then 4 byte - * RMWs (one per plane) write the row at byte-aligned dst. - * - * For transparent=0 this collapses to mask = p0|p1|p2|p3. - * Replaces the prior 64-iteration per-pixel SetPixel walker. */ - jlpTileSnap(src, srcBx, srcBy, scratch); - dstX0 = (uint16_t)((uint16_t)dstBx * TILE_PIXELS_PER_SIDE); - dstY0 = (uint16_t)((uint16_t)dstBy * TILE_PIXELS_PER_SIDE); - group = (uint16_t)(dstX0 >> 4); - halfOff = (uint16_t)((dstX0 & 8u) >> 3u); - dstByte = dstPd->base + dstY0 * ST_BYTES_PER_ROW - + group * ST_BYTES_PER_GROUP + halfOff; - - xK0 = (transparent & 1u) ? 0xFFu : 0u; - xK1 = (transparent & 2u) ? 0xFFu : 0u; - xK2 = (transparent & 4u) ? 0xFFu : 0u; - xK3 = (transparent & 8u) ? 0xFFu : 0u; - - for (row = 0; row < TILE_PIXELS_PER_SIDE; row++) { - p0 = scratch[row * 4 + 0]; - p1 = scratch[row * 4 + 1]; - p2 = scratch[row * 4 + 2]; - p3 = scratch[row * 4 + 3]; - mask = (uint8_t)((p0 ^ xK0) | (p1 ^ xK1) | (p2 ^ xK2) | (p3 ^ xK3)); - if (mask != 0u) { - notMask = (uint8_t)~mask; - dstByte[0] = (uint8_t)((dstByte[0] & notMask) | (p0 & mask)); - dstByte[2] = (uint8_t)((dstByte[2] & notMask) | (p1 & mask)); - dstByte[4] = (uint8_t)((dstByte[4] & notMask) | (p2 & mask)); - dstByte[6] = (uint8_t)((dstByte[6] & notMask) | (p3 & mask)); - } - dstByte += ST_BYTES_PER_ROW; - } -} - - // Phase 10 fast path: byte-aligned, fully-on-surface sprite draw. // Builds 4 plane bytes + 1 opacity byte from each tile-column row // in one pass, then does 4 word RMWs per group half. ~7x faster diff --git a/src/atarist/stPlanar.h b/src/atarist/stPlanar.h index fd04afe..4308851 100644 --- a/src/atarist/stPlanar.h +++ b/src/atarist/stPlanar.h @@ -12,6 +12,11 @@ #include "surfaceInternal.h" +// Interleaved-layout strides, shared with the stTile.h inlines: +// 160 bytes per scanline as 20 groups of 4 plane words (8 bytes). +#define ST_BYTES_PER_ROW 160 +#define ST_BYTES_PER_GROUP 8 // 4 plane words back-to-back + // jlSurfaceT.portData points to this on the ST. ST has word-interleaved // planar: ONE 32000-byte buffer per surface holds all 4 planes packed // 8 bytes per 16-pixel group. Per-scanline layout (160 bytes) is 20 diff --git a/src/atarist/stTile.h b/src/atarist/stTile.h new file mode 100644 index 0000000..e77f78d --- /dev/null +++ b/src/atarist/stTile.h @@ -0,0 +1,315 @@ +// Atari ST tile ops as always-inline functions (NATIVE-PERF Phase 2 +// R4a). Bodies moved verbatim from src/atarist/hal.c so the public +// wrappers in src/core/tile.c inline them, deleting the second full +// cdecl call layer (~130-190 cycles/op on the 68000). The jlpTile* +// dispatch names are macro-aliased to these in port.h, which also +// #includes this header at its tail -- AFTER the jlpGenericTile* +// prototypes the portData == NULL fallbacks call. Do not include this +// header directly; it comes with port.h on the ST. +#ifndef JOEY_ST_TILE_H +#define JOEY_ST_TILE_H + +#include +#include + +#include "joey/tile.h" +#include "stPlanar.h" + +// 68k asm entry (surface68kSt.s side): fills one 8-px half-group in +// all 4 planes. Declared here because stTileFill is its only caller. +extern void surface68kStTileFill8x8(uint8_t *firstGroupPtr, uint16_t mask, uint8_t color); + +// Phase 5 tile ops. 8x8 tiles at byte position (bx, by) start at +// pixel (bx*8, by*8). 8 pixels wide always covers exactly half a +// 16-pixel group: high half (bits 15..8) when bx is even, low half +// (bits 7..0) when bx is odd. Per-row work is 4 plane half-word RMWs. +static inline __attribute__((always_inline)) void stTileFill(jlSurfaceT *s, uint8_t bx, uint8_t by, uint8_t colorIndex) { + StPlanarT *pd; + uint16_t group; + uint16_t halfMask; + uint8_t *gp; + + // Core jlTileFill already proved s != NULL and clipped bx/by; only + // the non-planar (offscreen, no portData) fallback needs handling. + pd = (StPlanarT *)s->portData; + if (pd == NULL) { + jlpGenericTileFill(s, bx, by, colorIndex); + return; + } + group = (uint16_t)((uint16_t)bx >> 1); + halfMask = ((bx & 1u) == 0u) ? 0xFF00u : 0x00FFu; + gp = pd->base + (uint16_t)by * 8u * ST_BYTES_PER_ROW + group * ST_BYTES_PER_GROUP; + /* Phase 10 final: specialized 8x8 unrolled tile-fill skips the + * generic FRG_LOOP's per-row subq+bne overhead. */ + surface68kStTileFill8x8(gp, halfMask, colorIndex); +} + + +// Phase 10: tile paste/snap reuse the asm sprite save/restore +// helpers -- identical per-row work patterns at byte-aligned +// positions. Width 8 = single tile column = single half-group +// write per plane. The asm walker handles 8 rows just as well +// as a sprite's variable height. +static inline __attribute__((always_inline)) void stTilePaste(jlSurfaceT *dst, uint8_t bx, uint8_t by, const uint8_t *chunkyTile) { + StPlanarT *pd; + uint16_t group; + uint8_t *dstAddr; + int16_t row; + + if (dst->portData == NULL) { + jlpGenericTilePaste(dst, bx, by, chunkyTile); + return; + } + pd = (StPlanarT *)dst->portData; + /* Phase 10.5: jlTileT.pixels holds plane-major bytes (4 plane bytes + * per row * 8 rows = 32 bytes). Direct byte copy to the planar + * buffer; no chunky <-> planar conversion. Mirrors the sibling + * jlpTileCopy pattern but reads from the contiguous tile + * buffer. Drops the asm-walker entry/exit overhead. */ + group = (uint16_t)((uint16_t)bx >> 1); + dstAddr = pd->base + + (uint16_t)by * 8u * ST_BYTES_PER_ROW + + group * ST_BYTES_PER_GROUP + + (uint16_t)(bx & 1u); + (void)row; +#define ST_TILE_PASTE_ROW \ + do { \ + dstAddr[0] = chunkyTile[0]; \ + dstAddr[2] = chunkyTile[1]; \ + dstAddr[4] = chunkyTile[2]; \ + dstAddr[6] = chunkyTile[3]; \ + dstAddr += ST_BYTES_PER_ROW; \ + chunkyTile += TILE_BYTES_PER_ROW; \ + } while (0) + ST_TILE_PASTE_ROW; + ST_TILE_PASTE_ROW; + ST_TILE_PASTE_ROW; + ST_TILE_PASTE_ROW; + ST_TILE_PASTE_ROW; + ST_TILE_PASTE_ROW; + ST_TILE_PASTE_ROW; + ST_TILE_PASTE_ROW; +#undef ST_TILE_PASTE_ROW +} + + +// Planar monochrome paste. monoTile follows the cross-port mono +// contract (include/joey/tile.h + jlpGenericTilePasteMono): 32 chunky +// nibble-pair bytes, row-major, TILE_BYTES_PER_ROW bytes per row, two +// pixels per byte. A pixel renders fgColor when its source nibble is +// nonzero, bgColor when zero; the HIGH nibble is the LEFT pixel. Each +// row's 4 source bytes fold into an 8-bit shape mask (bit 7 = leftmost +// pixel, matching the planar bit order), then each plane k writes +// outPlaneK = (shape & maskFgK) | (~shape & maskBgK) +// at dstAddr + k*2 inside the 16-pixel group (4 plane words = 8 bytes +// per group). bx odd selects the second byte of each plane word (the +// low / right 8 pixels of the group). +static inline __attribute__((always_inline)) void stTilePasteMono(jlSurfaceT *dst, uint8_t bx, uint8_t by, const uint8_t *monoTile, uint8_t fgColor, uint8_t bgColor) { + StPlanarT *pd; + uint16_t group; + uint8_t *dstAddr; + uint8_t row; + uint8_t col; + uint8_t plane; + uint8_t masksFg[4]; + uint8_t masksBg[4]; + + if (dst->portData == NULL) { + jlpGenericTilePasteMono(dst, bx, by, monoTile, fgColor, bgColor); + return; + } + pd = (StPlanarT *)dst->portData; + for (plane = 0u; plane < 4u; plane++) { + masksFg[plane] = (uint8_t)((fgColor & (1u << plane)) ? 0xFFu : 0x00u); + masksBg[plane] = (uint8_t)((bgColor & (1u << plane)) ? 0xFFu : 0x00u); + } + group = (uint16_t)((uint16_t)bx >> 1); + dstAddr = pd->base + + (uint16_t)by * 8u * ST_BYTES_PER_ROW + + group * ST_BYTES_PER_GROUP + + (uint16_t)(bx & 1u); + for (row = 0u; row < TILE_PIXELS_PER_SIDE; row++) { + uint8_t shape = 0u; + + for (col = 0u; col < TILE_BYTES_PER_ROW; col++) { + uint8_t srcByte = *monoTile++; + + shape = (uint8_t)(shape << 2); + if (srcByte & 0xF0u) { + shape = (uint8_t)(shape | 0x02u); + } + if (srcByte & 0x0Fu) { + shape = (uint8_t)(shape | 0x01u); + } + } + dstAddr[0] = (uint8_t)((uint8_t)(shape & masksFg[0]) | + (uint8_t)((uint8_t)(~shape) & masksBg[0])); + dstAddr[2] = (uint8_t)((uint8_t)(shape & masksFg[1]) | + (uint8_t)((uint8_t)(~shape) & masksBg[1])); + dstAddr[4] = (uint8_t)((uint8_t)(shape & masksFg[2]) | + (uint8_t)((uint8_t)(~shape) & masksBg[2])); + dstAddr[6] = (uint8_t)((uint8_t)(shape & masksFg[3]) | + (uint8_t)((uint8_t)(~shape) & masksBg[3])); + dstAddr += ST_BYTES_PER_ROW; + } +} + + +static inline __attribute__((always_inline)) void stTileSnap(const jlSurfaceT *src, uint8_t bx, uint8_t by, uint8_t *chunkyOut) { + const StPlanarT *pd; + uint16_t group; + const uint8_t *srcAddr; + int16_t row; + + if (src->portData == NULL) { + jlpGenericTileSnap(src, bx, by, chunkyOut); + return; + } + pd = (const StPlanarT *)src->portData; + /* Phase 10.5: write plane-major bytes to jlTileT (4 per row * 8 rows). */ + group = (uint16_t)((uint16_t)bx >> 1); + srcAddr = pd->base + + (uint16_t)by * 8u * ST_BYTES_PER_ROW + + group * ST_BYTES_PER_GROUP + + (uint16_t)(bx & 1u); + (void)row; +#define ST_TILE_SNAP_ROW \ + do { \ + chunkyOut[0] = srcAddr[0]; \ + chunkyOut[1] = srcAddr[2]; \ + chunkyOut[2] = srcAddr[4]; \ + chunkyOut[3] = srcAddr[6]; \ + srcAddr += ST_BYTES_PER_ROW; \ + chunkyOut += TILE_BYTES_PER_ROW; \ + } while (0) + ST_TILE_SNAP_ROW; + ST_TILE_SNAP_ROW; + ST_TILE_SNAP_ROW; + ST_TILE_SNAP_ROW; + ST_TILE_SNAP_ROW; + ST_TILE_SNAP_ROW; + ST_TILE_SNAP_ROW; + ST_TILE_SNAP_ROW; +#undef ST_TILE_SNAP_ROW +} + + +// Phase 10: direct planar->planar tile copy. Each tile occupies one +// half-byte of one plane word per plane per row (8 rows total). +// We just byte-copy 4 plane bytes per row -- no chunky scratch, no +// bit transpose, no LUT. ~640 cyc per tile vs ~5000 cyc for the +// snap+paste path. +static inline __attribute__((always_inline)) void stTileCopy(jlSurfaceT *dst, uint8_t dstBx, uint8_t dstBy, const jlSurfaceT *src, uint8_t srcBx, uint8_t srcBy) { + StPlanarT *dstPd; + const StPlanarT *srcPd; + uint8_t *dstAddr; + const uint8_t *srcAddr; + uint16_t srcGroup; + uint16_t dstGroup; + int16_t row; + + if (dst->portData == NULL || src->portData == NULL) { + jlpGenericTileCopy(dst, dstBx, dstBy, src, srcBx, srcBy); + return; + } + dstPd = (StPlanarT *)dst->portData; + srcPd = (const StPlanarT *)src->portData; + srcGroup = (uint16_t)((uint16_t)srcBx >> 1); + dstGroup = (uint16_t)((uint16_t)dstBx >> 1); + srcAddr = srcPd->base + + (uint16_t)srcBy * 8u * ST_BYTES_PER_ROW + + srcGroup * ST_BYTES_PER_GROUP + + (uint16_t)(srcBx & 1u); + dstAddr = dstPd->base + + (uint16_t)dstBy * 8u * ST_BYTES_PER_ROW + + dstGroup * ST_BYTES_PER_GROUP + + (uint16_t)(dstBx & 1u); + /* gcc-mint -O2 does NOT unroll the 8-iter byte-copy loop, + * leaving cmpl + bnes loop overhead per row. Manual unroll + * drops ~150 cyc/call. (void)row keeps the unused decl quiet. */ + (void)row; +#define ST_TILE_COPY_ROW \ + do { \ + dstAddr[0] = srcAddr[0]; \ + dstAddr[2] = srcAddr[2]; \ + dstAddr[4] = srcAddr[4]; \ + dstAddr[6] = srcAddr[6]; \ + srcAddr += ST_BYTES_PER_ROW; \ + dstAddr += ST_BYTES_PER_ROW; \ + } while (0) + ST_TILE_COPY_ROW; /* row 0 */ + ST_TILE_COPY_ROW; /* row 1 */ + ST_TILE_COPY_ROW; /* row 2 */ + ST_TILE_COPY_ROW; /* row 3 */ + ST_TILE_COPY_ROW; /* row 4 */ + ST_TILE_COPY_ROW; /* row 5 */ + ST_TILE_COPY_ROW; /* row 6 */ + ST_TILE_COPY_ROW; /* row 7 */ +#undef ST_TILE_COPY_ROW +} + + +static inline __attribute__((always_inline)) void stTileCopyMasked(jlSurfaceT *dst, uint8_t dstBx, uint8_t dstBy, const jlSurfaceT *src, uint8_t srcBx, uint8_t srcBy, uint8_t transparent) { + StPlanarT *dstPd; + uint8_t scratch[TILE_BYTES]; + int16_t row; + uint16_t dstX0; + uint16_t dstY0; + uint16_t group; + uint16_t halfOff; + uint8_t *dstByte; + uint8_t p0; + uint8_t p1; + uint8_t p2; + uint8_t p3; + uint8_t xK0; + uint8_t xK1; + uint8_t xK2; + uint8_t xK3; + uint8_t mask; + uint8_t notMask; + + if (dst->portData == NULL) { + jlpGenericTileCopyMasked(dst, dstBx, dstBy, src, srcBx, srcBy, transparent); + return; + } + dstPd = (StPlanarT *)dst->portData; + /* Phase 10.5: bulk-plane fast path. scratch holds plane-major bytes + * (4 plane bytes per row * 8 rows). For each row, build a "non- + * transparent" mask = OR of (plane_byte XOR replicated transparent + * bit) -- 1s where the source pixel != transparent. Then 4 byte + * RMWs (one per plane) write the row at byte-aligned dst. + * + * For transparent=0 this collapses to mask = p0|p1|p2|p3. + * Replaces the prior 64-iteration per-pixel SetPixel walker. */ + stTileSnap(src, srcBx, srcBy, scratch); + dstX0 = (uint16_t)((uint16_t)dstBx * TILE_PIXELS_PER_SIDE); + dstY0 = (uint16_t)((uint16_t)dstBy * TILE_PIXELS_PER_SIDE); + group = (uint16_t)(dstX0 >> 4); + halfOff = (uint16_t)((dstX0 & 8u) >> 3u); + dstByte = dstPd->base + dstY0 * ST_BYTES_PER_ROW + + group * ST_BYTES_PER_GROUP + halfOff; + + xK0 = (transparent & 1u) ? 0xFFu : 0u; + xK1 = (transparent & 2u) ? 0xFFu : 0u; + xK2 = (transparent & 4u) ? 0xFFu : 0u; + xK3 = (transparent & 8u) ? 0xFFu : 0u; + + for (row = 0; row < TILE_PIXELS_PER_SIDE; row++) { + p0 = scratch[row * 4 + 0]; + p1 = scratch[row * 4 + 1]; + p2 = scratch[row * 4 + 2]; + p3 = scratch[row * 4 + 3]; + mask = (uint8_t)((p0 ^ xK0) | (p1 ^ xK1) | (p2 ^ xK2) | (p3 ^ xK3)); + if (mask != 0u) { + notMask = (uint8_t)~mask; + dstByte[0] = (uint8_t)((dstByte[0] & notMask) | (p0 & mask)); + dstByte[2] = (uint8_t)((dstByte[2] & notMask) | (p1 & mask)); + dstByte[4] = (uint8_t)((dstByte[4] & notMask) | (p2 & mask)); + dstByte[6] = (uint8_t)((dstByte[6] & notMask) | (p3 & mask)); + } + dstByte += ST_BYTES_PER_ROW; + } +} + +#endif diff --git a/src/codegen/spriteCompile.c b/src/codegen/spriteCompile.c index cbb5471..78af33d 100644 --- a/src/codegen/spriteCompile.c +++ b/src/codegen/spriteCompile.c @@ -368,464 +368,7 @@ bool jlSpriteCompile(jlSpriteT *sp) { // re-derivation is gone (#29). -#if defined(JOEYLIB_PLATFORM_IIGS) - -// SURFACE_ROW_OFFSET dispatches to the gRowOffsetLut lookup on IIgs; -// declared in surfaceInternal.h. Uses a single indexed long-mode read -// for the multiply. - -// IIgs sprite dispatch. The emitted routines (spriteEmitIigs.c) are ordinary -// C-ABI functions called through C function pointers, so the compiler emits the -// cross-bank JSL itself (large memory model -> 24-bit pointers). The DRAW -// routine takes the 24-bit dst pointer in A:X and sets up its own DBR (needed -// for `sta abs,Y`, whose high byte comes from DBR) + Y in a fixed prologue, then -// restores DBR before returning. SAVE/RESTORE take packed src/dst offsets and -// have only their MVN bank operands patched per call (a plain C pointer write -// into the arena from the dispatcher). No self-modifying call stub, no inline -// asm -- the machinery that hid the old un-instrumentable #19 corruption. -void spriteCompiledDraw(jlSurfaceT *dst, const jlSpriteT *sp, int16_t x, int16_t y, uint16_t routeOffset) { - uint8_t *dstRow0; - uint32_t fnAddr; - void (*drawFn)(uint8_t *); - - // Clean C-ABI dispatch. The emitted DRAW routine is a normal C function - // -- void draw(uint8_t *dstRow0) -- so the compiler emits the indirect - // JSL itself: NO self-modifying call stub, NO inline asm, NO hand-rolled - // per-call bank/offset patching (that machinery was the source of the - // old un-instrumentable corruption). dstRow0 (the sprite's top-left - // screen byte) is passed in A:X; the routine's prologue (spriteEmitIigs.c) - // sets up its own DBR/Y and restores DBR on exit. routeOffset is the - // gate-derived DRAW offset for this shift (#29). - dstRow0 = &dst->pixels[SURFACE_ROW_OFFSET(y) + ((uint16_t)x >> 1)]; - fnAddr = codegenArenaBaseAddr() - + sp->slot->offset - + (uint32_t)routeOffset; - drawFn = (void (*)(uint8_t *))fnAddr; - drawFn(dstRow0); -} - - -// patchMvnBanks stamps the dst/src bank operand bytes into each MVN inside an -// emitted C-ABI save/restore routine -- the routine's only runtime-dynamic -// bytes, since MVN's banks are immediate operands and the screen/backup banks -// vary at runtime. The patch is a plain C pointer write into the arena from -// the dispatcher (no self-modifying call stub). C-ABI routine layout from -// spriteEmitIigs.c::emitMvnCabiRoutine: -// prologue (4): 8B DA AA 7A -// row 0 (6): A9 lo hi 54 db sb -// row R (12): 8A/98 18 69 lo hi AA/A8 A9 lo hi 54 db sb -// epilogue (2): AB 6B -// The 4-byte prologue shifts MVN dstbk to routine offset (12*R + 8), srcbk to -// (12*R + 9). -static void patchMvnBanks(uint8_t *routine, uint16_t heightPx, uint8_t dstBank, uint8_t srcBank) { - uint16_t r; - - for (r = 0; r < heightPx; r++) { - routine[12u * r + 8u] = dstBank; - routine[12u * r + 9u] = srcBank; - } -} - - - -// Split a 24-bit pointer into its low 16 bits + bank byte. To avoid -// invoking a 32-bit-shift helper for the `>> 16` to extract the bank -// byte, we cast to uint32_t and then byte-alias the storage -- gets -// the same bytes with simple loads. -#define SPLIT_POINTER(_ptr, _outLo, _outBank) \ - do { \ - uint32_t spAddr_ = (uint32_t)(_ptr); \ - const uint8_t *spB_ = (const uint8_t *)&spAddr_; \ - *(_outLo) = (uint16_t)(spB_[0] | ((uint16_t)spB_[1] << 8)); \ - *(_outBank) = spB_[2]; \ - } while (0) - -// The backup->bytes pointer is split with SPLIT_POINTER directly. A -// single-entry global split cache used to live here (#59); it assumed -// one backup buffer per program, but the normal pattern is one buffer -// per moving sprite, so with 2+ sprites the cached pointer ping-ponged -// and every call paid a 24-bit compare AND the full split. The split -// itself is ~6 simple loads/stores -- cheaper than the compare on a -// hit -- so the cache was deleted rather than moved per-sprite. - - -void spriteCompiledSaveUnder(const jlSurfaceT *src, jlSpriteT *sp, int16_t x, int16_t y, jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { - int16_t clippedX; - uint16_t heightPx; - uint16_t copyBytes; - uint16_t screenLo; - uint16_t backupLo; - uint8_t screenBank; - uint8_t backupBank; - uint32_t fnAddr; - uint8_t *routine; - uint8_t *screenPtr; - uint16_t cacheIdx; /* shift * SPRITE_OP_COUNT + SPRITE_OP_SAVE, computed once */ - uint8_t *cachedDst; /* &sp->cachedDstBank[0][0] + cacheIdx */ - uint8_t *cachedSrc; /* &sp->cachedSrcBank[0][0] + cacheIdx */ - - // shift and routeOffset arrive pre-derived from the dispatcher - // gate (#29); widthPx/heightPx are create-time constants cached in - // the sprite struct. - clippedX = (int16_t)(x & ~1); - heightPx = sp->heightPx; - copyBytes = (uint16_t)((sp->widthPx >> 1) + shift); /* shift is x & 1, exactly 0 or 1 */ - - screenPtr = (uint8_t *)&src->pixels[SURFACE_ROW_OFFSET(y) + ((uint16_t)clippedX >> 1)]; - SPLIT_POINTER(screenPtr, &screenLo, &screenBank); - SPLIT_POINTER(backup->bytes, &backupLo, &backupBank); - - backup->sprite = sp; - backup->x = clippedX; - backup->y = y; - backup->width = (uint16_t)(copyBytes << 1); - backup->height = heightPx; - /* sizeBytes is constant per (sprite, shift); cache to dodge the - * per-call ~CUMUL2 (uint16_t * uint16_t) helper. The byte-pointer - * arithmetic avoids reintroducing ~MUL4 for the uint16_t array - * indexing. */ - { - uint16_t *sizeCachePtr = (uint16_t *)((uint8_t *)sp->cachedSizeBytes + ((uint16_t)shift << 1)); - if (*sizeCachePtr == 0) { - *sizeCachePtr = (uint16_t)(copyBytes * heightPx); - } - backup->sizeBytes = *sizeCachePtr; - } - - /* Compute the 1D index into the cached* bank arrays once; - * SPRITE_ROUTINE_INDEX (spriteInternal.h) owns the multiply-free - * derivation. */ - cacheIdx = SPRITE_ROUTINE_INDEX(shift, SPRITE_OP_SAVE); - cachedDst = (uint8_t *)sp->cachedDstBank + cacheIdx; - cachedSrc = (uint8_t *)sp->cachedSrcBank + cacheIdx; - - fnAddr = codegenArenaBaseAddr() - + sp->slot->offset - + (uint32_t)routeOffset; - - // Patch the MVN bank operands only if the dst/src bank pair changed since - // last call (the routine's only dynamic bytes). SAVE copies screen -> - // backup: dst = backup bank, src = screen bank. - if (*cachedDst != backupBank || *cachedSrc != screenBank) { - routine = (uint8_t *)fnAddr; - patchMvnBanks(routine, heightPx, /*dst*/backupBank, /*src*/screenBank); - *cachedDst = backupBank; - *cachedSrc = screenBank; - } - - // Clean C-ABI dispatch -- void save(uint32_t packed), packed = srcOff | - // (dstOff << 16). The compiler emits the indirect JSL: no self-modifying - // call stub, no inline asm. SAVE: src = screen offset, dst = backup offset. - { - uint32_t packed = (uint32_t)screenLo | ((uint32_t)backupLo << 16); - ((void (*)(uint32_t))fnAddr)(packed); - } -} - - -void spriteCompiledRestoreUnder(jlSurfaceT *dst, const jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { - uint16_t heightPx; - uint16_t screenLo; - uint16_t backupLo; - uint8_t screenBank; - uint8_t backupBank; - uint32_t fnAddr; - uint8_t *routine; - uint8_t *screenPtr; - jlSpriteT *sp; - uint16_t cacheIdx; /* shift * SPRITE_OP_COUNT + SPRITE_OP_RESTORE, computed once */ - uint8_t *cachedDst; - uint8_t *cachedSrc; - - // shift and routeOffset arrive pre-derived from the - // jlSpriteRestoreUnder gate (#29) -- no per-call re-inference of - // the shift from the backup here. - sp = backup->sprite; - heightPx = backup->height; - - screenPtr = (uint8_t *)&dst->pixels[SURFACE_ROW_OFFSET(backup->y) + ((uint16_t)backup->x >> 1)]; - SPLIT_POINTER(screenPtr, &screenLo, &screenBank); - SPLIT_POINTER(backup->bytes, &backupLo, &backupBank); - - /* Hoist 2D-array indexing -- see save-side comment. */ - cacheIdx = SPRITE_ROUTINE_INDEX(shift, SPRITE_OP_RESTORE); - cachedDst = (uint8_t *)sp->cachedDstBank + cacheIdx; - cachedSrc = (uint8_t *)sp->cachedSrcBank + cacheIdx; - - fnAddr = codegenArenaBaseAddr() - + sp->slot->offset - + (uint32_t)routeOffset; - - // Patch the MVN bank operands only if they changed (the routine's only - // dynamic bytes). RESTORE copies backup -> screen: dst = screen bank, - // src = backup bank. - if (*cachedDst != screenBank || *cachedSrc != backupBank) { - routine = (uint8_t *)fnAddr; - patchMvnBanks(routine, heightPx, /*dst*/screenBank, /*src*/backupBank); - *cachedDst = screenBank; - *cachedSrc = backupBank; - } - - // Clean C-ABI dispatch -- void restore(uint32_t packed), packed = srcOff | - // (dstOff << 16). RESTORE: src = backup offset, dst = screen offset. - { - uint32_t packed = (uint32_t)backupLo | ((uint32_t)screenLo << 16); - ((void (*)(uint32_t))fnAddr)(packed); - } -} - - -#elif defined(JOEYLIB_PLATFORM_AMIGA) - -// Amiga planar dispatchers. spriteEmitPlanar68k.c emits DRAW routines -// with a cdecl(p0, p1, p2, p3) signature that write directly to -// bitplanes, one pre-shifted variant per x % 8 phase (all 8 compiled -// since NATIVE-PERF Phase 1). Compute byteOff = y*40 + x/8 and pass -// plane[i]+byteOff as the 4 plane args. Every caller gate in -// src/core/sprite.c checks routineOffsets[shift][op] != -// SPRITE_NOT_COMPILED before calling here (#29); anything uncompiled -// (clip, degrade-tier truncation) falls back to the interpreted -// jlpSpriteDrawPlanes / jlpSpriteSavePlanes / jlpSpriteRestorePlanes. -// -// Plane pointers come straight from portData via amigaPlanar.h's -// inline accessor (NATIVE-PERF Phase 2) -- the old jlpSurfacePlanePtr -// cross-TU call cost ~190 cycles PER PLANE, x4 per dispatch. The -// pd == NULL gate preserves the "not a planar surface -> no-op" -// contract those calls provided. - -#include "amigaPlanar.h" - -void spriteCompiledDraw(jlSurfaceT *dst, const jlSpriteT *sp, int16_t x, int16_t y, uint16_t routeOffset) { - typedef void (*DrawFn)(uint8_t *p0, uint8_t *p1, uint8_t *p2, uint8_t *p3); - AmigaPlanarT *pd; - uint16_t byteOff; - DrawFn fn; - - pd = amigaSurfacePlanar(dst); - if (pd == NULL) { - return; - } - byteOff = (uint16_t)((uint16_t)y * AMIGA_BYTES_PER_ROW + ((uint16_t)x >> 3)); - fn = (DrawFn)(codegenArenaBase() + sp->slot->offset + routeOffset); - fn(pd->planes[0] + byteOff, pd->planes[1] + byteOff, pd->planes[2] + byteOff, pd->planes[3] + byteOff); -} - - -// Compiled save/restore (NATIVE-PERF Phase 1): the emitted routines are -// cdecl(p0, p1, p2, p3, buf) plane-run copiers over the 16-px-group -// window (bx = x & ~15, nGroups = SPRITE_GROUPS0 + class). shift here -// is the width CLASS from SPRITE_SAVE_CLASS / SPRITE_RESTORE_CLASS, -// not the draw bit-phase. Row displacements are baked into the -// routine; the window's x phase enters only through the plane base -// pointers (byteOff is even by construction -- bx % 16 == 0 -- so the -// emitted move.w/move.l are 68000-legal). -void spriteCompiledSaveUnder(const jlSurfaceT *src, jlSpriteT *sp, int16_t x, int16_t y, jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { - typedef void (*CopyFn)(uint8_t *p0, uint8_t *p1, uint8_t *p2, uint8_t *p3, uint8_t *buf); - AmigaPlanarT *pd; - int16_t bx; - uint16_t nGroups; - uint16_t byteOff; - CopyFn fn; - - pd = amigaSurfacePlanar(src); - if (pd == NULL) { - return; - } - bx = (int16_t)(x & ~15); - nGroups = (uint16_t)(SPRITE_GROUPS0(sp->widthPx) + shift); - byteOff = (uint16_t)((uint16_t)y * AMIGA_BYTES_PER_ROW + ((uint16_t)bx >> 3)); - - backup->x = bx; - backup->y = y; - backup->width = (uint16_t)(nGroups << 4); - backup->height = sp->heightPx; - backup->sizeBytes = (uint16_t)((uint16_t)(nGroups << 3) * sp->heightPx); - - fn = (CopyFn)(codegenArenaBase() + sp->slot->offset + routeOffset); - fn(pd->planes[0] + byteOff, pd->planes[1] + byteOff, pd->planes[2] + byteOff, pd->planes[3] + byteOff, backup->bytes); -} - - -void spriteCompiledRestoreUnder(jlSurfaceT *dst, const jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { - typedef void (*CopyFn)(uint8_t *p0, uint8_t *p1, uint8_t *p2, uint8_t *p3, uint8_t *buf); - AmigaPlanarT *pd; - uint16_t byteOff; - CopyFn fn; - - (void)shift; - pd = amigaSurfacePlanar(dst); - if (pd == NULL) { - return; - } - byteOff = (uint16_t)((uint16_t)backup->y * AMIGA_BYTES_PER_ROW + ((uint16_t)backup->x >> 3)); - fn = (CopyFn)(codegenArenaBase() + backup->sprite->slot->offset + routeOffset); - fn(pd->planes[0] + byteOff, pd->planes[1] + byteOff, pd->planes[2] + byteOff, pd->planes[3] + byteOff, backup->bytes); -} - -#elif defined(JOEYLIB_PLATFORM_ATARIST) - -// ST word-interleaved planar runtime dispatch. The JIT routine takes -// one arg: groupBase = pd->base + y*160 + (x>>4)*8 (the address of -// the first 16-pixel group the sprite touches). It walks rows by -// adda.w #160 at the end of each row. Per (row, tile_col, plane) it -// emits up to one move.b / clr.b / andi.b+ori.b / ori.b chain at -// d16(a0). -// -// shift selection (in spriteInternal.h SPRITE_SHIFT_INDEX): the full -// intra-group phase x & 15. Phases 0/8 are the unrolled byte-aligned -// JIT; the 14 sub-byte phases are thunk + pre-shifted data tables -// driven by stSpriteDrawShiftWalker (all compiled since NATIVE-PERF -// Phase 1). Anything uncompiled falls back to jlpSpriteDrawPlanes via -// the caller gates (#29). - -// ST word-interleaved row stride: 320 px * 4 planes / 8 bits = 160 -// bytes/row. Numerically equal to SURFACE_BYTES_PER_ROW but a distinct -// fact (group-interleaved layout, not chunky), so it gets its own -// name. Each 16-pixel group is 4 plane words = 8 bytes. -#define ST_GROUP_BYTES_PER_ROW (SURFACE_WIDTH * 4 / 8) -#define ST_BYTES_PER_PLANE_GROUP 8u - -// Buffer base comes straight from portData via stPlanar.h's inline -// accessor (NATIVE-PERF Phase 2); the pd == NULL gate preserves the -// "not a planar surface -> no-op" contract the old jlpSurfacePlanePtr -// cross-TU call provided. - -#include "stPlanar.h" - -void spriteCompiledDraw(jlSurfaceT *dst, const jlSpriteT *sp, int16_t x, int16_t y, uint16_t routeOffset) { - typedef void (*DrawFn)(uint8_t *groupBase); - StPlanarT *pd; - uint8_t *groupBase; - DrawFn fn; - - pd = stSurfacePlanar(dst); - if (pd == NULL) { - return; - } - groupBase = pd->base - + (uint16_t)y * ST_GROUP_BYTES_PER_ROW - + (uint16_t)((uint16_t)x >> 4) * ST_BYTES_PER_PLANE_GROUP; - fn = (DrawFn)(codegenArenaBase() + sp->slot->offset + routeOffset); - fn(groupBase); -} - - -// Compiled save/restore (NATIVE-PERF Phase 1): the emitted routines are -// cdecl(src, dst) raw-span copiers over the 16-px-group window -// (bx = x & ~15, nGroups = SPRITE_GROUPS0 + class); a group window's -// rows are CONTIGUOUS byte spans in the interleaved layout, so the -// routine is a pure move.l chain with the inter-row stride baked in. -// shift here is the width CLASS, not the draw bit-phase. rowStart is -// always even (bx % 16 == 0, base is long-aligned), so move.l is -// 68000-legal. -void spriteCompiledSaveUnder(const jlSurfaceT *src, jlSpriteT *sp, int16_t x, int16_t y, jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { - typedef void (*CopyFn)(const uint8_t *src, uint8_t *dst); - StPlanarT *pd; - int16_t bx; - uint16_t nGroups; - uint8_t *rowStart; - CopyFn fn; - - pd = stSurfacePlanar(src); - if (pd == NULL) { - return; - } - bx = (int16_t)(x & ~15); - nGroups = (uint16_t)(SPRITE_GROUPS0(sp->widthPx) + shift); - rowStart = pd->base - + (uint16_t)y * ST_GROUP_BYTES_PER_ROW - + (uint16_t)((uint16_t)bx >> 4) * ST_BYTES_PER_PLANE_GROUP; - - backup->x = bx; - backup->y = y; - backup->width = (uint16_t)(nGroups << 4); - backup->height = sp->heightPx; - backup->sizeBytes = (uint16_t)((uint16_t)(nGroups << 3) * sp->heightPx); - - fn = (CopyFn)(codegenArenaBase() + sp->slot->offset + routeOffset); - fn(rowStart, backup->bytes); -} - - -void spriteCompiledRestoreUnder(jlSurfaceT *dst, const jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { - typedef void (*CopyFn)(const uint8_t *src, uint8_t *dst); - StPlanarT *pd; - uint8_t *rowStart; - CopyFn fn; - - (void)shift; - pd = stSurfacePlanar(dst); - if (pd == NULL) { - return; - } - rowStart = pd->base - + (uint16_t)backup->y * ST_GROUP_BYTES_PER_ROW - + (uint16_t)((uint16_t)backup->x >> 4) * ST_BYTES_PER_PLANE_GROUP; - fn = (CopyFn)(codegenArenaBase() + backup->sprite->slot->offset + routeOffset); - fn(backup->bytes, rowStart); -} - -#else - -void spriteCompiledDraw(jlSurfaceT *dst, const jlSpriteT *sp, int16_t x, int16_t y, uint16_t routeOffset) { - typedef void (*DrawFn)(uint8_t *destRow); - uint8_t *destRow; - DrawFn fn; - - destRow = &dst->pixels[(uint16_t)y * SURFACE_BYTES_PER_ROW + ((uint16_t)x >> 1)]; - fn = (DrawFn)(codegenArenaBase() + sp->slot->offset + routeOffset); - fn(destRow); -} - - -// x86 / 68k compiled save: bytes are a cdecl -// void copy(const uint8_t *src, uint8_t *dst) -// that walks heightPx rows of copyBytes from screen (stride -// SURFACE_BYTES_PER_ROW) into the contiguous backup buffer. shift and -// routeOffset arrive pre-derived from the dispatcher gate (#29). -void spriteCompiledSaveUnder(const jlSurfaceT *src, jlSpriteT *sp, int16_t x, int16_t y, jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { - typedef void (*CopyFn)(const uint8_t *src, uint8_t *dst); - int16_t clippedX; - uint16_t heightPx; - uint16_t copyBytes; - uint8_t *screenPtr; - CopyFn fn; - - clippedX = (int16_t)(x & ~1); - heightPx = sp->heightPx; - copyBytes = (uint16_t)((sp->widthPx >> 1) + shift); /* shift is x & 1, exactly 0 or 1 */ - - screenPtr = (uint8_t *)&src->pixels[(uint16_t)y * SURFACE_BYTES_PER_ROW + ((uint16_t)clippedX >> 1)]; - - backup->sprite = sp; - backup->x = clippedX; - backup->y = y; - backup->width = (uint16_t)(copyBytes << 1); - backup->height = heightPx; - backup->sizeBytes = (uint16_t)(copyBytes * heightPx); - - fn = (CopyFn)(codegenArenaBase() + sp->slot->offset + routeOffset); - fn(screenPtr, backup->bytes); -} - - -// Mirror of save: caller swaps arg order so the same emitted shape -// drives backup -> screen. The screen-side stride lives inside the -// emitted bytes, so RESTORE has its own routine bytes (stride is -// applied to dst instead of src). routeOffset arrives pre-derived -// from the jlSpriteRestoreUnder gate (#29); shift is unused here (no -// bank-patch cache off-IIgs) but kept for the shared signature. -void spriteCompiledRestoreUnder(jlSurfaceT *dst, const jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { - typedef void (*CopyFn)(const uint8_t *src, uint8_t *dst); - jlSpriteT *sp; - uint8_t *screenPtr; - CopyFn fn; - - (void)shift; - sp = backup->sprite; - screenPtr = (uint8_t *)&dst->pixels[(uint16_t)backup->y * SURFACE_BYTES_PER_ROW + ((uint16_t)backup->x >> 1)]; - fn = (CopyFn)(codegenArenaBase() + sp->slot->offset + routeOffset); - fn(backup->bytes, screenPtr); -} - -#endif +// The spriteCompiled* runtime dispatchers moved to +// src/core/spriteDispatch.h as always-inline functions consumed by +// sprite.c (NATIVE-PERF Phase 2 R5). This TU keeps only the +// compile-time machinery. diff --git a/src/core/draw.c b/src/core/draw.c index 0726a0b..749df11 100644 --- a/src/core/draw.c +++ b/src/core/draw.c @@ -629,8 +629,20 @@ void jlDrawPixel(jlSurfaceT *s, int16_t x, int16_t y, uint8_t colorIndex) { if ((uint16_t)x >= SURFACE_WIDTH || (uint16_t)y >= SURFACE_HEIGHT) { return; } +#ifdef JOEYLIB_PLATFORM_IIGS + // Fused plot+mark in ONE JSL for the stage (NATIVE-PERF Phase 2 + // R10); non-stage surfaces carry no dirty state, so the plain plot + // suffices there. The batching plotters (plotPixelNoMark) keep the + // unfused inner. + if (s == gStage) { + iigsDrawPixelMark((uint16_t)x, (uint16_t)y, (uint16_t)(colorIndex & 0x0Fu)); + return; + } + jlpDrawPixel(s, (uint16_t)x, (uint16_t)y, colorIndex); +#else jlpDrawPixel(s, (uint16_t)x, (uint16_t)y, colorIndex); surfaceMarkDirtyRect(s, x, y, 1, 1); +#endif } diff --git a/src/core/port.h b/src/core/port.h index 4733ebf..00763d6 100644 --- a/src/core/port.h +++ b/src/core/port.h @@ -54,15 +54,19 @@ // ===================================================================== extern void iigsDrawPixelInner (uint8_t *pixels, uint16_t x, uint16_t y, uint16_t nibble); +extern void iigsDrawPixelMark (uint16_t x, uint16_t y, uint16_t nibble); extern void iigsDrawLineInner (uint8_t *pixels, uint16_t x0, uint16_t y0, uint16_t x1, uint16_t y1, uint16_t nibble); extern void iigsDrawCircleInner (uint8_t *pixels, uint16_t cx, uint16_t cy, uint16_t r, uint16_t nibble); extern void iigsFillCircleInner (uint8_t *pixels, uint16_t cx, uint16_t cy, uint16_t r, uint16_t nibble); extern void iigsSurfaceClearInner(uint8_t *pixels, uint16_t fillWord); extern void iigsSurfaceClearFastInner(uint8_t *pixels, uint16_t fillWord); extern void iigsTileFillInner (uint8_t *dstRow0, uint16_t nibble); +extern void iigsTileFillMark (uint16_t nibble, uint16_t bx, uint16_t by); extern void iigsTileCopyInner (uint8_t *dstRow0, const uint8_t *srcRow0); +extern void iigsTileCopyMark (const uint8_t *srcRow0, uint16_t dstBx, uint16_t dstBy); extern void iigsTileCopyMaskedInner(uint8_t *dstRow0, const uint8_t *srcRow0, uint16_t transparent); extern void iigsTilePasteInner (uint8_t *dstRow0, const uint8_t *srcTilePixels); +extern void iigsTilePasteMark (const uint8_t *srcTilePixels, uint16_t bx, uint16_t by); extern void iigsTilePasteMonoInner(uint8_t *dstRow0, const uint8_t *monoTile, uint16_t fgColor, uint16_t bgColor); extern void iigsTileSnapInner (uint8_t *dstTilePixels, const uint8_t *srcRow0); extern void iigsFillRectInner (uint8_t *pixels, uint16_t x, uint16_t y, uint16_t w, uint16_t h, uint16_t nibble); @@ -133,19 +137,32 @@ extern jlSurfaceT *gStage; // forward the args. by/bx are tile coords -> bx*4 + by*8*160 byte // offset within the surface. Use SURFACE_ROW_OFFSET (LUT lookup) to // dodge a software multiply helper for the *160 multiply. +// tileFill / tileCopy / tilePaste are STAGE-GATED onto the FUSED Mark +// entries (NATIVE-PERF Phase 2 Step D / R2): for the stage, ONE JSL +// derives the destination in asm from the pinned stage base + the +// gRowOffsetLut, runs the op, and widens the dirty band -- replacing +// C-side pointer math plus two JSLs (op inner + iigsMarkDirtyRowsInner +// via tile.c's tileMarkDirty, which is compiled out for these three +// ops on IIgs). Non-stage surfaces keep the pointer-based inners. // tileFill: the macro passes the raw nibble (the jlp layer owns the & 0x0F); // the asm derives the $NNNN fill word (PERF-AUDIT #41 -- the old C-side // * 0x1111 lowered to a software-multiply helper for a runtime color). #define jlpTileFill(_s, _bx, _by, _c) \ - iigsTileFillInner(&(_s)->pixels[SURFACE_ROW_OFFSET((uint16_t)(_by) << 3) \ - + ((uint16_t)(_bx) << 2)], \ - (uint16_t)((_c) & 0x0F)) + ((_s) == gStage \ + ? iigsTileFillMark((uint16_t)((_c) & 0x0F), (uint16_t)(_bx), (uint16_t)(_by)) \ + : iigsTileFillInner(&(_s)->pixels[SURFACE_ROW_OFFSET((uint16_t)(_by) << 3) \ + + ((uint16_t)(_bx) << 2)], \ + (uint16_t)((_c) & 0x0F))) // tileCopy / tileCopyMasked: take surface+tile-coords; the asm macro derives // row-0 pointers (row = by*8 -> SURFACE_ROW_OFFSET(by<<3); col byte = bx*4 -> bx<<2). +// tileCopy's src may be ANY surface, so its row-0 pointer stays C-derived +// even on the fused stage-dst arm. #define jlpTileCopy(_d, _dbx, _dby, _s, _sbx, _sby) \ - iigsTileCopyInner(&(_d)->pixels[SURFACE_ROW_OFFSET((uint16_t)(_dby) << 3) + ((uint16_t)(_dbx) << 2)], \ - &(_s)->pixels[SURFACE_ROW_OFFSET((uint16_t)(_sby) << 3) + ((uint16_t)(_sbx) << 2)]) + ((_d) == gStage \ + ? iigsTileCopyMark(&(_s)->pixels[SURFACE_ROW_OFFSET((uint16_t)(_sby) << 3) + ((uint16_t)(_sbx) << 2)], (uint16_t)(_dbx), (uint16_t)(_dby)) \ + : iigsTileCopyInner(&(_d)->pixels[SURFACE_ROW_OFFSET((uint16_t)(_dby) << 3) + ((uint16_t)(_dbx) << 2)], \ + &(_s)->pixels[SURFACE_ROW_OFFSET((uint16_t)(_sby) << 3) + ((uint16_t)(_sbx) << 2)])) #define jlpTileCopyMasked(_d, _dbx, _dby, _s, _sbx, _sby, _t) \ iigsTileCopyMaskedInner(&(_d)->pixels[SURFACE_ROW_OFFSET((uint16_t)(_dby) << 3) + ((uint16_t)(_dbx) << 2)], \ @@ -155,7 +172,9 @@ extern jlSurfaceT *gStage; // tilePaste / tileSnap: take surface+coords + the chunky tile buffer; the asm // macros derive row-0. #define jlpTilePaste(_d, _bx, _by, _ct) \ - iigsTilePasteInner(&(_d)->pixels[SURFACE_ROW_OFFSET((uint16_t)(_by) << 3) + ((uint16_t)(_bx) << 2)], (_ct)) + ((_d) == gStage \ + ? iigsTilePasteMark((_ct), (uint16_t)(_bx), (uint16_t)(_by)) \ + : iigsTilePasteInner(&(_d)->pixels[SURFACE_ROW_OFFSET((uint16_t)(_by) << 3) + ((uint16_t)(_bx) << 2)], (_ct))) #define jlpTileSnap(_s, _bx, _by, _out) \ iigsTileSnapInner((_out), &(_s)->pixels[SURFACE_ROW_OFFSET((uint16_t)(_by) << 3) + ((uint16_t)(_bx) << 2)]) @@ -296,6 +315,36 @@ extern uint16_t iigsGetTickWord(void); #endif /* JOEYLIB_PLATFORM_IIGS */ +// ===================================================================== +// 68k tile ops as always-inline headers (NATIVE-PERF Phase 2 R4a). +// +// The ST/Amiga tile overrides live in stTile.h / amigaTile.h as +// static inline functions named stTile* / amigaTile*. Aliasing the +// jlpTile* dispatch names to them HERE (before the dispatch tail) +// makes the tail's #if !defined(jlpTile*) guards skip their extern +// prototypes -- same mechanism as the IIgs macro block above -- and +// every caller (tile.c wrappers, genericTile.c fallback routing) +// inlines the op instead of paying a second cdecl call. The headers +// themselves are #included at the END of this file so their +// portData == NULL fallbacks can see the jlpGenericTile* prototypes +// declared in the dispatch tail. +// ===================================================================== +#if defined(JOEYLIB_PLATFORM_ATARIST) + #define jlpTileFill stTileFill + #define jlpTileCopy stTileCopy + #define jlpTileCopyMasked stTileCopyMasked + #define jlpTilePaste stTilePaste + #define jlpTilePasteMono stTilePasteMono + #define jlpTileSnap stTileSnap +#elif defined(JOEYLIB_PLATFORM_AMIGA) + #define jlpTileFill amigaTileFill + #define jlpTileCopy amigaTileCopy + #define jlpTileCopyMasked amigaTileCopyMasked + #define jlpTilePaste amigaTilePaste + #define jlpTilePasteMono amigaTilePasteMono + #define jlpTileSnap amigaTileSnap +#endif + // ===================================================================== // Chunky planar-hook elision for the other chunky ports (PERF-AUDIT #42). // @@ -833,4 +882,13 @@ void jlpGenericBigFree(void *p); #define jlpBigFree(_p) jlpGenericBigFree((_p)) #endif #endif + +// 68k tile-op inline bodies (see the alias block above the dispatch +// tail). Included LAST so they can call the jlpGenericTile* +// prototypes declared above. +#if defined(JOEYLIB_PLATFORM_ATARIST) + #include "stTile.h" +#elif defined(JOEYLIB_PLATFORM_AMIGA) + #include "amigaTile.h" +#endif #endif diff --git a/src/core/sprite.c b/src/core/sprite.c index 5b176df..8a13cba 100644 --- a/src/core/sprite.c +++ b/src/core/sprite.c @@ -12,6 +12,7 @@ #include "codegenArenaInternal.h" #include "port.h" #include "spriteInternal.h" +#include "spriteDispatch.h" #include "surfaceInternal.h" diff --git a/src/core/spriteDispatch.h b/src/core/spriteDispatch.h new file mode 100644 index 0000000..1533cb4 --- /dev/null +++ b/src/core/spriteDispatch.h @@ -0,0 +1,488 @@ +// Compiled-sprite runtime dispatchers as always-inline functions +// (NATIVE-PERF Phase 2 R5). Bodies moved verbatim from +// src/codegen/spriteCompile.c so the hot entries in src/core/sprite.c +// inline them -- deleting a cross-TU call per compiled sprite op +// (~130 cycles on the 68000, more on the 65816 with its JSL + frame) +// and letting the compiler share the slot/field loads the caller +// already performed. sprite.c is the ONLY consumer; include this +// header nowhere else. The compile-time machinery (jlSpriteCompile, +// emitters, arena plumbing) stays in spriteCompile.c. +#ifndef JOEY_SPRITE_DISPATCH_H +#define JOEY_SPRITE_DISPATCH_H + +#include +#include + +#include "codegenArenaInternal.h" +#include "spriteInternal.h" +#include "surfaceInternal.h" + +#if defined(JOEYLIB_PLATFORM_AMIGA) +#include "amigaPlanar.h" +#elif defined(JOEYLIB_PLATFORM_ATARIST) +#include "stPlanar.h" +#endif + +#if defined(JOEYLIB_PLATFORM_IIGS) + +// SURFACE_ROW_OFFSET dispatches to the gRowOffsetLut lookup on IIgs; +// declared in surfaceInternal.h. Uses a single indexed long-mode read +// for the multiply. + +// IIgs sprite dispatch. The emitted routines (spriteEmitIigs.c) are ordinary +// C-ABI functions called through C function pointers, so the compiler emits the +// cross-bank JSL itself (large memory model -> 24-bit pointers). The DRAW +// routine takes the 24-bit dst pointer in A:X and sets up its own DBR (needed +// for `sta abs,Y`, whose high byte comes from DBR) + Y in a fixed prologue, then +// restores DBR before returning. SAVE/RESTORE take packed src/dst offsets and +// have only their MVN bank operands patched per call (a plain C pointer write +// into the arena from the dispatcher). No self-modifying call stub, no inline +// asm -- the machinery that hid the old un-instrumentable #19 corruption. +static inline __attribute__((always_inline)) void spriteCompiledDraw(jlSurfaceT *dst, const jlSpriteT *sp, int16_t x, int16_t y, uint16_t routeOffset) { + uint8_t *dstRow0; + uint32_t fnAddr; + void (*drawFn)(uint8_t *); + + // Clean C-ABI dispatch. The emitted DRAW routine is a normal C function + // -- void draw(uint8_t *dstRow0) -- so the compiler emits the indirect + // JSL itself: NO self-modifying call stub, NO inline asm, NO hand-rolled + // per-call bank/offset patching (that machinery was the source of the + // old un-instrumentable corruption). dstRow0 (the sprite's top-left + // screen byte) is passed in A:X; the routine's prologue (spriteEmitIigs.c) + // sets up its own DBR/Y and restores DBR on exit. routeOffset is the + // gate-derived DRAW offset for this shift (#29). + dstRow0 = &dst->pixels[SURFACE_ROW_OFFSET(y) + ((uint16_t)x >> 1)]; + fnAddr = codegenArenaBaseAddr() + + sp->slot->offset + + (uint32_t)routeOffset; + drawFn = (void (*)(uint8_t *))fnAddr; + drawFn(dstRow0); +} + + +// patchMvnBanks stamps the dst/src bank operand bytes into each MVN inside an +// emitted C-ABI save/restore routine -- the routine's only runtime-dynamic +// bytes, since MVN's banks are immediate operands and the screen/backup banks +// vary at runtime. The patch is a plain C pointer write into the arena from +// the dispatcher (no self-modifying call stub). C-ABI routine layout from +// spriteEmitIigs.c::emitMvnCabiRoutine: +// prologue (4): 8B DA AA 7A +// row 0 (6): A9 lo hi 54 db sb +// row R (12): 8A/98 18 69 lo hi AA/A8 A9 lo hi 54 db sb +// epilogue (2): AB 6B +// The 4-byte prologue shifts MVN dstbk to routine offset (12*R + 8), srcbk to +// (12*R + 9). +static inline void patchMvnBanks(uint8_t *routine, uint16_t heightPx, uint8_t dstBank, uint8_t srcBank) { + uint16_t r; + + for (r = 0; r < heightPx; r++) { + routine[12u * r + 8u] = dstBank; + routine[12u * r + 9u] = srcBank; + } +} + + + +// Split a 24-bit pointer into its low 16 bits + bank byte. To avoid +// invoking a 32-bit-shift helper for the `>> 16` to extract the bank +// byte, we cast to uint32_t and then byte-alias the storage -- gets +// the same bytes with simple loads. +#define SPLIT_POINTER(_ptr, _outLo, _outBank) \ + do { \ + uint32_t spAddr_ = (uint32_t)(_ptr); \ + const uint8_t *spB_ = (const uint8_t *)&spAddr_; \ + *(_outLo) = (uint16_t)(spB_[0] | ((uint16_t)spB_[1] << 8)); \ + *(_outBank) = spB_[2]; \ + } while (0) + +// The backup->bytes pointer is split with SPLIT_POINTER directly. A +// single-entry global split cache used to live here (#59); it assumed +// one backup buffer per program, but the normal pattern is one buffer +// per moving sprite, so with 2+ sprites the cached pointer ping-ponged +// and every call paid a 24-bit compare AND the full split. The split +// itself is ~6 simple loads/stores -- cheaper than the compare on a +// hit -- so the cache was deleted rather than moved per-sprite. + + +static inline __attribute__((always_inline)) void spriteCompiledSaveUnder(const jlSurfaceT *src, jlSpriteT *sp, int16_t x, int16_t y, jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { + int16_t clippedX; + uint16_t heightPx; + uint16_t copyBytes; + uint16_t screenLo; + uint16_t backupLo; + uint8_t screenBank; + uint8_t backupBank; + uint32_t fnAddr; + uint8_t *routine; + uint8_t *screenPtr; + uint16_t cacheIdx; /* shift * SPRITE_OP_COUNT + SPRITE_OP_SAVE, computed once */ + uint8_t *cachedDst; /* &sp->cachedDstBank[0][0] + cacheIdx */ + uint8_t *cachedSrc; /* &sp->cachedSrcBank[0][0] + cacheIdx */ + + // shift and routeOffset arrive pre-derived from the dispatcher + // gate (#29); widthPx/heightPx are create-time constants cached in + // the sprite struct. + clippedX = (int16_t)(x & ~1); + heightPx = sp->heightPx; + copyBytes = (uint16_t)((sp->widthPx >> 1) + shift); /* shift is x & 1, exactly 0 or 1 */ + + screenPtr = (uint8_t *)&src->pixels[SURFACE_ROW_OFFSET(y) + ((uint16_t)clippedX >> 1)]; + SPLIT_POINTER(screenPtr, &screenLo, &screenBank); + SPLIT_POINTER(backup->bytes, &backupLo, &backupBank); + + backup->sprite = sp; + backup->x = clippedX; + backup->y = y; + backup->width = (uint16_t)(copyBytes << 1); + backup->height = heightPx; + /* sizeBytes is constant per (sprite, shift); cache to dodge the + * per-call ~CUMUL2 (uint16_t * uint16_t) helper. The byte-pointer + * arithmetic avoids reintroducing ~MUL4 for the uint16_t array + * indexing. */ + { + uint16_t *sizeCachePtr = (uint16_t *)((uint8_t *)sp->cachedSizeBytes + ((uint16_t)shift << 1)); + if (*sizeCachePtr == 0) { + *sizeCachePtr = (uint16_t)(copyBytes * heightPx); + } + backup->sizeBytes = *sizeCachePtr; + } + + /* Compute the 1D index into the cached* bank arrays once; + * SPRITE_ROUTINE_INDEX (spriteInternal.h) owns the multiply-free + * derivation. */ + cacheIdx = SPRITE_ROUTINE_INDEX(shift, SPRITE_OP_SAVE); + cachedDst = (uint8_t *)sp->cachedDstBank + cacheIdx; + cachedSrc = (uint8_t *)sp->cachedSrcBank + cacheIdx; + + fnAddr = codegenArenaBaseAddr() + + sp->slot->offset + + (uint32_t)routeOffset; + + // Patch the MVN bank operands only if the dst/src bank pair changed since + // last call (the routine's only dynamic bytes). SAVE copies screen -> + // backup: dst = backup bank, src = screen bank. + if (*cachedDst != backupBank || *cachedSrc != screenBank) { + routine = (uint8_t *)fnAddr; + patchMvnBanks(routine, heightPx, /*dst*/backupBank, /*src*/screenBank); + *cachedDst = backupBank; + *cachedSrc = screenBank; + } + + // Clean C-ABI dispatch -- void save(uint32_t packed), packed = srcOff | + // (dstOff << 16). The compiler emits the indirect JSL: no self-modifying + // call stub, no inline asm. SAVE: src = screen offset, dst = backup offset. + { + uint32_t packed = (uint32_t)screenLo | ((uint32_t)backupLo << 16); + ((void (*)(uint32_t))fnAddr)(packed); + } +} + + +static inline __attribute__((always_inline)) void spriteCompiledRestoreUnder(jlSurfaceT *dst, const jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { + uint16_t heightPx; + uint16_t screenLo; + uint16_t backupLo; + uint8_t screenBank; + uint8_t backupBank; + uint32_t fnAddr; + uint8_t *routine; + uint8_t *screenPtr; + jlSpriteT *sp; + uint16_t cacheIdx; /* shift * SPRITE_OP_COUNT + SPRITE_OP_RESTORE, computed once */ + uint8_t *cachedDst; + uint8_t *cachedSrc; + + // shift and routeOffset arrive pre-derived from the + // jlSpriteRestoreUnder gate (#29) -- no per-call re-inference of + // the shift from the backup here. + sp = backup->sprite; + heightPx = backup->height; + + screenPtr = (uint8_t *)&dst->pixels[SURFACE_ROW_OFFSET(backup->y) + ((uint16_t)backup->x >> 1)]; + SPLIT_POINTER(screenPtr, &screenLo, &screenBank); + SPLIT_POINTER(backup->bytes, &backupLo, &backupBank); + + /* Hoist 2D-array indexing -- see save-side comment. */ + cacheIdx = SPRITE_ROUTINE_INDEX(shift, SPRITE_OP_RESTORE); + cachedDst = (uint8_t *)sp->cachedDstBank + cacheIdx; + cachedSrc = (uint8_t *)sp->cachedSrcBank + cacheIdx; + + fnAddr = codegenArenaBaseAddr() + + sp->slot->offset + + (uint32_t)routeOffset; + + // Patch the MVN bank operands only if they changed (the routine's only + // dynamic bytes). RESTORE copies backup -> screen: dst = screen bank, + // src = backup bank. + if (*cachedDst != screenBank || *cachedSrc != backupBank) { + routine = (uint8_t *)fnAddr; + patchMvnBanks(routine, heightPx, /*dst*/screenBank, /*src*/backupBank); + *cachedDst = screenBank; + *cachedSrc = backupBank; + } + + // Clean C-ABI dispatch -- void restore(uint32_t packed), packed = srcOff | + // (dstOff << 16). RESTORE: src = backup offset, dst = screen offset. + { + uint32_t packed = (uint32_t)backupLo | ((uint32_t)screenLo << 16); + ((void (*)(uint32_t))fnAddr)(packed); + } +} + + +#elif defined(JOEYLIB_PLATFORM_AMIGA) + +// Amiga planar dispatchers. spriteEmitPlanar68k.c emits DRAW routines +// with a cdecl(p0, p1, p2, p3) signature that write directly to +// bitplanes, one pre-shifted variant per x % 8 phase (all 8 compiled +// since NATIVE-PERF Phase 1). Compute byteOff = y*40 + x/8 and pass +// plane[i]+byteOff as the 4 plane args. Every caller gate in +// src/core/sprite.c checks routineOffsets[shift][op] != +// SPRITE_NOT_COMPILED before calling here (#29); anything uncompiled +// (clip, degrade-tier truncation) falls back to the interpreted +// jlpSpriteDrawPlanes / jlpSpriteSavePlanes / jlpSpriteRestorePlanes. +// +// Plane pointers come straight from portData via amigaPlanar.h's +// inline accessor (NATIVE-PERF Phase 2) -- the old jlpSurfacePlanePtr +// cross-TU call cost ~190 cycles PER PLANE, x4 per dispatch. The +// pd == NULL gate preserves the "not a planar surface -> no-op" +// contract those calls provided. + +#include "amigaPlanar.h" + +static inline __attribute__((always_inline)) void spriteCompiledDraw(jlSurfaceT *dst, const jlSpriteT *sp, int16_t x, int16_t y, uint16_t routeOffset) { + typedef void (*DrawFn)(uint8_t *p0, uint8_t *p1, uint8_t *p2, uint8_t *p3); + AmigaPlanarT *pd; + uint16_t byteOff; + DrawFn fn; + + pd = amigaSurfacePlanar(dst); + if (pd == NULL) { + return; + } + byteOff = (uint16_t)((uint16_t)y * AMIGA_BYTES_PER_ROW + ((uint16_t)x >> 3)); + fn = (DrawFn)(codegenArenaBase() + sp->slot->offset + routeOffset); + fn(pd->planes[0] + byteOff, pd->planes[1] + byteOff, pd->planes[2] + byteOff, pd->planes[3] + byteOff); +} + + +// Compiled save/restore (NATIVE-PERF Phase 1): the emitted routines are +// cdecl(p0, p1, p2, p3, buf) plane-run copiers over the 16-px-group +// window (bx = x & ~15, nGroups = SPRITE_GROUPS0 + class). shift here +// is the width CLASS from SPRITE_SAVE_CLASS / SPRITE_RESTORE_CLASS, +// not the draw bit-phase. Row displacements are baked into the +// routine; the window's x phase enters only through the plane base +// pointers (byteOff is even by construction -- bx % 16 == 0 -- so the +// emitted move.w/move.l are 68000-legal). +static inline __attribute__((always_inline)) void spriteCompiledSaveUnder(const jlSurfaceT *src, jlSpriteT *sp, int16_t x, int16_t y, jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { + typedef void (*CopyFn)(uint8_t *p0, uint8_t *p1, uint8_t *p2, uint8_t *p3, uint8_t *buf); + AmigaPlanarT *pd; + int16_t bx; + uint16_t nGroups; + uint16_t byteOff; + CopyFn fn; + + pd = amigaSurfacePlanar(src); + if (pd == NULL) { + return; + } + bx = (int16_t)(x & ~15); + nGroups = (uint16_t)(SPRITE_GROUPS0(sp->widthPx) + shift); + byteOff = (uint16_t)((uint16_t)y * AMIGA_BYTES_PER_ROW + ((uint16_t)bx >> 3)); + + backup->x = bx; + backup->y = y; + backup->width = (uint16_t)(nGroups << 4); + backup->height = sp->heightPx; + backup->sizeBytes = (uint16_t)((uint16_t)(nGroups << 3) * sp->heightPx); + + fn = (CopyFn)(codegenArenaBase() + sp->slot->offset + routeOffset); + fn(pd->planes[0] + byteOff, pd->planes[1] + byteOff, pd->planes[2] + byteOff, pd->planes[3] + byteOff, backup->bytes); +} + + +static inline __attribute__((always_inline)) void spriteCompiledRestoreUnder(jlSurfaceT *dst, const jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { + typedef void (*CopyFn)(uint8_t *p0, uint8_t *p1, uint8_t *p2, uint8_t *p3, uint8_t *buf); + AmigaPlanarT *pd; + uint16_t byteOff; + CopyFn fn; + + (void)shift; + pd = amigaSurfacePlanar(dst); + if (pd == NULL) { + return; + } + byteOff = (uint16_t)((uint16_t)backup->y * AMIGA_BYTES_PER_ROW + ((uint16_t)backup->x >> 3)); + fn = (CopyFn)(codegenArenaBase() + backup->sprite->slot->offset + routeOffset); + fn(pd->planes[0] + byteOff, pd->planes[1] + byteOff, pd->planes[2] + byteOff, pd->planes[3] + byteOff, backup->bytes); +} + +#elif defined(JOEYLIB_PLATFORM_ATARIST) + +// ST word-interleaved planar runtime dispatch. The JIT routine takes +// one arg: groupBase = pd->base + y*160 + (x>>4)*8 (the address of +// the first 16-pixel group the sprite touches). It walks rows by +// adda.w #160 at the end of each row. Per (row, tile_col, plane) it +// emits up to one move.b / clr.b / andi.b+ori.b / ori.b chain at +// d16(a0). +// +// shift selection (in spriteInternal.h SPRITE_SHIFT_INDEX): the full +// intra-group phase x & 15. Phases 0/8 are the unrolled byte-aligned +// JIT; the 14 sub-byte phases are thunk + pre-shifted data tables +// driven by stSpriteDrawShiftWalker (all compiled since NATIVE-PERF +// Phase 1). Anything uncompiled falls back to jlpSpriteDrawPlanes via +// the caller gates (#29). + +// ST word-interleaved row stride: 320 px * 4 planes / 8 bits = 160 +// bytes/row. Numerically equal to SURFACE_BYTES_PER_ROW but a distinct +// fact (group-interleaved layout, not chunky), so it gets its own +// name. Each 16-pixel group is 4 plane words = 8 bytes. +#define ST_GROUP_BYTES_PER_ROW (SURFACE_WIDTH * 4 / 8) +#define ST_BYTES_PER_PLANE_GROUP 8u + +// Buffer base comes straight from portData via stPlanar.h's inline +// accessor (NATIVE-PERF Phase 2); the pd == NULL gate preserves the +// "not a planar surface -> no-op" contract the old jlpSurfacePlanePtr +// cross-TU call provided. + +#include "stPlanar.h" + +static inline __attribute__((always_inline)) void spriteCompiledDraw(jlSurfaceT *dst, const jlSpriteT *sp, int16_t x, int16_t y, uint16_t routeOffset) { + typedef void (*DrawFn)(uint8_t *groupBase); + StPlanarT *pd; + uint8_t *groupBase; + DrawFn fn; + + pd = stSurfacePlanar(dst); + if (pd == NULL) { + return; + } + groupBase = pd->base + + (uint16_t)y * ST_GROUP_BYTES_PER_ROW + + (uint16_t)((uint16_t)x >> 4) * ST_BYTES_PER_PLANE_GROUP; + fn = (DrawFn)(codegenArenaBase() + sp->slot->offset + routeOffset); + fn(groupBase); +} + + +// Compiled save/restore (NATIVE-PERF Phase 1): the emitted routines are +// cdecl(src, dst) raw-span copiers over the 16-px-group window +// (bx = x & ~15, nGroups = SPRITE_GROUPS0 + class); a group window's +// rows are CONTIGUOUS byte spans in the interleaved layout, so the +// routine is a pure move.l chain with the inter-row stride baked in. +// shift here is the width CLASS, not the draw bit-phase. rowStart is +// always even (bx % 16 == 0, base is long-aligned), so move.l is +// 68000-legal. +static inline __attribute__((always_inline)) void spriteCompiledSaveUnder(const jlSurfaceT *src, jlSpriteT *sp, int16_t x, int16_t y, jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { + typedef void (*CopyFn)(const uint8_t *src, uint8_t *dst); + StPlanarT *pd; + int16_t bx; + uint16_t nGroups; + uint8_t *rowStart; + CopyFn fn; + + pd = stSurfacePlanar(src); + if (pd == NULL) { + return; + } + bx = (int16_t)(x & ~15); + nGroups = (uint16_t)(SPRITE_GROUPS0(sp->widthPx) + shift); + rowStart = pd->base + + (uint16_t)y * ST_GROUP_BYTES_PER_ROW + + (uint16_t)((uint16_t)bx >> 4) * ST_BYTES_PER_PLANE_GROUP; + + backup->x = bx; + backup->y = y; + backup->width = (uint16_t)(nGroups << 4); + backup->height = sp->heightPx; + backup->sizeBytes = (uint16_t)((uint16_t)(nGroups << 3) * sp->heightPx); + + fn = (CopyFn)(codegenArenaBase() + sp->slot->offset + routeOffset); + fn(rowStart, backup->bytes); +} + + +static inline __attribute__((always_inline)) void spriteCompiledRestoreUnder(jlSurfaceT *dst, const jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { + typedef void (*CopyFn)(const uint8_t *src, uint8_t *dst); + StPlanarT *pd; + uint8_t *rowStart; + CopyFn fn; + + (void)shift; + pd = stSurfacePlanar(dst); + if (pd == NULL) { + return; + } + rowStart = pd->base + + (uint16_t)backup->y * ST_GROUP_BYTES_PER_ROW + + (uint16_t)((uint16_t)backup->x >> 4) * ST_BYTES_PER_PLANE_GROUP; + fn = (CopyFn)(codegenArenaBase() + backup->sprite->slot->offset + routeOffset); + fn(backup->bytes, rowStart); +} + +#else + +static inline __attribute__((always_inline)) void spriteCompiledDraw(jlSurfaceT *dst, const jlSpriteT *sp, int16_t x, int16_t y, uint16_t routeOffset) { + typedef void (*DrawFn)(uint8_t *destRow); + uint8_t *destRow; + DrawFn fn; + + destRow = &dst->pixels[(uint16_t)y * SURFACE_BYTES_PER_ROW + ((uint16_t)x >> 1)]; + fn = (DrawFn)(codegenArenaBase() + sp->slot->offset + routeOffset); + fn(destRow); +} + + +// x86 / 68k compiled save: bytes are a cdecl +// void copy(const uint8_t *src, uint8_t *dst) +// that walks heightPx rows of copyBytes from screen (stride +// SURFACE_BYTES_PER_ROW) into the contiguous backup buffer. shift and +// routeOffset arrive pre-derived from the dispatcher gate (#29). +static inline __attribute__((always_inline)) void spriteCompiledSaveUnder(const jlSurfaceT *src, jlSpriteT *sp, int16_t x, int16_t y, jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { + typedef void (*CopyFn)(const uint8_t *src, uint8_t *dst); + int16_t clippedX; + uint16_t heightPx; + uint16_t copyBytes; + uint8_t *screenPtr; + CopyFn fn; + + clippedX = (int16_t)(x & ~1); + heightPx = sp->heightPx; + copyBytes = (uint16_t)((sp->widthPx >> 1) + shift); /* shift is x & 1, exactly 0 or 1 */ + + screenPtr = (uint8_t *)&src->pixels[(uint16_t)y * SURFACE_BYTES_PER_ROW + ((uint16_t)clippedX >> 1)]; + + backup->sprite = sp; + backup->x = clippedX; + backup->y = y; + backup->width = (uint16_t)(copyBytes << 1); + backup->height = heightPx; + backup->sizeBytes = (uint16_t)(copyBytes * heightPx); + + fn = (CopyFn)(codegenArenaBase() + sp->slot->offset + routeOffset); + fn(screenPtr, backup->bytes); +} + + +// Mirror of save: caller swaps arg order so the same emitted shape +// drives backup -> screen. The screen-side stride lives inside the +// emitted bytes, so RESTORE has its own routine bytes (stride is +// applied to dst instead of src). routeOffset arrives pre-derived +// from the jlSpriteRestoreUnder gate (#29); shift is unused here (no +// bank-patch cache off-IIgs) but kept for the shared signature. +static inline __attribute__((always_inline)) void spriteCompiledRestoreUnder(jlSurfaceT *dst, const jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset) { + typedef void (*CopyFn)(const uint8_t *src, uint8_t *dst); + jlSpriteT *sp; + uint8_t *screenPtr; + CopyFn fn; + + (void)shift; + sp = backup->sprite; + screenPtr = (uint8_t *)&dst->pixels[(uint16_t)backup->y * SURFACE_BYTES_PER_ROW + ((uint16_t)backup->x >> 1)]; + fn = (CopyFn)(codegenArenaBase() + sp->slot->offset + routeOffset); + fn(backup->bytes, screenPtr); +} + +#endif + +#endif /* JOEY_SPRITE_DISPATCH_H */ diff --git a/src/core/spriteInternal.h b/src/core/spriteInternal.h index 0d63fb5..82018be 100644 --- a/src/core/spriteInternal.h +++ b/src/core/spriteInternal.h @@ -176,9 +176,10 @@ struct jlSpriteT { // gate-checked entry for the callee's own op; shift feeds the IIgs // bank-patch cache index and the chunky save copyBytes (draw needs // neither, so it only takes routeOffset). -void spriteCompiledDraw (jlSurfaceT *dst, const jlSpriteT *sp, int16_t x, int16_t y, uint16_t routeOffset); -void spriteCompiledSaveUnder (const jlSurfaceT *src, jlSpriteT *sp, int16_t x, int16_t y, jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset); -void spriteCompiledRestoreUnder (jlSurfaceT *dst, const jlSpriteBackupT *backup, uint8_t shift, uint16_t routeOffset); +// The spriteCompiled* runtime dispatchers are always-inline functions +// in spriteDispatch.h, included ONLY by sprite.c (NATIVE-PERF Phase 2 +// R5) -- no extern declarations here or the static inline definitions +// would conflict. // Initialize a freshly-allocated jlSpriteT to the uncompiled/default // state (tile-data fields, 0xFF routine-offset fill, bank/size cache diff --git a/src/core/surface.c b/src/core/surface.c index 24f3139..41711ad 100644 --- a/src/core/surface.c +++ b/src/core/surface.c @@ -345,6 +345,14 @@ bool stageAlloc(void) { #endif stageDirtyClearAll(); paletteInitDefault(gStage); + // Fresh stage = fresh SCB/palette; force the first present after + // (re-)init to upload them. The load-time initializers cover the + // first jlInit, but a jlShutdown/jlInit cycle in one process would + // otherwise leave stale false flags and skip the upload + // (NATIVE-PERF Phase 2 R9 -- ports now trust these flags instead + // of memcmp'ing cached SCB/palette mirrors every present). + gStageScbDirty = true; + gStagePaletteDirty = true; return true; } diff --git a/src/core/tile.c b/src/core/tile.c index 9fdf330..e54befa 100644 --- a/src/core/tile.c +++ b/src/core/tile.c @@ -47,7 +47,15 @@ typedef char tileMarkBandProof[(SURFACE_WORD_INDEX(TILE_PIXELS_PER_SIDE) == 2) ? #ifdef JOEYLIB_PLATFORM_IIGS #define tileMarkDirty(_pixelY, _bx) iigsMarkDirtyRowsInner((uint16_t)(_pixelY), (uint16_t)((_pixelY) + TILE_PIXELS_PER_SIDE), (uint16_t)((uint16_t)(_bx) << 1), (uint16_t)(((uint16_t)(_bx) << 1) + 1u)) +// Fill/copy/paste FUSE the mark into their stage-arm asm entries on +// the IIgs (port.h jlpTile* macros -> iigsTile*Mark, NATIVE-PERF +// Phase 2 Step D), so their C-side mark compiles out here. +// copyMasked/pasteMono still mark separately via tileMarkDirty. +#define tileMarkDirtyFused(_pixelY, _bx) ((void)0) #else +// On non-IIgs ports nothing fuses the mark, so the "fused" spelling +// is the plain helper. +#define tileMarkDirtyFused(_pixelY, _bx) tileMarkDirty((_pixelY), (_bx)) // Explicit 8-row unroll: gcc-6.5 -O2 keeps the equivalent loop rolled // (8 iterations of compare + spill-reload, ~670 cycles); the d16(An) // unroll is ~450. Same widen semantics, row by row, as the loop. @@ -198,7 +206,7 @@ void jlTileCopy(jlSurfaceT *dst, uint8_t dstBx, uint8_t dstBy, const jlSurfaceT // Single op: machine asm/planar override or the generic chunky default. jlpTileCopy(dst, dstBx, dstBy, src, srcBx, srcBy); if (dst == gStage) { - tileMarkDirty(TILE_BLOCK_TO_PIXEL(dstBy), dstBx); + tileMarkDirtyFused(TILE_BLOCK_TO_PIXEL(dstBy), dstBx); } } @@ -229,7 +237,7 @@ void jlTileFill(jlSurfaceT *s, uint8_t bx, uint8_t by, uint8_t colorIndex) { // generic chunky default. jlpTileFill(s, bx, by, colorIndex); if (s == gStage) { - tileMarkDirty(TILE_BLOCK_TO_PIXEL(by), bx); + tileMarkDirtyFused(TILE_BLOCK_TO_PIXEL(by), bx); } } @@ -244,7 +252,7 @@ void jlTilePaste(jlSurfaceT *dst, uint8_t bx, uint8_t by, const jlTileT *in) { // Single op: machine asm/planar override or the generic chunky default. jlpTilePaste(dst, bx, by, &in->pixels[0]); if (dst == gStage) { - tileMarkDirty(TILE_BLOCK_TO_PIXEL(by), bx); + tileMarkDirtyFused(TILE_BLOCK_TO_PIXEL(by), bx); } } diff --git a/src/dos/hal.c b/src/dos/hal.c index d346096..9dc9fc9 100644 --- a/src/dos/hal.c +++ b/src/dos/hal.c @@ -54,10 +54,9 @@ static FILE *gCrashLog = NULL; // (The old $3DA edge-poll frame-counter state lived here; removed with // the poll-rate-dependent counter itself -- see jlpFrameCount.) -// Cached palette from the last present. VGA DAC programming is ~768 -// outportb calls; skip it when the source palette is unchanged. -static uint16_t gCachedPalette[SURFACE_PALETTE_COUNT][SURFACE_COLORS_PER_PALETTE]; -static bool gCacheValid = false; +// VGA DAC programming is ~768 outportb calls; skipped when the core +// gStagePaletteDirty flag (surfaceInternal.h) says the stage palette +// is unchanged -- see uploadPaletteIfNeeded. // Expand LUT for jlpPresent. Maps a 4bpp-packed source byte to the // 2-byte little-endian mode-13h output pair for the current scanline's @@ -222,13 +221,17 @@ static void installCrashLog(void) { } +// `src` is always gStage (jlpPresent's contract), so the core palette +// dirty flag is the change signal, replacing a 512-byte cached-mirror +// memcmp per present (NATIVE-PERF Phase 2 R9). The SCB flag is not +// consumed here: mode 13h has one DAC; the per-line SCB choice is +// applied in the pixel expansion, which reads src->scb directly. static void uploadPaletteIfNeeded(const jlSurfaceT *src) { - if (gCacheValid && memcmp(gCachedPalette, src->palette, sizeof(gCachedPalette)) == 0) { + if (!gStagePaletteDirty) { return; } uploadPalette(src); - memcpy(gCachedPalette, src->palette, sizeof(gCachedPalette)); - gCacheValid = true; + gStagePaletteDirty = false; } diff --git a/src/iigs/joeyDraw.s b/src/iigs/joeyDraw.s index fe74fdf..f9d11f9 100644 --- a/src/iigs/joeyDraw.s +++ b/src/iigs/joeyDraw.s @@ -1270,6 +1270,43 @@ tpmNibPairIdx: ; transparent's low byte from the stack (@12,s in M=8) during ; setup. [0],y = dst, [4],y = src in the per-byte subroutine. ; ==================================================================== +; TMASK_BYTE: one masked byte, expanded inline (NATIVE-PERF Phase 2 +; R6). Body identical to the old tmaskByte subroutine; expanding at +; each of the 32 sites deletes the jsr/rts pair (12 cyc x 32 = 384 +; cyc/call, ~9% of the op). GAS numeric labels stay local per +; expansion. In: M=8, X=16, Y = byte offset, [0]=dst / [4]=src frame. +.macro TMASK_BYTE + lda [4], y ; A = src byte + cmp tmaskDoubled + beq 9f ; both nibbles == transparent + sta tmaskSrc + and #0xF0 + cmp tmaskTHi + beq 7f + sta tmaskOutHi ; src hi (kept in $F0 form) + bra 8f +7: + lda [0], y + and #0xF0 + sta tmaskOutHi +8: + lda tmaskSrc + and #0x0F + cmp tmaskTLo + beq 5f + sta tmaskOutLo ; src lo + bra 6f +5: + lda [0], y + and #0x0F + sta tmaskOutLo +6: + lda tmaskOutHi + ora tmaskOutLo + sta [0], y +9: +.endm + .section .text.iigsTileCopyMaskedInner,"ax" .globl iigsTileCopyMaskedInner iigsTileCopyMaskedInner: @@ -1329,13 +1366,13 @@ iigsTileCopyMaskedInner: ldy #0 ; byte offset tmaskRowLoop: - jsr tmaskByte ; byte 0 + TMASK_BYTE ; byte 0 iny - jsr tmaskByte ; byte 1 + TMASK_BYTE ; byte 1 iny - jsr tmaskByte ; byte 2 + TMASK_BYTE ; byte 2 iny - jsr tmaskByte ; byte 3 + TMASK_BYTE ; byte 3 rep #0x20 ; M=16 for arithmetic on Y .a16 @@ -1347,7 +1384,11 @@ tmaskRowLoop: .a8 dex - bne tmaskRowLoop + beq tmaskRowsDone + brl tmaskRowLoop ; loop body is 4 inlined + ; TMASK_BYTE expansions -- + ; past 8-bit branch range +tmaskRowsDone: rep #0x20 ; M=16 before epilogue .a16 @@ -1361,56 +1402,336 @@ tmaskRowLoop: plp rtl -; tmaskByte: in M=8 X=16, Y holds byte offset. -; - load src byte -; - if both nibbles transparent: skip -; - else: assemble output from src/dst nibbles per nibble's -; non-transparency, store -tmaskByte: - ; Entered (via JSR from the M=8 row loop) with M=8, X=16. - ; Tell the assembler so the immediates below size as 1 - ; byte. RTS returns to the still-M=8 row loop, so no - ; width restore is needed here. + + +; ==================================================================== +; NATIVE-PERF Phase 2 Step D (R2): FUSED stage-tile entries. +; +; The C tile path used to pay pointer math (far LUT read + s->pixels +; double-deref) plus TWO JSLs per stage tile op (op inner + dirty +; marker). These entries do all three in ONE JSL: derive the stage +; destination from the pinned stage base ($01:2000, IIGS_STAGE_PIXELS +; in hal.c) + gRowOffsetLut[by*8] + bx*4, run the op via +; absolute-long,X stores with the stage literal folded into each +; offset, then widen the dirty band via stageTileMark below. The +; jlpTile* macros in port.h gate on s == gStage; non-stage surfaces +; keep taking the pointer-based inners above. +; ==================================================================== + +; stageTileMark: shared fused dirty-mark tail. +; In: A = bx (word), Y = by (word); M=16, X=16. +; Out: M=16, X=16. Clobbers A, X, Y. +; Widens gStageMinWord/gStageMaxWord rows by*8..by*8+7 to the tile +; band [bx*2, bx*2+1] -- byte compares in M=8 through the same +; DBR-inherited abs,x write path as iigsMarkDirtyRowsInner (the +; proven asm-writes-C-global pattern; see the trailer note). +; ONE shared section for the helper + all three fused entries: the +; multi-seg OMF packer places sections atomically, so the 16-bit jsr +; from each entry to stageTileMark stays same-bank by construction +; (same reason tmaskByte lives inside the copyMasked section; 16-bit +; self-references only resolve intra-segment -- see clang-build.sh). +.section .text.iigsFusedTile,"ax" +.globl stageTileMark +stageTileMark: + asl a ; A = bx*2 = minWord + pha ; park minWord + tya + asl a + asl a + asl a ; A = by*8 = first row + tax + pla ; A = minWord (high byte 0) + sep #0x20 .a8 - .i16 - lda [4], y ; A = src byte - cmp tmaskDoubled - beq tmaskSkip ; both nibbles == transparent - - sta tmaskSrc - - ; Decide hi nibble. - and #0xF0 - cmp tmaskTHi - beq tmaskUseDstHi - sta tmaskOutHi ; src hi (kept in $F0 form) - bra tmaskHiDone -tmaskUseDstHi: - lda [0], y - and #0xF0 - sta tmaskOutHi -tmaskHiDone: - - ; Decide lo nibble. - lda tmaskSrc - and #0x0F - cmp tmaskTLo - beq tmaskUseDstLo - sta tmaskOutLo ; src lo - bra tmaskLoDone -tmaskUseDstLo: - lda [0], y - and #0x0F - sta tmaskOutLo -tmaskLoDone: - - lda tmaskOutHi - ora tmaskOutLo - sta [0], y - -tmaskSkip: + ldy #8 ; row counter (X stays 16-bit) +stmLoop: + cmp gStageMinWord,x + bcs stmNoMin ; A >= existing -> already covered + sta gStageMinWord,x ; widen down +stmNoMin: + inc a ; A = maxWord = minWord + 1 + cmp gStageMaxWord,x + bcc stmNoMax ; A < existing -> already covered + beq stmNoMax + sta gStageMaxWord,x ; widen up +stmNoMax: + dec a ; back to minWord + inx + dey + bne stmLoop + rep #0x20 + .a16 rts +; ==================================================================== +; iigsTileFillMark(uint16_t nibble, uint16_t bx, uint16_t by) +; +; ABI: arg0 nibble (word, pre-masked to 0x0..0xF by the jlpTileFill +; macro) in A; bx @4,s, by @6,s on JSL entry -> @8,s / @10,s +; after php+phb+phd. No D-frame: the destination is reached via +; $012000+off,x (absolute long indexed), X = tile byte offset. +; fillWord = nibble * $1111 derived in asm (PERF-AUDIT #41); +; tileScratchPtrs is free scratch (no src pointer here). +; ==================================================================== +.globl iigsTileFillMark +iigsTileFillMark: + php + phb + phd + rep #0x30 + .a16 + .i16 + sta tileScratchPtrs+0 ; nibble + asl a + asl a + asl a + asl a ; nibble << 4 + ora tileScratchPtrs+0 ; nibble * 0x11 + sta tileScratchPtrs+2 + xba ; doubled byte << 8 + ora tileScratchPtrs+2 ; nibble * 0x1111 = fillWord + sta tileScratchPtrs+4 ; park fillWord + ; dst offset = gRowOffsetLut[by*8] + bx*4 + lda 10,s ; by + asl a + asl a + asl a + asl a ; by*16 = (by*8) * 2 = LUT byte index + tax + lda 8,s ; bx + asl a + asl a ; bx*4 + clc + adc gRowOffsetLut,x + tax ; X = stage byte offset of tile row 0 + lda tileScratchPtrs+4 ; fillWord + sta 0x012000+0,x + sta 0x012000+2,x + sta 0x012000+160,x + sta 0x012000+162,x + sta 0x012000+320,x + sta 0x012000+322,x + sta 0x012000+480,x + sta 0x012000+482,x + sta 0x012000+640,x + sta 0x012000+642,x + sta 0x012000+800,x + sta 0x012000+802,x + sta 0x012000+960,x + sta 0x012000+962,x + sta 0x012000+1120,x + sta 0x012000+1122,x + ; fused dirty mark + lda 10,s + tay ; Y = by + lda 8,s ; A = bx + jsr stageTileMark + pld + plb + plp + rtl + +; ==================================================================== +; iigsTilePasteMark(const uint8_t *srcTilePixels, uint16_t bx, uint16_t by) +; +; ABI: arg0 src in A:X (packed 32-byte tile, stride 4); bx @4,s, +; by @6,s on JSL entry -> @8,s / @10,s after php+phb+phd -> +; @12,s / @14,s after the 4-byte src D-frame (phx+pha). The src +; D-frame comes straight from the registers -- no tileScratchPtrs +; round-trip (the fill-style single-pointer frame). Dst is +; $012000+off,x like iigsTileFillMark. +; ==================================================================== +.globl iigsTilePasteMark +iigsTilePasteMark: + php + phb + phd + rep #0x30 + .a16 + .i16 + phx ; src high16 -> 3,s + pha ; src low16 -> 1,s + tsc + inc a + tcd ; [0] = srcTilePixels + ; dst offset = gRowOffsetLut[by*8] + bx*4 + lda 14,s ; by + asl a + asl a + asl a + asl a + tax + lda 12,s ; bx + asl a + asl a + clc + adc gRowOffsetLut,x + tax ; X = stage byte offset of tile row 0 + ; 8 rows: src stride 4 via [0],y; dst stride 160 folded + ; into the long,x literals. + ldy #0 + lda [0], y + sta 0x012000+0,x + ldy #2 + lda [0], y + sta 0x012000+2,x + ldy #4 + lda [0], y + sta 0x012000+160,x + ldy #6 + lda [0], y + sta 0x012000+162,x + ldy #8 + lda [0], y + sta 0x012000+320,x + ldy #10 + lda [0], y + sta 0x012000+322,x + ldy #12 + lda [0], y + sta 0x012000+480,x + ldy #14 + lda [0], y + sta 0x012000+482,x + ldy #16 + lda [0], y + sta 0x012000+640,x + ldy #18 + lda [0], y + sta 0x012000+642,x + ldy #20 + lda [0], y + sta 0x012000+800,x + ldy #22 + lda [0], y + sta 0x012000+802,x + ldy #24 + lda [0], y + sta 0x012000+960,x + ldy #26 + lda [0], y + sta 0x012000+962,x + ldy #28 + lda [0], y + sta 0x012000+1120,x + ldy #30 + lda [0], y + sta 0x012000+1122,x + ; fused dirty mark + lda 14,s + tay ; Y = by + lda 12,s ; A = bx + jsr stageTileMark + ; drop the src D-frame, restore + tsc + clc + adc #4 + tcs + pld + plb + plp + rtl + +; ==================================================================== +; iigsTileCopyMark(const uint8_t *srcRow0, uint16_t dstBx, uint16_t dstBy) +; +; Same shape as iigsTilePasteMark but the source is a SURFACE row +; (stride 160, any bank -- the jlpTileCopy macro derives it in C +; because src may be any surface), so the [0],y offsets mirror the +; dst deltas. Dst is the stage, derived in asm. +; ==================================================================== +.globl iigsTileCopyMark +iigsTileCopyMark: + php + phb + phd + rep #0x30 + .a16 + .i16 + phx ; src high16 -> 3,s + pha ; src low16 -> 1,s + tsc + inc a + tcd ; [0] = srcRow0 + ; dst offset = gRowOffsetLut[dstBy*8] + dstBx*4 + lda 14,s ; dstBy + asl a + asl a + asl a + asl a + tax + lda 12,s ; dstBx + asl a + asl a + clc + adc gRowOffsetLut,x + tax ; X = stage byte offset of tile row 0 + ; 8 rows: src stride 160 via [0],y; dst stride 160 in + ; the long,x literals (parallel offsets = block copy). + ldy #0 + lda [0], y + sta 0x012000+0,x + ldy #2 + lda [0], y + sta 0x012000+2,x + ldy #160 + lda [0], y + sta 0x012000+160,x + ldy #162 + lda [0], y + sta 0x012000+162,x + ldy #320 + lda [0], y + sta 0x012000+320,x + ldy #322 + lda [0], y + sta 0x012000+322,x + ldy #480 + lda [0], y + sta 0x012000+480,x + ldy #482 + lda [0], y + sta 0x012000+482,x + ldy #640 + lda [0], y + sta 0x012000+640,x + ldy #642 + lda [0], y + sta 0x012000+642,x + ldy #800 + lda [0], y + sta 0x012000+800,x + ldy #802 + lda [0], y + sta 0x012000+802,x + ldy #960 + lda [0], y + sta 0x012000+960,x + ldy #962 + lda [0], y + sta 0x012000+962,x + ldy #1120 + lda [0], y + sta 0x012000+1120,x + ldy #1122 + lda [0], y + sta 0x012000+1122,x + ; fused dirty mark + lda 14,s + tay ; Y = dstBy + lda 12,s ; A = dstBx + jsr stageTileMark + ; drop the src D-frame, restore + tsc + clc + adc #4 + tcs + pld + plb + plp + rtl + + ; -------------------------------------------------------------------- ; Scratch for the masked-tile inner loop. Lives outside the @@ -1939,6 +2260,93 @@ dpxlDone: rtl +; ==================================================================== +; iigsDrawPixelMark(uint16_t x, uint16_t y, uint16_t nibble) +; +; FUSED stage-only jlDrawPixel (NATIVE-PERF Phase 2 R10): one JSL does +; the single-row dirty widen (word x>>2, row y -- same abs,x +; DBR-inherited write path as iigsMarkDirtyRowsInner) and the nibble +; RMW plot against the pinned stage base ($01:2000 + gRowOffsetLut[y] +; + (x>>1), long,x with the literal folded in). Caller (jlDrawPixel) +; has clipped and owns the s == gStage gate; the batching plotters +; (plotPixelNoMark for lines/circles) stay on iigsDrawPixelInner. +; +; ABI: x in A; y @4,s, nibble @6,s on JSL entry. Frame is php (+1) +; plus the parked x (+2): x @1,s, y @7,s, nibble @9,s. No PHB (DBR +; never written, PERF-AUDIT #53 precedent), no D use. Exits M=16/X=16. +; No intra-asm jsr, so the section is self-contained (cross-section +; 16-bit jsr is a wild jump -- see the fused-tile block note). +; ==================================================================== +.section .text.iigsDrawPixelMark,"ax" +.globl iigsDrawPixelMark +iigsDrawPixelMark: + php + rep #0x30 + .a16 + .i16 + pha ; park x; y -> 7,s nib -> 9,s + ; --- single-row widen: A still holds x --- + lsr a + lsr a + tay ; Y = wordIdx = x>>2 (<= 79) + lda 7,s ; y + tax ; X = row + tya ; A = wordIdx + sep #0x20 + .a8 + cmp gStageMinWord,x + bcs dpxmNoMin + sta gStageMinWord,x ; widen down +dpxmNoMin: + cmp gStageMaxWord,x + bcc dpxmNoMax + beq dpxmNoMax + sta gStageMaxWord,x ; widen up +dpxmNoMax: + rep #0x20 + .a16 + ; --- plot: byte offset = gRowOffsetLut[y] + (x>>1) --- + lda 7,s ; y + asl a + tax + lda 1,s ; x + lsr a + clc + adc gRowOffsetLut,x + tax ; X = stage byte offset + sep #0x20 + .a8 + lda 1,s ; x low byte (parity in LSB) + and #1 + bne dpxmOdd + lda 9,s ; nibble -> high half (left pixel) + asl a + asl a + asl a + asl a + sta dpxlNibPart + lda 0x012000,x + and #0x0F + ora dpxlNibPart + sta 0x012000,x + bra dpxmDone +dpxmOdd: + lda 9,s ; nibble -> low half (right pixel) + and #0x0F + sta dpxlNibPart + lda 0x012000,x + and #0xF0 + ora dpxlNibPart + sta 0x012000,x +dpxmDone: + rep #0x30 + .a16 + .i16 + pla ; drop parked x + plp + rtl + + .section .bss.dpxlNibPart,"aw" .globl dpxlNibPart dpxlNibPart: