Author SHA1 Message Date
ed 4fbf550d3c hot-reload attempt (unreviewed, not working) 2026-08-06 10:44:34 -04:00
ed 01f7ceba7c buzzing brain. 2026-08-05 02:48:05 -04:00
ed 6f2eff920d some more review before bed. 2026-08-05 02:00:41 -04:00
ed f25765a7b7 Preparing for camera transformation chapter. 2026-08-05 01:21:25 -04:00
ed 2757aa4330 Fix bug with pad input processing (needed mac_yield load fallthrough case) 2026-08-05 01:08:01 -04:00
ed 748b58c5c5 Codebase overhaul. Metaprogram proofread (part 2). Starting to get serious.
Need to rewrite the ps1 lua metaprogram sometime soonish. Getting too bloated... need to consolidate code paths.

In this push codebase structure is starting to get a bit more realized. Decided todo now to match Pikuma's linking module files vods beginning to reorganize its codebase as well.
Atoms & atom components are not in their on *.atom.c files. (Not calling it tape.c as I don't really bake tapes like that outside of the unity c file so far...)

The lua metaprogram has had additional features added to it yet again to avoid hardcoding module handling and supporting multiple atom files per-module.
Either after the camera or cd-rom section I'll be most likely pausing to fully refactor the metaprogram. Possibly as a full re-write to get the loc minimal.
2026-08-04 23:34:00 -04:00
ed 6441dbc23e Proof-reading lua metaprogram (part 1) 2026-08-04 19:32:43 -04:00
ed 57fdb9e037 improvmenets to delay slot modeling (lua metaprogram) 2026-08-04 18:27:00 -04:00
ed b5953a723b add ac_yield_load and ac_yield_tail for delay slot optimization opportunities. 2026-08-04 17:25:02 -04:00
ed 888ffce859 Finished: Pikuma Linking multiple files (not applying to codebase only watched) 2026-08-04 16:49:02 -04:00
ed 7289e7c89c Added jump_rel (can't use abs jump with asm dsl). Fixes + improvements to ps1 asm meta passes. 2026-08-04 16:01:01 -04:00
ed 54a5bb9a31 starting to optimize 2026-08-04 12:59:51 -04:00
ed e0f4ac873d spamming load delay slots for now as a fix... 2026-08-04 09:07:53 -04:00
ed f17fa9165e wip: input was working... messed it up (bios snapshot reads) 2026-08-04 00:50:12 -04:00
ed 8282f8e902 overkill sio cruft, not keeping. 2026-08-03 10:12:06 -04:00
ed 9eb696ece8 drafting 2026-08-02 21:58:57 -04:00
ed 858e57f293 preparing to overhaul input handling 2026-08-02 17:49:24 -04:00
ed afcd9b86f0 Gaining clarity on tape abi.. screen_init atoms done. Time to finish rest of joypad course vods... 2026-08-02 15:19:52 -04:00
ed 43cd4e0344 WIP: working towards minimizing C-ABI & PsyQ CRT usage 2026-08-01 23:11:10 -04:00
ed 09dde54030 Finished(Controller Input): Reading Joypad State 2026-07-31 15:15:51 -04:00
ed 315e1b2c5e Fix(lua atom tape dsl): Bad-hardcode for source file line-table mapping in dwarf injection pass. 2026-07-31 14:28:50 -04:00
ed 02658d3609 Prepare for hello joypad! 2026-07-28 00:35:02 -04:00
ed dbc459b7e0 gte_hello -> hello_gte. gte is done, moving on to controller! 2026-07-28 00:17:23 -04:00
ed a704341fc6 Testing out the metaprogram with some optimization, need to remove some hardcoding later.. 2026-07-27 23:35:03 -04:00
ed 7421b32fd7 redundant nop reduction 2026-07-27 22:49:41 -04:00
ed e2eb74be19 Remove gte component result contracts (was a bad bodge in, for a later directive thats TODO) 2026-07-27 22:49:26 -04:00
ed 338f1fe46e Better reports from dsl metaprogram 2026-07-27 10:06:23 -04:00
ed 27a9038e0d req c11, 2026-07-26 17:36:46 -04:00
ed 8c8d2e54aa remove cruft 2026-07-26 14:40:57 -04:00
ed 80a35aa23a WIP: Better step debug on atom components, better db_skip annotation, lots of curation passes on lua.
Still don't have this thing in its final state for  the curse but its close.
2026-07-26 13:55:47 -04:00
ed f247d56c32 Debug vis ergonomics 2026-07-25 13:19:35 -04:00
ed 590ff1e2ec Curation pass: reduce nested conditional branching in some defnitions. 2026-07-25 13:00:36 -04:00
ed 653e18ee28 remove code related to dry run and dep graph rendering (ps1 meta) 2026-07-25 11:59:41 -04:00
ed ebb876fe89 report.lua: Remove redudnant section formatting/header 2026-07-25 11:25:12 -04:00
ed 1b40b16c0e Review pass. 2026-07-25 11:20:53 -04:00
ed 9ffd6592bc Better static analysis for C0 <-> C2 data race hazards. 2026-07-25 04:09:48 -04:00
ed d56adab38f branch delay slot better support.
Still reviewing. Need to see if gte is handled properly.
2026-07-23 18:35:02 -04:00
ed 08af73d0d2 Lua Metaprogram: Improvements to static analysis + others. 2026-07-23 10:18:30 -04:00
ed 67d54debfa offset corections (dwarf) 2026-07-22 18:00:09 -04:00
ed 3c25306070 fixes 2026-07-22 09:47:01 -04:00
ed c3cf05950e good enough for now 2026-07-21 22:29:22 -04:00
ed f6b4d9895e Adjustments to offset convention (don't want 1s based addresssing to mess with the spec defined encoding) 2026-07-21 20:52:13 -04:00
ed e70361b548 curation: first pass 2026-07-21 19:20:30 -04:00
ed ed3eb45b1d Fixes atom component gdb stepping. New phase/ctx annotations for atoms. Attempt at type views on registers (gdb pretty print failures).
Needs heavy curation and problably simplicication.
2026-07-18 10:29:04 -04:00
ed d7770b6e1d review pass on c code. 2026-07-15 08:56:37 -04:00
ed 137549b1c8 First pass review 2026-07-14 22:55:16 -04:00
ed 7d5b13aadb TODO: need to review snapshot 2026-07-14 12:16:00 -04:00
ed 2d901003f9 Fix off by one ahead issue with stepping into atoms. Support for local register symbols used in atoms + atom bindings locals in gdb. 2026-07-13 12:41:48 -04:00
ed b43d22008e improve step-debug latency 2026-07-12 15:39:35 -04:00
ed 904889b483 general review post-dwarf_injection.lua working 2026-07-12 15:14:59 -04:00
ed f7aa7b75e7 doing dwarf injectiion/mods for the tape atoms. syncs with vscode cursor. 2026-07-12 12:51:41 -04:00
ed aca6e30e20 better debug support 2026-07-11 22:44:32 -04:00
ed 8b0fb1d4e4 exploring gdb support for the atom asm dsl. 2026-07-11 21:02:34 -04:00
ed 9f7a4a00ce final pass on metaprogram 2026-07-11 19:53:12 -04:00
ed 277af1c901 update readme 2026-07-11 17:49:55 -04:00
ed 97d2f66c5a eliminated most lag (runs in ms) 2026-07-11 17:34:52 -04:00
ed d9406553b3 finally starting to approach decent performance. 2026-07-11 17:30:16 -04:00
ed e662d175ab lifting tokenize_body, using lfs package 2026-07-11 16:47:09 -04:00
ed 5387a07b84 progress on static analysis 2026-07-11 15:18:27 -04:00
ed 65d805e3ba start to generalize check rules.. 2026-07-11 14:57:48 -04:00
ed 987f4dee1e preparing for a big refactor 2026-07-11 14:48:57 -04:00
ed df723c691d progress 2026-07-11 14:25:40 -04:00
ed 45ac85c038 lua metaprogram: Delete dead code, some more lifting to duffle 2026-07-11 14:16:29 -04:00
ed 072231c46b Lua Metaprogram: Scan codepaths collapse + more reviews. 2026-07-11 13:45:22 -04:00
ed 2b00956862 Corrections, flatting nested branches (lua metaprogram) 2026-07-11 10:24:34 -04:00
ed 1ffad6cf98 lua metaprogram: more cruft removal. 2026-07-11 09:45:51 -04:00
ed 318516a354 adding comments for scan progress 2026-07-11 02:00:05 -04:00
ed 91a91b3495 mostly comment review (lua metaprogram) 2026-07-11 01:47:38 -04:00
ed a0d22700db lots of cruft to still sift thru 2026-07-11 00:27:28 -04:00
ed 51bdf7106b update_deps.ps1 properly gets lpeg now without jank 2026-07-11 00:14:59 -04:00
ed 531e1cbd58 update readme 2026-07-11 00:11:49 -04:00
ed 541e52de2b adjsutments for the old graphics hello module. 2026-07-11 00:11:32 -04:00
ed eccf17d21c update readme 2026-07-10 23:49:00 -04:00
ed 0d94632edf dealing with this mess still. 2026-07-10 23:36:44 -04:00
ed 798807a9c2 some saved by cahcing git path resolution. 2026-07-10 21:32:53 -04:00
ed e9f26f89b8 review pass on lua scripts related to tape atom metaprogram
script running is slow need to fix.
2026-07-10 21:15:21 -04:00
ed a226b45d18 more adjustments 2026-07-10 21:14:16 -04:00
ed c22e4baa41 minor adjustmnets to some headers (doing a review pass) 2026-07-10 19:51:41 -04:00
ed fa598a41c6 readability pass on word_count_eval.lua 2026-07-10 18:51:45 -04:00
ed a928d06ac9 more improvments to static pass. reduce cruft in build/gen 2026-07-10 17:46:32 -04:00
ed 91c2218471 more static analysis 2026-07-10 14:50:32 -04:00
ed 27a5f8029f improvmenets for atom components 2026-07-10 13:18:23 -04:00
ed 7a168137fc static analysis first pass 2026-07-10 12:01:17 -04:00
ed 2ceb2f2a05 minor changes preparing for static analysis metaprogram and revewing cube_g4_face code. 2026-07-10 09:33:35 -04:00
ed 6103f47f05 reduce cruft 2026-07-10 09:23:02 -04:00
ed c824c998eb broken. 2026-07-10 09:08:29 -04:00
ed 9d066ae292 nesting reduction 2026-07-09 20:10:01 -04:00
ed c9b7f8c08b almost ready for static analysis additions 2026-07-09 19:48:02 -04:00
ed 59903546d7 rework of metaprogram 2026-07-09 19:30:32 -04:00
ed 1ffdda45e5 Adjustments to formatting 2026-07-09 19:28:56 -04:00
ed 1209172649 wip: lua metaprogram rework 2026-07-09 18:45:36 -04:00
ed 98e27c2815 fixed. 2026-07-09 17:21:47 -04:00
ed 88aa1b8b59 wip 2026-07-09 16:57:52 -04:00
ed ca3dc4aff0 wip: cube_g4_face is bugged 2026-07-09 16:17:34 -04:00
ed 1fb4883138 gitignore update 2026-07-09 15:47:11 -04:00
ed 0ad609e7c2 cookin 2026-07-09 15:38:03 -04:00
ed ccdf1b832b more intiution... 2026-07-09 13:30:13 -04:00
ed 4d177bc34d refactor 2026-07-09 11:14:11 -04:00
ed 32a754cd06 FACK. 2026-07-09 10:48:59 -04:00
ed 407c7d352a cube_tri 2026-07-09 10:48:55 -04:00
ed 8541713d0c metaprogram improvements 2026-07-09 01:17:35 -04:00
ed 602a0b46d8 Still learning/de-obfuscating 2026-07-08 21:17:16 -04:00
ed 74f390c3b1 Reviewing post-dsl refactors, more pseudo instructions 2026-07-08 17:14:11 -04:00
ed 5e7da32387 Adjustments to gp docs 2026-07-08 13:38:25 -04:00
ed 5375478044 gp.h improvements 2026-07-08 10:21:31 -04:00
ed 10c8dcdc07 improving dsl: gte. 2026-07-08 00:37:27 -04:00
ed d0b1bae896 improving dsl. 2026-07-08 00:30:02 -04:00
ed 0b147a8b0c need to change symbols.. 2026-07-07 23:28:20 -04:00
ed 101b07fe71 fixes, de-obufscation... still confused about formating color... 2026-07-07 22:12:06 -04:00
77 changed files with 23172 additions and 3772 deletions
+9 -2
View File
@@ -1,8 +1,9 @@
build
toolchain/armips
toolchain/luajit-2.1
toolchain/pcsx-redux
# toolchain/psyq_iwyu
# toolchain/PSn00bSDK
toolchain/psyq_iwyu
toolchain/PSn00bSDK
*.exe
*.elf
@@ -14,3 +15,9 @@ toolchain/pcsx-redux
*.a
.sentry-native
.vscode/settings.json
toolchain/lfs
toolchain/lpeg
scratch
toolchain/libpsn00b
scripts/pcsx_debug_helper.zip
+163 -27
View File
@@ -4,7 +4,7 @@
// For more information, visit: https://go.microsoft.com/fwlink/?linkid=830387
"version": "0.2.0",
"configurations": [
{
{
"name": "Debug: Hello Psy-Q!",
"type": "gdb",
"request": "attach",
@@ -12,6 +12,10 @@
"remote": true,
"cwd": "${workspaceRoot}/build",
"valuesFormatting": "parseText",
"registerLimit": "1-32",
"frameFilters": false,
"showDevDebugOutput": false,
"printCalls": false,
"stopAtConnect": true,
"gdbpath": "gdb-multiarch",
"windows": {
@@ -20,10 +24,17 @@
"osx": {
"gdbpath": "gdb"
},
"executable": "${workspaceRoot}/build/hello_psyq.elf",
"executable": "${workspaceRoot}/build/hello_gte.elf",
"setupCommands": [
{ "text": "set mi-async off" },
{ "text": "set remotetimeout 0" },
{ "text": "set logging file build/gen/hello_gte.gdb.log" },
{ "text": "set logging redirect on" }
],
"autorun": [
"monitor reset shellhalt",
"load hello_psyq.elf",
"load hello_gte.elf",
"source scripts/gdb/gdb_tape_atoms.gdb",
"tbreak main",
"continue"
]
@@ -36,30 +47,10 @@
"remote": true,
"cwd": "${workspaceRoot}/build",
"valuesFormatting": "parseText",
"stopAtConnect": true,
"gdbpath": "gdb-multiarch",
"windows": {
"gdbpath": "gdb-multiarch.exe"
},
"osx": {
"gdbpath": "gdb"
},
"executable": "${workspaceRoot}/build/hello_gpu.elf",
"autorun": [
"monitor reset shellhalt",
"load hello_gpu.elf",
"tbreak main",
"continue"
]
},
{
"name": "Debug: Hello GTE Psy-Q!",
"type": "gdb",
"request": "attach",
"target": "localhost:3333",
"remote": true,
"cwd": "${workspaceRoot}/build",
"valuesFormatting": "parseText",
"registerLimit": "1-32",
"frameFilters": false,
"showDevDebugOutput": false,
"printCalls": false,
"stopAtConnect": true,
"gdbpath": "gdb-multiarch",
"windows": {
@@ -69,12 +60,157 @@
"gdbpath": "gdb"
},
"executable": "${workspaceRoot}/build/hello_gte.elf",
"setupCommands": [
{ "text": "set mi-async off" },
{ "text": "set remotetimeout 0" },
{ "text": "set logging file build/gen/hello_gte.gdb.log" },
{ "text": "set logging redirect on" }
],
"autorun": [
"monitor reset shellhalt",
"load hello_gte.elf",
"tbreak main",
"continue"
]
},
{
"name": "Debug: Hello GTE!",
"type": "gdb",
"request": "attach",
"target": "localhost:3333",
"remote": true,
"cwd": "${workspaceRoot}",
"valuesFormatting": "parseText",
"registerLimit": "1-32",
"frameFilters": false,
"showDevDebugOutput": false,
"printCalls": false,
"stopAtConnect": true,
"gdbpath": "gdb-multiarch",
"windows": {
"gdbpath": "gdb-multiarch.exe"
},
"osx": {
"gdbpath": "gdb"
},
"executable": "${workspaceRoot}/build/hello_gte.dwarf-injected.elf",
"setupCommands": [
{ "text": "set mi-async off" },
{ "text": "set remotetimeout 0" },
{ "text": "set logging file build/gen/hello_gte.gdb.log" },
{ "text": "set logging redirect on" }
],
"autorun": [
"monitor reset shellhalt",
"load build/hello_gte.dwarf-injected.elf",
"source scripts/gdb/gdb_tape_atoms.gdb",
"tbreak main",
"continue"
]
},
{
"name": "Debug: Hello Joypad!",
"type": "gdb",
"request": "attach",
"target": "localhost:3333",
"remote": true,
"cwd": "${workspaceRoot}",
"valuesFormatting": "parseText",
"registerLimit": "1-32",
"frameFilters": false,
"showDevDebugOutput": false,
"printCalls": false,
"stopAtConnect": true,
"gdbpath": "gdb-multiarch",
"windows": {
"gdbpath": "gdb-multiarch.exe"
},
"osx": {
"gdbpath": "gdb"
},
"executable": "${workspaceRoot}/build/hello_joypad.dwarf-injected.elf",
"setupCommands": [
{ "text": "set mi-async off" },
{ "text": "set remotetimeout 0" },
{ "text": "set logging file build/gen/hello_joypad.gdb.log" },
{ "text": "set logging redirect on" }
],
"autorun": [
"monitor reset shellhalt",
"load build/hello_joypad.dwarf-injected.elf",
"source scripts/gdb/gdb_tape_atoms.gdb",
"tbreak main",
"continue"
]
},
{
"name": "Debug: Hello Camera!",
"type": "gdb",
"request": "attach",
"target": "localhost:3333",
"remote": true,
"cwd": "${workspaceRoot}",
"valuesFormatting": "parseText",
"registerLimit": "1-32",
"frameFilters": false,
"showDevDebugOutput": false,
"printCalls": false,
"stopAtConnect": true,
"gdbpath": "gdb-multiarch",
"windows": {
"gdbpath": "gdb-multiarch.exe"
},
"osx": {
"gdbpath": "gdb"
},
"executable": "${workspaceRoot}/build/hello_camera.dwarf-injected.elf",
"setupCommands": [
{ "text": "set mi-async off" },
{ "text": "set remotetimeout 0" },
{ "text": "set logging file build/gen/hello_camera.gdb.log" },
{ "text": "set logging redirect on" }
],
"autorun": [
"monitor reset shellhalt",
"load build/hello_camera.dwarf-injected.elf",
"source scripts/gdb/gdb_tape_atoms.gdb",
"tbreak main",
"continue"
]
},
{
"name": "Debug: Hello Camera! (attach only)",
"type": "gdb",
"request": "attach",
"target": "localhost:3333",
"remote": true,
"cwd": "${workspaceRoot}",
"valuesFormatting": "parseText",
"registerLimit": "1-32",
"frameFilters": false,
"showDevDebugOutput": false,
"printCalls": false,
"stopAtConnect": true,
"gdbpath": "gdb-multiarch",
"windows": {
"gdbpath": "gdb-multiarch.exe"
},
"osx": {
"gdbpath": "gdb"
},
"executable": "${workspaceRoot}/build/hello_camera.dwarf-injected.elf",
"setupCommands": [
{ "text": "set mi-async off" },
{ "text": "set remotetimeout 0" },
{ "text": "set logging file build/gen/hello_camera.gdb.log" },
{ "text": "set logging redirect on" }
],
"autorun": [
"source scripts/gdb/gdb_tape_atoms.gdb",
"tbreak hot_reload_entry",
"continue"
]
}
]
}
-462
View File
@@ -1,462 +0,0 @@
/*
* atom_dsl.h
* ============================================================================
*
* ATOM DSL — annotation layer for tape atoms (lottes_tape.h).
*
* This header turns `__attribute__((annotate(...)))` and `_Pragma(...)` into
* a small named DSL that the metaprogram can validate against.
*
* The C compiler treats every macro below as a no-op:
* - atom_init / atom_terminate / atom_bind / atom_setup / atom_commit /
* atom_annot all expand to `__attribute__((annotate("..."))) MipsAtom_(name)`
* — accepted by GCC (with -Wno-attributes), absent at runtime.
* - atom_resource / atom_region / atom_group / atom_cadence / atom_async
* expand to `_Pragma("...")` — accepted by any C11 preprocessor.
*
* The metaprogram (tape_atom_annotation_pass.lua) reads the source-as-written
* and validates:
* - every MipsAtom_ has one atom_*() annotation (no orphans)
* - phase is recognized (init/bind/setup/work/commit/terminate)
* - reads/writes reference canonical wave-context registers
* - rbind atoms reference a real Binds_* struct declaration
* - word-counts in tapre metadata agree with the body's actual .word count
* - resource/region/group/cadence/async pragmas are spelled correctly and
* reference known enum values
*
* ============================================================================
*
* PUTTING IT ON AN ATOM — the canonical pattern
*
* _tape_resources_
* atom_resource(cube_tri, "model_ship_cube")
* atom_region (cube_tri, PRIM_ARENA)
* atom_group (cube_tri, GROUP_RENDER_PRIMS)
* atom_cadence (cube_tri, CADENCE_FRAME)
*
* atom_annot(cube_tri, phase_work,
* tape_regs(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase),
* tape_regs(R_PrimCursor, R_FaceCursor))
* internal MipsAtom_(cube_tri) {
* atom_label(culling),
* // ... atom body ...
* atom_label(bounds_chk),
* };
*
* atom_offset(culling, bounds_chk) // ← branch target, validated
*
* RBIND pattern — `Binds_*` is the contract
*
* // Wave-context register layout (declarative):
* typedef struct Binds_TrackFaceBatch {
* U4 R_PrimCursor, R_FaceCursor,
* R_VertBase, R_OtBase;
* } Binds_TrackFaceBatch;
*
* atom_resource(rbind_track_face_batch, "track_face_batch_42")
* atom_region (rbind_track_face_batch, HEAP_3D)
* atom_group (rbind_track_face_batch, GROUP_LOAD_FACES)
* atom_cadence (rbind_track_face_batch, CADENCE_ONDEMAND)
* atom_async (rbind_track_face_batch, true)
*
* atom_bind(rbind_track_face_batch, Binds_TrackFaceBatch,
* tape_regs(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase))
* internal MipsAtom_(rbind_track_face_batch) { ... };
*
* Annotation rules
* ----------------
* 1. Each MipsAtom_(name) needs EXACTLY ONE atom_*() macro on the line
* immediately above. No annotation = orphan (warning). Two annotations
* on the same name = duplicate (error).
*
* 2. atom_init and atom_terminate take only the name.
*
* 3. atom_setup and atom_commit take name + reads.
*
* 4. atom_bind takes name + Binds_* type + writes.
*
* 5. atom_annot takes name + phase token + reads + writes.
* Phase tokens: phase_init / phase_bind / phase_setup / phase_work /
* phase_commit / phase_terminate.
*
* 6. Optional pragmas (atom_resource / atom_region / atom_group /
* atom_cadence / atom_async) attach metadata to the atom. They can
* appear in any order, with one per atom. They're independent of the
* atom_*() macro — multiple pragmatics are fine.
*
* ============================================================================
*
* WHY A SEPARATE LAYER (not just put everything in source comments)?
*
* Source comments are invisible to the compiler. Annotations live in the
* source as actual C tokens, so:
* - they can never silently get out of sync with the code (the build
* fails at preprocessing if the metaprogram disagrees)
* - they can be cross-validated against metadata (build fails if a
* WORD_COUNT entry drifts away from the .word count in source)
* - they make the C compiler a witness ("there's a marker here, and
* it's labelled, and it has arguments") without making the C compile
* itself do any work
*
* ============================================================================
*/
#ifdef INTELLISENSE_DIRECTIVES
#pragma once
// #include <stdint.h>
#endif
/* ============================================================================
* PHASE TOKENS — strings, used as the second arg to atom_annot(...)
*
* Why strings? They preserve the metaprogram's ability to read phase directly
* from the source-as-written, even when the macro isn't expanded. The Lua
* tool also has a MACRO_EXPANSION table for resolving phase_* source-level
* references.
*
* atom_annot(cube_tri, phase_work, ...) ← legal
* atom_annot(cube_tri, "work", ...) ← legal (and equivalent)
* atom_annot(cube_tri, phase_setup, ...) ← legal
*
* ============================================================================*/
#define phase_init "init"
#define phase_bind "bind"
#define phase_setup "setup"
#define phase_work "work"
#define phase_commit "commit"
#define phase_terminate "terminate"
/* ============================================================================
* WAVE-CONTEXT REGISTERS — canonical register set for the tape wave model.
*
* The tape-atom runtime carries four registers across a wave:
*
* R_PrimCursor output pointer into the prim arena (next OT entry to write)
* R_FaceCursor input pointer into the face array (next face to consume)
* R_VertBase base pointer into the vertex arena (this wave's vertices)
* R_OtBase base pointer into the ordering table (this wave's OT slot)
*
* Each atom declares its reads/writes against this canonical set. The Lua
* tool rejects wave-context positions that reference any other register
* (warning today — the C compiler's R_T4..R_T7 / R_RA / etc. aliases are
* implementation details and not part of the typed surface).
*
* If your atom needs to touch GTE / SP / DMA / other side state, declare it
* at the source level as you normally would — but DO NOT put those registers
* in tape_regs(...). Wave-context is a closed set.
*
* ============================================================================*/
/* ============================================================================
* REGION TOKENS — memory regions atoms may allocate from or write into.
*
* Use atom_region(name, REGION) to declare. The Lua tool validates that the
* region is in this set, AND that:
* - rbind atoms declare the source region (usually HEAP_3D or CDROM_STREAM)
* - work atoms declare the destination region (the arena they push to)
* - commit atoms must declare a region equal to what setup wrote, so the
* C-side mirror is consistent
*
* Add new regions by extending this list and the metaprogram's KNOWN_REGIONS.
* Don't add regions ad-hoc — every new region becomes part of the contract.
*
* ============================================================================*/
#define REGION_PRIM_ARENA prim_arena /* OT/prim packet arena */
#define REGION_FACE_ARENA face_arena /* face index array */
#define REGION_VERTEX_ARENA vertex_arena /* vertex pool */
#define REGION_OT_ARENA ot_arena /* ordering-table array */
#define REGION_HEAP_3D heap_3d_models /* loaded model heap */
#define REGION_CDROM_STREAM cdrom_stream /* CDROM read buffer */
#define REGION_VRAM vram_heap /* VRAM texture/GPU buffer */
/* ============================================================================
* CADENCE TOKENS — how often the atom runs.
*
* frame runs every vsync (rendering, input poll)
* once runs exactly once per process lifetime (init, terminate)
* ondemand runs when triggered by event (CDROM load, async DMA complete)
*
* Used as a hint for the metaprogram to flag:
* - frame-cadence atoms that have side effects (they'll be hit many times,
* so avoid global state mutation unless it's idempotent)
* - once-cadence atoms inside "if (frame_count == 0)" guards (the guard
* is then provably one-shot, the metaprogram can lift initialization)
* - ondemand atoms that are missed by the wave scheduler (forces async
* and discards yield results without further processing)
*
* ============================================================================*/
#define CADENCE_FRAME frame
#define CADENCE_ONCE once
#define CADENCE_ONDEMAND ondemand
/* ============================================================================
* tape_regs(...) — wave-context register list
*
* tape_regs(R_PrimCursor, R_FaceCursor) → (R_PrimCursor, R_FaceCursor)
*
* The macro produces a comma-evaluated expression that the C compiler
* silently discards (it's wrapped in parentheses in the call argument
* position — the result is never bound). The Lua tool pattern-matches the
* "tape_regs(...)" token to extract the list.
*
* You can have at most one tape_regs(...) in the reads slot and one in the
* writes slot of atom_annot. To declare multiple disjoint sets (rare), just
* declare the union — the metaprogram doesn't track which reads need which
* writes at this granularity.
*
* ============================================================================*/
#define atom_reads(...) (__VA_ARGS__)
#define atom_writes(...) (__VA_ARGS__)
/* ============================================================================
* ATOM ANNOTATION MACROS
*
* Each expands to `__attribute__((annotate("kind"))) MipsAtom_(name)` —
* the GCC attribute is accepted under -Wno-attributes (already in your
* build flags) and stripped at runtime. The annotation string is just the
* macro kind ("atom_annot", "atom_bind", etc.) — the metaprogram reads
* the macro call's full args list from the source-as-written.
*
* ============================================================================*/
/* ----------------------------------------------------------------------------
* atom_init — entry into tape_runtime_main
*
* atom_init(tape_main)
* internal MipsAtom_(tape_main) { ... };
*
* Implies: no reads, no writes (wave-context not established yet).
* ----------------------------------------------------------------------------*/
#define atom_init(name) __attribute__((annotate("atom_init")))
/* ----------------------------------------------------------------------------
* atom_terminate — exit from tape_runtime_main
*
* atom_terminate(tape_exit)
* internal MipsAtom_(tape_exit) { ... };
*
* Implies: no reads, no writes (wave-context destroyed at this point).
* ----------------------------------------------------------------------------*/
#define atom_terminate(name) __attribute__((annotate("atom_terminate")))
/* ----------------------------------------------------------------------------
* atom_setup — pre-work atom: prepares engine state (e.g., set_gte_world)
*
* atom_setup(set_gte_world, tape_regs(R_TapePtr))
* internal MipsAtom_(set_gte_world) { ... };
*
* Reads: anything (the engine state you're reading)
* Writes: engine state (GTE / DMA / etc. — declared in source, not part of
* wave-context, so doesn't go in tape_regs)
*
* The metaprogram checks that setup is followed (in atomic order) by a work
* atom in the same wave — there's no point in setting up state if no one
* reads it.
* ----------------------------------------------------------------------------*/
#define atom_setup(name, reads) __attribute__((annotate("atom_setup")))
/* ----------------------------------------------------------------------------
* atom_commit — post-work atom: flushes wave-context back to C-side state
*
* atom_commit(sync_prim_cursor, tape_regs(R_PrimCursor))
* internal MipsAtom_(sync_prim_cursor) { ... };
*
* Reads: wave-context registers (the ones you sync back to C)
* Writes: C-side mirror (declared in source — not part of wave-context)
*
* The metaprogram checks that commit is preceded (in atomic order) by a
* work atom that wrote the registers this commit is reading.
* ----------------------------------------------------------------------------*/
#define atom_commit(name, reads) __attribute__((annotate("atom_commit")))
/* ----------------------------------------------------------------------------
* atom_bind — rbind atom: read wave-context registers from tape pointer
*
* atom_bind(rbind_cube_tri, Binds_CubeTri,
* tape_regs(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase))
* internal MipsAtom_(rbind_cube_tri) { ... };
*
* The binds_struct MUST be a typedef'd type (declared via
* `typedef struct Binds_X { ... } Binds_X;` somewhere in the source).
* The Lua tool cross-references this. Missing struct = error.
*
* Implicit: reads R_TapePtr, writes the four wave-context registers.
* ----------------------------------------------------------------------------*/
#define atom_bind(name, binds_struct, writes) __attribute__((annotate("atom_bind")))
/* ----------------------------------------------------------------------------
* atom_annot — generic work atom with explicit phase
*
* atom_annot(cube_tri, phase_work,
* tape_regs(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase),
* tape_regs(R_PrimCursor, R_FaceCursor))
* internal MipsAtom_(cube_tri) { ... };
*
* Use this for the bulk of your atoms. For init/setup/commit/bind, prefer
* the convenience macros above — they pin the phase for you.
*
* The phase arg is one of: phase_init / phase_bind / phase_setup /
* phase_work / phase_commit / phase_terminate. Spelling mistakes are errors.
* ----------------------------------------------------------------------------*/
#define atom_annot(name, phase, reads, writes) __attribute__((annotate("atom_annot")))
/* ============================================================================
* RESOURCE / GROUP / CADENCE / REGION / ASYNC — optional atom metadata
*
* These don't annotate the atom semantically (phase/reads/writes do that).
* They attach extra context that the metaprogram uses to catch:
* - same resource loaded twice in different ways
* - atoms that span multiple regions (likely bug — pick one)
* - frame-cadence atoms that should be once-cadence (perf / correctness)
* - ondemand atoms that aren't async (CDROM races)
*
* You can use as many as apply to a given atom, in any order, immediately
* above the atom_*() macro.
*
* ============================================================================*/
/* ----------------------------------------------------------------------------
* atom_resource — name the logical resource the atom references
*
* atom_resource(cube_tri, "model_ship_cube")
* atom_resource(load_track_faces, "track_lavender_field_0x42")
* atom_resource(play_engine_sfx, "sfx_engine_loop")
*
* Use any human-readable string. The metaprogram:
* - validates resource strings are non-empty and don't contain control chars
* - flags duplicates across atoms with the same name (two atoms claiming
* ownership of a resource is usually a refactor artifact or bug)
* - flags references to resources that no atom actually defines
*
* The arg is a STRING LITERAL, so it can't accidentally alias a variable.
* ----------------------------------------------------------------------------*/
#define atom_resource(name, res_id) //_Pragma("atom " #name " resource=" res_id)
/* ----------------------------------------------------------------------------
* atom_region — name the memory region the atom touches
*
* atom_region(cube_tri, REGION_PRIM_ARENA)
* atom_region(load_faces, REGION_HEAP_3D)
* atom_region(load_tex, REGION_VRAM)
*
* Use REGION_* tokens above. The metaprogram enforces the closed set.
*
* Edge cases the metaprogram catches:
* - rbind atom that doesn't declare a SOURCE region (where is it loading from?)
* - work atom with no destination region (where is it pushing to?)
* - region that disagrees with the Binds_* struct layout (you said it's a
* prim_arena rbind but the struct has 4 faces in it — wait, that's wrong)
* ----------------------------------------------------------------------------*/
#define atom_region(name, region) //_Pragma("atom " #name " region=" #region)
/* ----------------------------------------------------------------------------
* atom_group — bundle atoms into a logical batch (track-load, sound-load, etc.)
*
* atom_group(load_track_face_42, GROUP_LOAD_FACES)
* atom_group(load_track_face_43, GROUP_LOAD_FACES)
* atom_group(swap_face_42_43, GROUP_VISIBILITY_SWAP)
*
* Use any token as the group id. The metaprogram:
* - validates all atoms in a group emit their waves in the same tb_group
* (no spawning other waves inside a group)
* - flags groups with only one member (probably a typo — meant to be a group?)
* - validates cross-group edges (no atom reads what another group writes,
* unless explicitly grouped together)
*
* Useful when:
* - subdivisible work (track-face batches, polygon subdivision) needs to
* confirm that all batches of one logical visible scene are emitted
* together
* - async loads (CDROM -> VRAM) need to be grouped so all batches complete
* before the swap
*
* Use GROUPS for sound effects to track which sound plays during which atom,
* which is needed if the sound tool ever has to validate "this atom is the
* trigger for an audio play".
* ----------------------------------------------------------------------------*/
#define atom_group(name, group_id) //_Pragma("atom " #name " group=" #group_id)
/* ----------------------------------------------------------------------------
* atom_cadence — declare execution frequency
*
* atom_cadence(render_frame, CADENCE_FRAME) // every vsync
* atom_cadence(load_track_faces, CADENCE_ONDEMAND) // on demand
* atom_cadence(init_heap, CADENCE_ONCE) // process lifetime
*
* Default (no atom_cadence call) is CADENCE_FRAME — most atoms run every
* frame. Override explicitly when not.
*
* The metaprogram's checks:
* - CADENCE_ONCE atoms inside `if (frame == 0)` or `if (!initialized)` are
* tagged, validating that guards are required (or warning if missing)
* - CADENCE_FRAME atoms that mutate state outside the wave context get
* flagged (likely a bug — state should persist through commits)
* - CADENCE_ONDEMAND atoms must have atom_async — otherwise the trigger
* mechanism is undefined
* ----------------------------------------------------------------------------*/
#define atom_cadence(name, cadence) //_Pragma("atom " #name " cadence=" #cadence)
/* ----------------------------------------------------------------------------
* atom_async — declare whether the atom yields / interacts with CDROM DMA
*
* atom_async(load_track_tex, true) // CDROM read yield
* atom_async(load_vram, true) // VRAM upload DMA
* atom_async(render_frame, false) // pure compute, no async
*
* The metaprogram requires this for CADENCE_ONDEMAND atoms. For
* CADENCE_FRAME, it's optional but documents intent.
*
* Note: CDROM ATOMS in Psy-Q are typically implemented as a chain of
* "async-init" atom followed by a "wait-for-completion" atom. Both atoms
* should be marked async=true, and both should have the same resource/group
* tag (so the metaprogram can verify they're paired).
* ----------------------------------------------------------------------------*/
#define atom_async(name, is_async) //_Pragma("atom " #name " async=" #is_async)
/* ============================================================================
* WORD-COUNT ANNOTATION FOR A #define MAC
*
* tape_words(mac_yield, 1)
* #define mac_yield() \
* load_word(R_AtomJmp, R_TapePtr, 0), \
* add_ui_1(R_TapePtr, 4), \
* jump_reg(R_AtomJmp), \
* nop
*
* The compiler accepts the unknown _Pragma. The Lua tool reads it and
* cross-checks against WORD_COUNT(mac_yield, 1) in tape_atom.metadata.h.
* If they disagree, build fails.
*
* Use sparingly — only on multi-word macros (single-word ones don't need
* drift tracking; they're checked by the .word-count pass anyway).
*
* ============================================================================*/
#define tape_words(name, n) //_Pragma(#name " tape_atom words=" #n)
/* ============================================================================
* atom_label / atom_offset — branch target machinery
*
* atom_label(culling) ← nothing in C; anchor only
* ... body ...
* atom_label(bounds_chk) ← another anchor
*
* atom_offset(culling, bounds_chk) ← resolved by gen/.offsets.h
*
* The metaprogram generates gen/atom_offsets.h with one
* #define atom_offset__culling__bounds_chk ((target - branch_pos - 1))
* per atom_offset(F, T) call. The preprocessor then expands your call to
* the right immediate value.
*
* If gen/atom_offsets.h is stale (or atom_label(name) is undefined),
* `atom_offset__F__T` becomes an undefined macro and the C build fails.
* This catches:
* - typo in atom_label (no anchor → metaprogram doesn't emit the macro)
* - .offsets.h not regenerated after body edits
* - body edit that broke the offset math (recompile + retest picks it up
* in CPU emulator)
*
* ============================================================================*/
#define atom_offset(F, T) atom_offset_ ## F ## _ ## T
#define atom_label(name) /* anchor — see metaprogram documentation */
+159
View File
@@ -0,0 +1,159 @@
/*
* dsl.atom.h
* ============================================================================
*
* ATOM DSL: Annotation layer for tape atoms (lottes_tape.h).
* The metaprogram (scripts/passes/annotation.lua) reads source-as-written and validates:
* - atom_info(...) shape: up to three sub-calls (atom_bind(Binds_X), atom_reads(...), atom_writes(...)) in any order and are optional.
* - rbind atoms (atom_info(..., atom_bind(Binds_X), ...)) reference a real Binds_* struct declaration.
* - atom word-counts in word_counts.metadata.h match the body's actual .word count.
*
* Pure macro anntation.
* ---------------
* Don't want to constraint the macro usage to some attribute placment constraint, etc, don't want ot dela with the compiler.
* atom_info, atom_bind, atom_reads, atom_writes, atom_label, atom_dbg_skip each expand to a C comment or to nothing
* (C preprocessor strips them to whitespace).
*
* ============================================================================
* Usage:
* MipsAtom_(cube_tri) atom_info(
* atom_reads (R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase)
* , atom_writes(R_PrimCursor, R_FaceCursor)
* ){
* atom_label(culling),
* // ... atom body ...
* atom_offset(culling, bounds_chk) // branch target, validated
* // ... atom body ...
* atom_label(bounds_chk),
* };
*
*
* Data Binding pattern -- atom_bind as a sub-call of atom_info
*
* // Wave-context register layout (declarative):
* typedef Struct_(Binds_TrackFaceBatch) {
* U4 PrimCursor;
* U4 FaceCursor;
* U4 VertBase;
* U4 OtBase;
* };
* MipsAtom_(rbind_track_face_batch) atom_info(
* atom_bind(Binds_TrackFaceBatch)
* , atom_writes(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase)
* ){ ... };
*
* Annotation rules
* ----------------
* 1. atom_info(...) is OPTIONAL. Atoms without atom_info are silently skipped by the metaprogram.
* 2. If present, atom_info takes up to three sub-calls, all order-independent within the arg list:
* - atom_bind(Binds_X)
* - atom_reads(...)
* - atom_writes(...)
* 3. atom_bind(Binds_X): metaprogram cross-references Binds_X against the `typedef struct Binds_X { ... } Binds_X;` declaration.
* 4. atom_reads(...) and atom_writes(...): Used to to check if registers are used correctly in macros: R_PrimCursor / R_FaceCursor / R_VertBase / R_OtBase.
* 5. atom_label(name: Utilize with atom_offset as a target location.
* 6. atom_offset(F, T): Resolved by gen/atom_offsets.h, generated from the atom_label markers. Calculated during the offset pass of the lua metaprogram.
*/
#ifdef INTELLISENSE_DIRECTIVES
#pragma once
#endif
/* ============================================================================
* atom_reads(...) / atom_writes(...)
*
* Used during the static analysis pass of the metaprogram to do
* ============================================================================*/
#define atom_reads(...) (__VA_ARGS__)
#define atom_writes(...) (__VA_ARGS__)
/* ----------------------------------------------------------------------------
* atom_reg (per-enum opt-in marker for the DWARF register-alias registry)
*
* The bare `atom_reg` token adjacent to an enum entry in mips.h / lottes_tape.h flags that alias as debug-visible for scan_source's register_alias_registry.
* The C preprocessor strips it to a comment so no runtime symbol is created; the Lua scanner reads the bare token.
* ----------------------------------------------------------------------------*/
#define atom_reg /* atom_reg: opt the preceding enum entry into the DWARF registry */
/* ============================================================================
* atom_info :
* MipsAtom_(cube_tri) atom_info(
* atom_reads (R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase)
* , atom_writes(R_PrimCursor, R_FaceCursor)
* ){ ... };
*
* - atom_bind(Binds_X): metaprogram cross-references Binds_X against the `typedef struct Binds_X { ... } Binds_X;` declaration.
* - atom_reads(...): comma-list of registers
* - atom_writes(...): comma-list of registers
* ============================================================================*/
#define atom_info(...) /* atom_info(__VA_ARGS__) */
/* ----------------------------------------------------------------------------
* DEBUG SOURCE-STEP MARKER
*
* Place `atom_dbg_skip` (BARE) before a MipsAtom_, MipsAtomComp_, or MipsAtomComp_Proc_.
* The following declaration kind determines whether the marker selects a whole atom or a component inline view.
* The source scanner associates the marker with that declaration; placement diagnostics are handled by the annotation pass.
*
* Example:
* atom_dbg_skip MipsAtom_(tape_exit) { jump_reg(rret_addr), nop };
* atom_dbg_skip MipsAtomComp_(ac_yield) { ... };
* atom_dbg_skip MipsAtomComp_Proc_(ac_format_f3_color, { ... });
* ----------------------------------------------------------------------------*/
#define atom_dbg_skip /* atom_dbg_skip: skip the following atom or component source view */
/* ----------------------------------------------------------------------------
* Typed-view annotations (Registry for DWARF RR_<R_X> chain resolution)
* atom_type(<T>) -- overloaded:
* (a) enum-site default: `R_Foo = R_Tn, atom_reg atom_type(T)`
* Sets the per-alias default typed view in the register_alias_registry.
* Consumed by the DWARF chain step (e) when no per-atom atom_ctx / atom_phase / atom_type callsite provides a stronger resolution.
* (b) callsite override: `atom_reads(R_Foo atom_type(T), ...)` Overrides the per-alias default for THIS atom only.
* Last-write-wins per R_Name; conflict -> error.
* atom_ctx(<atom_name>) -- atom-info sub-call:
* Propagate another atom's atom.rbind.fields (its Binds_* typed fields) into THIS atom's typed-view resolution.
* The named atom must be an rbind atom (have `atom_bind(Binds_X)` in its `atom_info`).
* Used as the escape hatch when atom_phase is not the natural correlation.
* atom_phase(<label>) -- atom-info sub-call:
* Free-form C-identifier label for grouping atoms.
* Within a phase, the FIRST atom in source-order that owns its own atom.rbind provides
* the Binds_* field types used by all other atoms in the same phase.
* The preferred correlation mechanism; atom_ctx is the escape hatch for non-natural cases.
*
* All three expand to C comments
* (the bare-token convention matching `atom_reg` and `atom_dbg_skip`).
* The Lua scanner reads the bare tokens in source-as-written; the C preprocessor strips them.
* ----------------------------------------------------------------------------*/
#define atom_type(T) /* atom_type: associate <T> with the preceding enum entry (enum site) or this register (atom-info site) */
#define atom_ctx(atom_name) /* atom_ctx: propagate <atom_name>'s Binds_* field types into this atom's typed views */
#define atom_phase(label) /* atom_phase: tag this atom with <label> for grouped typed-view resolution */
/* ----------------------------------------------------------------------------
* atom_bind(Binds_X) -- rbind sub-call of atom_info
*
* MipsAtom_(rbind_cube_tri) atom_info(
* atom_bind(Binds_CubeTri)
* , atom_writes(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase)
* ){ ... };
*
* The Binds_X MUST be a typedef'd type (declared via `typedef struct Binds_X { ... } Binds_X;` somewhere in the source).
* ----------------------------------------------------------------------------*/
#define atom_bind(binds_struct) /* atom_bind(binds_struct) */
/* ============================================================================
* atom_label / atom_offset — branch target machinery
*
* atom_label(culling) ← nothing in C; anchor only
* ... body ...
* atom_label(bounds_chk) ← another anchor
*
* atom_offset(culling, bounds_chk) ← resolved by gen/offsets.h
*
* The metaprogram generates gen/offsets.h with one #define with the offset value per atom_offset(F, T) call.
* The preprocessor then expands the call to the right immediate value.
*
* If gen/offsets.h is stale (or atom_label(name) is undefined), `atom_offset_F_T` becomes an undefined macro and the C build fails.
* ============================================================================*/
#define atom_offset(F, T) atom_offset_ ## F ## _ ## T
// atom_label is a pure annotation for the metaprogram's offset calculations.
#define atom_label(name) /* atom_label anchor: name */
+23 -16
View File
@@ -43,7 +43,9 @@
#define R_ restrict
#define V_ volatile
// Fictional, used for intiution.
#pragma region Fictional //, used for intiution
#define EUB_ restrict // Execute Unit Bound: Data is siloed in the ALU Register File. The Load/Store Unit is bypassed. (Route to Execution Unit. Keep in registers)
#define ISO_ restrict // Isolated Provenance: Alternative to Exu_. Guarantees electrical memory isolation,
// unlocking the compilers ability to safely pack data across multiple parallel SIMD lanes (vectorization).
@@ -67,7 +69,8 @@
#define latch_load_anchor(ptr) //__atomic_load_n(ptr, ooo_anchor_)
#define latch_store_drain(ptr, val) //__atomic_store_n(ptr, val, ooo_drain_)
#define pulse_xchg_weld(ptr, val) //__atomic_exchange_n(ptr, val, ooo_weld_)
//end of: Fictional.
#pragma endreigon Fictional
// R_ (restrict) establishes an "Eigen" or "Proprius" mapping.
@@ -99,10 +102,10 @@
#define Struct_(symbol) struct symbol TSet_(symbol); struct symbol
#define Union_(symbol) union symbol TSet_(symbol); union symbol
#define Opt_(proc) Struct_(tmpl(Opt,proc))
#define opt_(symbol, ...) (tmpl(Opt,symbol)){__VA_ARGS__}
#define Ret_(proc) Struct_(tmpl(Ret,proc))
#define ret_(proc) tmpl(Ret,proc) proc
#define Opt_(proc) Struct_(tmpl(Opt,proc))
#define opt_(symbol, ...) (tmpl(Opt,symbol)){__VA_ARGS__}
#define Ret_(proc) Struct_(tmpl(Ret,proc))
#define ret_(proc) tmpl(Ret,proc) proc
// Using Byte-Width convention for the fundamental types.
typedef __UINT8_TYPE__ TSet_(U1);
@@ -135,16 +138,17 @@ enum { false = 0, true = 1, true_overflow, };
typedef void Proc_(VoidFn) (void);
#define kilo(n) (C_(U4, n) << 10)
#define mega(n) (C_(U4, n) << 20)
#define giga(n) (C_(U4, n) << 30)
#define tera(n) (C_(U4, n) << 40)
#define null C_(U4, 0)
#define nullptr C_(void*, 0)
#define O_(type, field) (C_(U4, & C_(type*,0)->field))
#define OT_(field) O_(typeof_ptr(& field), filed))
#define S_(data) C_(U4, sizeof(data))
#define kilo(n) (C_(U4, n) << 10)
#define mega(n) (C_(U4, n) << 20)
#define giga(n) (C_(U4, n) << 30)
#define tera(n) (C_(U4, n) << 40)
#define null C_(U4, 0)
#define nullptr C_(void*, 0)
#define O_(type, field) C_(U4, & C_(type*,0)->field)
#define OA_(type, member, idx) C_(U4, & C_(type*,0)->member[idx])
#define OT_(field) O_(typeof_ptr(& field), filed))
#define S_(data) C_(U4, sizeof(data))
#define sop_1(op,a,b) C_(U1, s1_(a) op s1_(b))
#define sop_2(op,a,b) C_(U2, s2_(a) op s2_(b))
@@ -218,3 +222,6 @@ IA_ void assert(U8 cond) { if(cond){return;} else{debug_trap(); ms_exit_process(
#endif
#pragma endregion Debug
#endif
#define GCC_OPTIMIZATION_DISABLE _Pragma("GCC push_options") _Pragma("GCC optimize(\"O0\")")
#define GCC_OPTIMIZATION_ENABLE _Pragma("GCC pop_options")
+15 -23
View File
@@ -50,17 +50,13 @@
#define asm_words(...) m_expand(glue(GCC_ASM_INL_, GCC_ASM_COUNT_ARGS(__VA_ARGS__))(__VA_ARGS__))
// Very nasty macro expansion. See the Cruft pragma region after all the DSL defines
/* reg_str(n) — Stringify an integer register id into the GCC asm
* string form (e.g. 12 → "$12"). Use this anywhere GCC's parser
* expects a literal string identifying a register: clobber lists,
* asm templates, etc. The two-level macro is the standard preprocessor
* idiom for forcing one level of expansion before stringify — without
* it, `#n` would stringify the macro name `R_T4` to `"R_T4"` instead
* of expanding `R_T4` to its value first.
/* reg_str(n) — Stringify an integer register id into the GCC asm string form (e.g. 12 → "$12").
* Use this anywhere GCC's parser expects a literal string identifying a register: clobber lists,
* asm templates, etc. The two-level macro is the standard preprocessor idiom for forcing one level of expansion before stringify —
* without it, `#n` would stringify the macro name `R_T4` to `"R_T4"` instead of expanding `R_T4` to its value first.
*
* For declaring a register variable bound to a specific GPR, use the
* `rgcc(n)` bundle from gcc_asm.h instead — it adds the `__asm__()`
* qualifier around the string.
* For declaring a register variable bound to a specific GPR, use the `rgcc(n)` bundle from gcc_asm.h instead —
* it adds the `__asm__()` qualifier around the string.
*
* register V3_S2* p0 __asm__(reg_str(R_T4)) = ...; // verbose
* register V3_S2* p0 rgcc(R_T4) = ...; // bundled
@@ -85,21 +81,19 @@
* - The string "$12" is derived from it via reg_str, so they cannot drift apart.
* - Spelling `__asm__(reg_str(R_T4_Code))` at every call site is noise.
*
* tmpl defined in dsl.h (the token-paste glue).
* tmpl defined in dsl.h (token-paste glue).
* rgcc define here (gcc_asm.h) because the `__asm__` keyword is GCC-specific.
* Anyone porting to a different compiler's asm dialect overrides rgcc,
* Anyone porting to a different compiler's asm dialect overrides rgcc,
* and the integer→string derivation in rlit can be retargeted in one place.
*
* For clobber lists and asm-template strings, use the bare `rlit(R_T4_Code)`.
* ------------------------------------------------------------------------ */
#define rgcc(n) __asm__(rlit(n))
/* rgcc_ref(n) — GCC operand-reference form "%N". Not currently used
* by the placeholder-pun macros (the .word bodies are fully baked
* at compile time and have no runtime operand references), but kept
* here for completeness in case a future asm template needs to refer
* to a runtime input by position. Mirror of rgcc but produces "%N"
* instead of "$N". */
/* rgcc_ref(n) — GCC operand-reference form "%N". Not currently used by the placeholder-pun macros
* (the .word bodies are fully baked at compile time and have no runtime operand references),
* but kept here for completeness in case a future asm template needs to refer to a runtime input by position.
* Mirror of rgcc but produces "%N" instead of "$N". */
#define rgcc_ref_(n) "%" #n
#define rgcc_ref(n) rgcc_ref_(n)
@@ -147,11 +141,9 @@
9, 8, 7, 6, 5, 4, 3, 2, 1, 0))
/* --- 2. String Concatenation Helpers --- *
* NOTE: we use `%0`, `%1`, ... not `%c0`, `%c1`, ... because GCC's
* asm-parser rejects `%cN` in this position with "invalid use of '%c'".
* The `%cN` form is for printing *character* constants; for arbitrary
* integer immediates (the only kind `"i"(...)` produces), the plain
* `%N` form is the right one. Both expand to the bare immediate.
* NOTE: we use `%0`, `%1`, ... not `%c0`, `%c1`, ... because GCC's asm-parser rejects `%cN` in this position with "invalid use of '%c'".
* The `%cN` form is for printing *character* constants; for arbitrary integer immediates (the only kind `"i"(...)` produces),
* the plain `%N` form is the right one. Both expand to the bare immediate.
*/
#define GCC_ASM_W1 "%0"
#define GCC_ASM_W2 GCC_ASM_W1 ", %1"
-9
View File
@@ -1,9 +0,0 @@
// Auto-generated by tape_atom_offset_gen.meta.lua — DO NOT EDIT
// Source: C:\projects\Pikuma\ps1\code\duffle\lottes_tape.h
#pragma once
#pragma region lottes_tape
#pragma endregion lottes_tape
+184
View File
@@ -0,0 +1,184 @@
#ifdef INTELLISENSE_DIRECTIVES
#pragma once
#endif
// Auto-generated by ps1_meta.lua — DO NOT EDIT
// Directory: C:\projects\Pikuma\ps1\code\duffle/
// source: C:\projects\Pikuma\ps1\code\duffle\word_count.metadata.h
// source: C:\projects\Pikuma\ps1\code\duffle\dsl.h
// source: C:\projects\Pikuma\ps1\code\duffle\memory.h
// source: C:\projects\Pikuma\ps1\code\duffle\math.h
// source: C:\projects\Pikuma\ps1\code\duffle\gcc_asm.h
// source: C:\projects\Pikuma\ps1\code\duffle\mips.h
// source: C:\projects\Pikuma\ps1\code\duffle\gp.h
// source: C:\projects\Pikuma\ps1\code\duffle\gte.h
// source: C:\projects\Pikuma\ps1\code\duffle\pad.h
// source: C:\projects\Pikuma\ps1\code\duffle\dsl.atom.h
// source: C:\projects\Pikuma\ps1\code\duffle\lottes_tape.h
// source: C:\projects\Pikuma\ps1\code\duffle\psyq.h
// source: C:\projects\Pikuma\ps1\code\duffle\math.atom.c
// source: C:\projects\Pikuma\ps1\code\duffle\mips.atom.c
// source: C:\projects\Pikuma\ps1\code\duffle\gte.atom.c
// source: C:\projects\Pikuma\ps1\code\duffle\gp.atom.c
// source: C:\projects\Pikuma\ps1\code\duffle\pad.atom.c
// source: C:\projects\Pikuma\ps1\code\duffle\psyq.atom.c
// Component atoms (MipsAtomComp_(ac_*)) -> macro variants (mac_*)
#ifndef WORD_COUNT
#define WORD_COUNT(name, count) enum { words_##name = (count) };
#endif
/* atom_dbg_skip */
/* ---------------------------------------------------------------------------
* MACRO ATOM Components (Reusable Assembly Components)
* These do NOT yield. They are expanded inline inside Tape Atoms.
* ---------------------------------------------------------------------------*/
// The 'Yield' sequence for Tape Atoms (mac_yield).
// - mac_yield() is the safe default for atom-endings: 4 words, BD-slot of jr is mandatory nop.
// - mac_yield_load() + mac_yield_tail():
// - unconditional branch: mac_yield_load fills the branch's BD-slot (replaces a nop);
// - mac_yield_tail runs at the branch target (does NOT re-load R_AtomJmp).
#define mac_yield(...) \
load_word(R_AtomJmp, R_TapePtr, 0) \
, add_ui_self( R_TapePtr, S_(MipsCode)) \
, jump_reg( R_AtomJmp) \
, nop
WORD_COUNT(mac_yield, 4)
/* atom_dbg_skip */
#define mac_yield_load(...) \
load_word(R_AtomJmp, R_TapePtr, 0)
WORD_COUNT(mac_yield_load, 1)
/* atom_dbg_skip */
#define mac_yield_tail(...) \
add_ui_self(R_TapePtr, S_(MipsCode)) \
, jump_reg( R_AtomJmp) \
, nop
WORD_COUNT(mac_yield_tail, 3)
/* atom_dbg_skip */
#define mac_load_v2s2(rs_x, rs_y, r_base, offset) \
load_half( rs_x, r_base, O_(V3_S2,x)) \
, load_half( rs_y, r_base, O_(V3_S2,y))
WORD_COUNT(mac_load_v2s2, 2)
/* atom_dbg_skip */
#define mac_store_v2s2(rt_x, rt_y, base, offset) \
store_half(rt_x, base, offset + O_(V2_S2,x)) \
, store_half(rt_y, base, offset + O_(V2_S2,y))
WORD_COUNT(mac_store_v2s2, 2)
/* atom_dbg_skip */
#define mac_store_rects2(rt_x, rt_y, rt_width, rt_height, base, offset) \
store_half(rt_x, base, offset + O_(Rect_S2,x)) \
, store_half(rt_y, base, offset + O_(Rect_S2,y)) \
, store_half(rt_width, base, offset + O_(Rect_S2,width)) \
, store_half(rt_height, base, offset + O_(Rect_S2,height))
WORD_COUNT(mac_store_rects2, 4)
/* atom_dbg_skip */
#define mac_load_tri_indices(r_face_cusor, r_i0, r_i1, r_i2) \
load_half_u(r_i0, r_face_cusor, 0 * S_(S2)) \
, load_half_u(r_i1, r_face_cusor, 1 * S_(S2)) \
, load_half_u(r_i2, r_face_cusor, 2 * S_(S2))
WORD_COUNT(mac_load_tri_indices, 3)
/* atom_dbg_skip */
#define mac_gte_store_f3(r_primitive_cursor) \
gte_sw(C2_SXY0, r_primitive_cursor, O_(Poly_F3,p0)) \
, gte_sw(C2_SXY1, r_primitive_cursor, O_(Poly_F3,p1)) \
, gte_sw(C2_SXY2, r_primitive_cursor, O_(Poly_F3,p2))
WORD_COUNT(mac_gte_store_f3, 3)
/* atom_dbg_skip */
#define mac_gte_load_tri_verts(r_vert_base, r_v0, r_v1, r_v2) \
shift_lleft(R_AT, r_v0, v3s2_byteoff) \
, add_u_self(R_AT, r_vert_base) \
, load_word(R_V0, R_AT, O_(V3_S2,x)) \
, load_word(R_V1, R_AT, O_(V3_S2,z)) \
, gte_mv_to_data_r(R_V0, C2_VXY0) \
, gte_mv_to_data_r(R_V1, C2_VZ0) \
, shift_lleft(R_AT, r_v1, v3s2_byteoff) \
, add_u_self(R_AT, r_vert_base) \
, load_word(R_V0, R_AT, O_(V3_S2,x)) \
, load_word(R_V1, R_AT, O_(V3_S2,z)) \
, gte_mv_to_data_r(R_V0, C2_VXY1) \
, gte_mv_to_data_r(R_V1, C2_VZ1) \
, shift_lleft(R_AT, r_v2, v3s2_byteoff) \
, add_u_self(R_AT, r_vert_base) \
, load_word(R_V0, R_AT, O_(V3_S2,x)) \
, load_word(R_V1, R_AT, O_(V3_S2,z)) \
, gte_mv_to_data_r(R_V0, C2_VXY2) \
, gte_mv_to_data_r(R_V1, C2_VZ2)
WORD_COUNT(mac_gte_load_tri_verts, 18)
/* atom_dbg_skip */
#define mac_gte_store_g4_p012(r_primitive_cursor) \
gte_sw(C2_SXY0, r_primitive_cursor, O_(Poly_G4,p0)) \
, gte_sw(C2_SXY1, r_primitive_cursor, O_(Poly_G4,p1)) \
, gte_sw(C2_SXY2, r_primitive_cursor, O_(Poly_G4,p2))
WORD_COUNT(mac_gte_store_g4_p012, 3)
/* atom_dbg_skip */
#define mac_gte_store_g4_p3(r_primitive_cursor) \
gte_sw(C2_SXY2, r_primitive_cursor, O_(Poly_G4,p3))
WORD_COUNT(mac_gte_store_g4_p3, 1)
#define mac_gcmd_push(cmd, reg_transfer, reg_base, port) \
load_upper_i(reg_transfer, cmd >> 16) \
, or_i_self( reg_transfer, cmd & 0xFFFF) \
, store_word( reg_transfer, reg_base, port)
WORD_COUNT(mac_gcmd_push, 3)
/* atom_dbg_skip */
#define mac_store_rgb8(rr, rg, rb, base, offset) \
store_byte(rr, base, offset + O_(RGB8,r)) \
, store_byte(rg, base, offset + O_(RGB8,g)) \
, store_byte(rb, base, offset + O_(RGB8,b))
WORD_COUNT(mac_store_rgb8, 3)
/* atom_dbg_skip */
#define mac_pack_color_word(r_base, off, cmd, r, g, b) \
load_upper_i(R_AT, (cmd) << 8 | (b)) \
, or_i_self( R_AT, ((g) << 8) | (r)) \
, store_word( R_AT, r_base, (off))
WORD_COUNT(mac_pack_color_word, 3)
/* atom_dbg_skip */
#define mac_format_f3_color(r_base, r, g, b) \
mac_pack_color_word(r_base, O_(Poly_F3,color), gp0_cmd_poly_f3, r, g, b)
WORD_COUNT(mac_format_f3_color, 3)
#define mac_format_g4_color(r_prim_cursor, r0, g0, b0, r1, g1, b1, r2, g2, b2, r3, g3, b3) \
mac_pack_color_word(r_prim_cursor, O_(Poly_G4,c0), gp0_cmd_poly_g4, r0,g0,b0) \
, mac_pack_color_word(r_prim_cursor, O_(Poly_G4,c1), 0, r1,g1,b1) \
, mac_pack_color_word(r_prim_cursor, O_(Poly_G4,c2), 0, r2,g2,b2) \
, mac_pack_color_word(r_prim_cursor, O_(Poly_G4,c3), 0, r3,g3,b3)
WORD_COUNT(mac_format_g4_color, 12)
#define mac_insert_ot_tag_f3(r_ot_base, r_prim_cursor) \
shift_lleft( R_T1, R_T1, S_(U4)/2) /* T1 = otz * S_(U4) (otz arg is implicit R_T1) */ \
, add_u_self( R_T1, r_ot_base) /* T1 = & OrderingTable[OTZ] */ \
, load_word( R_AT, R_T1, O_(PolyTag,code)) /* AT = old_ot_head */ \
, load_upper_i(R_V0, (S_(Poly_F3)/S_(U4) - S_(PolyTag)/S_(U4)) << PolyTag_len_bits) /* V0 = (5 - 1) << 24 = 4 << 24 */ \
, mask_upper( R_AT, R_AT, S_(PolyTag_len_bits)) /* Strip upper 8 bits (length from prev cell) → keep only low 24 */ \
, or_u( R_AT, R_AT, R_V0) /* Merge length */ \
, store_word( R_AT, r_prim_cursor, O_(PolyTag,code)) /* prim->tag = packed(prim_length, old_addr) */ \
, shift_lleft( R_AT, r_prim_cursor, S_(PolyTag_len_bits)) /* AT = (prim_length << 24) | old_addr */ \
, shift_lright(R_AT, R_AT, S_(PolyTag_len_bits)) \
, store_word( R_AT, R_T1, O_(PolyTag,code)) /* OrderingTable[OTZ] = PrimCursor */
WORD_COUNT(mac_insert_ot_tag_f3, 11)
#define mac_insert_ot_tag_g4(r_ot_base, r_prim_cursor) \
shift_lleft( R_T1, R_T1, S_(U4)/2) /* T1 = otz * S_(U4) (otz arg is implicit R_T1) */ \
, add_u_self( R_T1, r_ot_base) /* T1 = & OrderingTable[OTZ] */ \
, load_word( R_AT, R_T1, O_(PolyTag,code)) /* AT = old_ot_head */ \
, load_upper_i(R_V0, (S_(Poly_G4)/S_(U4) - S_(PolyTag)/S_(U4)) << PolyTag_len_bits) /* V0 = (9 - 1) << 24 = 8 << 24 */ \
, mask_upper( R_AT, R_AT, S_(PolyTag_len_bits)) /* Strip upper 8 bits (length from prev cell) → keep only low 24 */ \
, or_u( R_AT, R_AT, R_V0) /* Merge length */ \
, store_word( R_AT, r_prim_cursor, O_(PolyTag,code)) /* prim->tag = packed(prim_length, old_addr) */ \
, shift_lleft( R_AT, r_prim_cursor, S_(PolyTag_len_bits)) /* AT = (prim_length << 24) | old_addr */ \
, shift_lright(R_AT, R_AT, S_(PolyTag_len_bits)) \
, store_word( R_AT, R_T1, O_(PolyTag,code)) /* OrderingTable[OTZ] = PrimCursor */
WORD_COUNT(mac_insert_ot_tag_g4, 11)
+53
View File
@@ -0,0 +1,53 @@
// Auto-generated by ps1_meta.lua (passes/offsets.lua) — DO NOT EDIT
// Directory: C:\projects\Pikuma\ps1\code\duffle\
// source: C:\projects\Pikuma\ps1\code\duffle\word_count.metadata.h
// source: C:\projects\Pikuma\ps1\code\duffle\dsl.h
// source: C:\projects\Pikuma\ps1\code\duffle\memory.h
// source: C:\projects\Pikuma\ps1\code\duffle\math.h
// source: C:\projects\Pikuma\ps1\code\duffle\gcc_asm.h
// source: C:\projects\Pikuma\ps1\code\duffle\mips.h
// source: C:\projects\Pikuma\ps1\code\duffle\gp.h
// source: C:\projects\Pikuma\ps1\code\duffle\gte.h
// source: C:\projects\Pikuma\ps1\code\duffle\pad.h
// source: C:\projects\Pikuma\ps1\code\duffle\dsl.atom.h
// source: C:\projects\Pikuma\ps1\code\duffle\lottes_tape.h
// source: C:\projects\Pikuma\ps1\code\duffle\psyq.h
// source: C:\projects\Pikuma\ps1\code\duffle\math.atom.c
// source: C:\projects\Pikuma\ps1\code\duffle\mips.atom.c
// source: C:\projects\Pikuma\ps1\code\duffle\gte.atom.c
// source: C:\projects\Pikuma\ps1\code\duffle\gp.atom.c
// source: C:\projects\Pikuma\ps1\code\duffle\pad.atom.c
// source: C:\projects\Pikuma\ps1\code\duffle\psyq.atom.c
#pragma once
#pragma region duffle
// --- atom: pad_bios_snapshot (78 words) ---
#define _atom_offset_snap_root_skip_disconnected 8
#define _atom_offset_disconnected_snap_end 61
#define _atom_offset_case_2_id_dispatch 8
#define _atom_offset_pending_snap_end 51
#define _atom_offset_id_dispatch_try_analog_stick 11
#define _atom_offset_id_dispatch_snap_end 38
#define _atom_offset_try_analog_stick_try_analog_pad 12
#define _atom_offset_analog_stick_snap_end 24
#define _atom_offset_try_analog_pad_try_unsupported 11
#define _atom_offset_analog_pad_snap_end 10
enum {
atom_offset_snap_root_skip_disconnected = _atom_offset_snap_root_skip_disconnected,
atom_offset_disconnected_snap_end = _atom_offset_disconnected_snap_end,
atom_offset_case_2_id_dispatch = _atom_offset_case_2_id_dispatch,
atom_offset_pending_snap_end = _atom_offset_pending_snap_end,
atom_offset_id_dispatch_try_analog_stick = _atom_offset_id_dispatch_try_analog_stick,
atom_offset_id_dispatch_snap_end = _atom_offset_id_dispatch_snap_end,
atom_offset_try_analog_stick_try_analog_pad = _atom_offset_try_analog_stick_try_analog_pad,
atom_offset_analog_stick_snap_end = _atom_offset_analog_stick_snap_end,
atom_offset_try_analog_pad_try_unsupported = _atom_offset_try_analog_pad_try_unsupported,
atom_offset_analog_pad_snap_end = _atom_offset_analog_pad_snap_end,
};
#pragma endregion duffle
+82
View File
@@ -0,0 +1,82 @@
#ifdef INTELLISENSE_DIRECTIVES
# include "dsl.h"
# include "gp.h"
# include "lottes_tape.h"
#endif
ATOM_FILE_DEBUGGER_LINE_MARKER(gp_atom_c);
#pragma region MACs (Mips Atom Components)
FI_ Slice_MipsCode ac_gcmd_push(U4 cmd, U4 reg_transfer, U4 reg_base, U2 port)
MipsAtomComp_Proc_(ac_gcmd_push, {
load_upper_i(reg_transfer, cmd >> 16),
or_i_self( reg_transfer, cmd & 0xFFFF),
store_word( reg_transfer, reg_base, port),
})
FI_ Slice_MipsCode ac_store_rgb8(U1 rr, U1 rg, U1 rb, U4 base, U4 offset) atom_dbg_skip MipsAtomComp_Proc_(ac_store_rgb8, {
store_byte(rr, base, offset + O_(RGB8,r)),
store_byte(rg, base, offset + O_(RGB8,g)),
store_byte(rb, base, offset + O_(RGB8,b)),
})
/* Words: 3; Emits one (cmd|color) word to R_PrimCursor at the given
* byte offset. Internal helper used by the *_format_*_color macros. */
FI_ Slice_MipsCode ac_pack_color_word(U4 r_base, U4 off, U4 cmd, U1 r, U1 g, U1 b)
atom_dbg_skip MipsAtomComp_Proc_(ac_pack_color_word, {
load_upper_i(R_AT, (cmd) << 8 | (b)),
or_i_self( R_AT, ((g) << 8) | (r)),
store_word( R_AT, r_base, (off)),
})
/* Words: 3; Emits the F3 command+color word (cmd byte | BLUE | GREEN | RED)
* Args: _r, _g, _b are 8-bit RGB byte values (not raw 16-bit fields). */
FI_ Slice_MipsCode ac_format_f3_color(U4 r_base, U1 r, U1 g, U1 b)
atom_dbg_skip MipsAtomComp_Proc_(ac_format_f3_color, { mac_pack_color_word(r_base, O_(Poly_F3,color), gp0_cmd_poly_f3, r, g, b) })
/* Words: 12; Emits the four (code|color) words of a Poly_G4.
* Args: rN,gN,bN are 8-bit RGB byte values for each of the 4 vertices. */
FI_ Slice_MipsCode ac_format_g4_color(U4 r_prim_cursor,
U1 r0, U1 g0, U1 b0,
U1 r1, U1 g1, U1 b1,
U1 r2, U1 g2, U1 b2,
U1 r3, U1 g3, U1 b3)
MipsAtomComp_Proc_(ac_format_g4_color, {
mac_pack_color_word(r_prim_cursor, O_(Poly_G4,c0), gp0_cmd_poly_g4, r0,g0,b0),
mac_pack_color_word(r_prim_cursor, O_(Poly_G4,c1), 0, r1,g1,b1),
mac_pack_color_word(r_prim_cursor, O_(Poly_G4,c2), 0, r2,g2,b2),
mac_pack_color_word(r_prim_cursor, O_(Poly_G4,c3), 0, r3,g3,b3),
})
/* Words: 11; Correctly inserts a primitive into the Ordering Table linked list.
* Hardcoded for Poly_F3 (5 words). For Poly_G4, use ac_insert_ot_tag_g4. */
I_ Slice_MipsCode ac_insert_ot_tag_f3(U4 r_ot_base, U4 r_prim_cursor) MipsAtomComp_Proc_(ac_insert_ot_tag_f3, {
shift_lleft( R_T1, R_T1, S_(U4)/2), // T1 = otz * S_(U4) (otz arg is implicit R_T1)
add_u_self( R_T1, r_ot_base), // T1 = & OrderingTable[OTZ]
load_word( R_AT, R_T1, O_(PolyTag,code)), // AT = old_ot_head
load_upper_i(R_V0, (S_(Poly_F3)/S_(U4) - S_(PolyTag)/S_(U4)) << PolyTag_len_bits), // V0 = (5 - 1) << 24 = 4 << 24
mask_upper( R_AT, R_AT, S_(PolyTag_len_bits)), // Strip upper 8 bits (length from prev cell) → keep only low 24
or_u( R_AT, R_AT, R_V0), // Merge length
store_word( R_AT, r_prim_cursor, O_(PolyTag,code)), // prim->tag = packed(prim_length, old_addr)
shift_lleft( R_AT, r_prim_cursor, S_(PolyTag_len_bits)), // AT = (prim_length << 24) | old_addr
shift_lright(R_AT, R_AT, S_(PolyTag_len_bits)),
store_word( R_AT, R_T1, O_(PolyTag,code)), // OrderingTable[OTZ] = PrimCursor
})
/* Words: 11; Correctly inserts a primitive into the Ordering Table linked list.
* Hardcoded for Poly_G4 (9 words). For Poly_F3, use ac_insert_ot_tag_f3. */
I_ Slice_MipsCode ac_insert_ot_tag_g4(U4 r_ot_base, U4 r_prim_cursor) MipsAtomComp_Proc_(ac_insert_ot_tag_g4, {
shift_lleft( R_T1, R_T1, S_(U4)/2), // T1 = otz * S_(U4) (otz arg is implicit R_T1)
add_u_self( R_T1, r_ot_base), // T1 = & OrderingTable[OTZ]
load_word( R_AT, R_T1, O_(PolyTag,code)), // AT = old_ot_head
load_upper_i(R_V0, (S_(Poly_G4)/S_(U4) - S_(PolyTag)/S_(U4)) << PolyTag_len_bits), // V0 = (9 - 1) << 24 = 8 << 24
mask_upper( R_AT, R_AT, S_(PolyTag_len_bits)), // Strip upper 8 bits (length from prev cell) → keep only low 24
or_u( R_AT, R_AT, R_V0), // Merge length
store_word( R_AT, r_prim_cursor, O_(PolyTag,code)), // prim->tag = packed(prim_length, old_addr)
shift_lleft( R_AT, r_prim_cursor, S_(PolyTag_len_bits)), // AT = (prim_length << 24) | old_addr
shift_lright(R_AT, R_AT, S_(PolyTag_len_bits)),
store_word( R_AT, R_T1, O_(PolyTag,code)), // OrderingTable[OTZ] = PrimCursor
})
#pragma endregion MACs (Mips Atom Components)
+690 -98
View File
@@ -1,122 +1,714 @@
/* ============================================================================
* duffle DSL Suffix Conventions
* ============================================================================
* Every mnemonic in this header follows the same suffix grammar:
*
* Primitive commands: gp0_cmd_poly_f3 = 0x20 (byte opcode)
* Packed 32-bit cmd: gp0_word_poly_f3(r, g, b) (32-bit, shifted)
*
* Type ordering: domain?_(direction)?_action_target_modifier_type?
* Examples: add_ui (add + unsigned + immediate)
* add_s (add + signed, R-type implicit)
* shift_lleft (shift + logical + left)
* shift_aright (shift + arithmetic + right)
* call_reg(rs) (call + register, $ra implicit)
* gte_mv_to_data_r (gte + mv + to + data + register)
* gte_lw_v0_xy(base) (gte + lw + v0 + xy)
* load_upper_i (load-upper + immediate, unique verb)
*
* --- GPU-domain layer cake ---
* Every gp.h macro follows the same 4-layer composition as mips.h and gte.h:
* 4. Semantic encoders gp0_word_poly_f3(r,g,b)
* 3. Composite encoders enc_color_word(cmd, r, g, b)
* 2. Per-field encoders enc_gp0_color_r(r), enc_gp0_color_g(g), ...
* 1. Bitfield layout consts gp0_color_red_shift = 0, gp0_color_red_mask = 0xFF
* 0. Opcode IDs gp0_cmd_poly_f3 = 0x20
*
* Vendor mnemonics (gte_mtc2, gte_mfc2, etc.) are NOT in this header.
* They live in the opt-in `gp_vendor_sym.h` for users who prefer the PSYQ-style names.
* ============================================================================ */
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "dsl.h"
# include "math.h"
# include "mips.h"
#endif
typedef Enum_(U4, gp_Commands) {
gcmd_Reset = 0b000,
gcmd_Polygon = 0b001,
gcmd_Line = 0b010,
gcmd_Rect = 0b011,
gcmd_VM_to_VM = 0b100,
gcmd_CPU_to_VM = 0b101,
gcmd_VM_to_CPU = 0b110,
gcmd_Environment = 0b111,
gcmd_SetDrawMode = 0xE1,
gcmd_SetTextureWindow = 0xE2,
gcmd_SetDrawArea_TopLeft = 0xE3,
gcmd_SetDrawArea_BotRight = 0xE4,
gcmd_SetDrawOffset = 0xE5,
gcmd_SetMaskBit = 0xE6,
gcmd_ResetCommandBuffer = 0x01,
gcmd_AcknowledgeGPUInterrupt = 0x02,
gcmd_DisplayEnable = 0x03,
gcmd_DMA_Request = 0x04,
gcmd_DispArea_Start = 0x05,
gcmd_HorizontalDisplayRange = 0x06,
gcmd_VerticalDisplayRange = 0x07,
gcmd_DisplayMode = 0x08,
gcmd_SetVramSize = 0x09,
};
#pragma region GPU Ports & Commands
/* ============================================================================
* Hardware MMIO Addresses
* ============================================================================
* PSX GPU has two 32-bit ports in the I/O register region at KSEG2 0x1F800000+.
* GP0 (offset 0x10) is the data port (commands + params).
* GP1 (offset 0x14) is the control port (status, ctrl writes).
* ============================================================================ */
/* IO base address (KSEG2 0x1F800000+ for the I/O register region).
* The 16-bit upper half `IO_BASE_ADDR_HI16` is the form used by tape-side macros that pin a register
* to hold the IO base and access ports via offsets:
* `lui $reg, 0x1F80` (1 word) then `sw $data, GPIO_PORT*_OFFSET($reg)` (1 word).
* Mirrors the `IO_BASE_ADDR equ 0x1F80` + `gpio_port0 equ 0x1810` pattern from graphics_hello/gp.s. */
enum {
gpio_port_0 = 0x1810,
gpio_port_1 = 0x1814,
IO_BASE_ADDR = 0x1F800000, /* full 32-bit I/O region base */
IO_BASE_ADDR_HI16 = 0x1F80, /* fits in a single `lui $reg, 0x1F80` */
gcmd_offset = 24,
/* Offsets from IO_BASE_ADDR to each port. Used by tape-side macros
* that pin a register to IO_BASE_ADDR and access ports via offsets:
* sw $data, GPIO_PORT0_OFFSET($io_base) ; write GP0
* sw $data, GPIO_PORT1_OFFSET($io_base) ; write GP1 */
GPIO_PORT0_OFFSET = 0x1810,
GPIO_PORT1_OFFSET = 0x1814,
gp_Reset = (gcmd_Reset << gcmd_offset),
gp_DisplayEnabled = (gcmd_DisplayEnable << gcmd_offset | 0x0),
gp_DisplayDisabled = (gcmd_DisplayEnable << gcmd_offset | 0x1),
gp_DMA_FIFO = 1,
gp_DMA_CPU_to_GPU = 2,
gp_DMA_GPU_to_CPU = 3,
gp_DMA_Request = (gcmd_DMA_Request << gcmd_offset),
gp_HorizontalDisplayRange_3168_608 = (gcmd_HorizontalDisplayRange << gcmd_offset | 0xC60 << 12 | 0x260),
gp_VerticalDiplayRange = (gcmd_VerticalDisplayRange << gcmd_offset),
gp_VerticalDisplayRange_264_24 = (gp_VerticalDiplayRange | 264 << 10 | 24),
gp_VerticalDisplayRange_504_24 = (gp_VerticalDiplayRange | 504 << 10 | 24),
gp_DisplayMode = (gcmd_DisplayMode << gcmd_offset),
gp_Disp_HRes_256 = (0x0),
gp_Disp_HRes_320 = (0x1),
gp_Disp_HRes_512 = (0x2),
gp_Disp_HRes_640 = (0x3),
gp_Disp_VRes_240 = (0x0 << 2),
gp_Disp_VRes_480 = (0x1 << 2),
gp_Disp_Color15 = (0x0 << 4),
gp_Disp_Color24 = (0x1 << 4),
gp_Disp_VInterlace = (0x1 << 5),
gp_DisplayMode_320x240_15bit_NTSC = (gp_DisplayMode | gp_Disp_HRes_320 | gp_Disp_VRes_240 | gp_Disp_Color15),
gp_DisplayMOde_640x480_24bbp_NTSC = (gp_DisplayMode | gp_Disp_HRes_640 | gp_Disp_VRes_480 | gp_Disp_Color24 | gp_Disp_VInterlace),
gp_DrawMode_DrawAllowed = 10,
gp_SetDrawMode_DrawAllowed = (gcmd_SetDrawMode << gcmd_offset | 0x1 << gp_DrawMode_DrawAllowed),
gp_SetArea_TopLeft = (gcmd_SetDrawArea_TopLeft << gcmd_offset),
gp_SetArea_BottomRight = (gcmd_SetDrawArea_BotRight << gcmd_offset),
HW_GP0_ADDR = (IO_BASE_ADDR_HI16 << 16) | GPIO_PORT0_OFFSET,
HW_GP1_ADDR = (IO_BASE_ADDR_HI16 << 16) | GPIO_PORT1_OFFSET,
};
#define HW_GP0 C_(U4 V_*, HW_GP0_ADDR)
#define HW_GP1 C_(U4 V_*, HW_GP1_ADDR)
#define gp0_send(word) (HW_GP0[0] = (word))
#define gp1_send(word) (HW_GP1[0] = (word))
/* ============================================================================
* GP0 command byte constants + Layer 1 (GPU bitfield shifts)
* ============================================================================
* 8-bit GP0 opcodes (the upper byte of a primitive's first word). These are the BYTE only.
* NO macro body past this point uses a raw shift or raw mask.
* Mirrors the OPCODE_SHIFT / RS_SHIFT / REG_MASK convention from mips.h.
* ============================================================================ */
enum {
gp0_cmd_Nop = 0x00,
/* Cache management */
gp0_cmd_ClearCache = 0x01,
gp0_cmd_FillVram = 0x02,
gp0_cmd_CopyVram = 0x80,
gp0_cmd_CopyVramChained = 0x81,
gp0_cmd_ReadVram = 0xC0,
/* Polygons */
gp0_cmd_poly_f3 = 0x20, /* Flat Triangle */
gp0_cmd_poly_ft3 = 0x24, /* Flat Textured Triangle */
gp0_cmd_poly_g3 = 0x30, /* Gouraud Triangle */
gp0_cmd_poly_gt3 = 0x34, /* Gouraud Textured Tri */
gp0_cmd_poly_f4 = 0x28, /* Flat Quad */
gp0_cmd_poly_ft4 = 0x2C, /* Flat Textured Quad */
gp0_cmd_poly_g4 = 0x38, /* Gouraud Quad */
gp0_cmd_poly_gt4 = 0x3C, /* Gouraud Textured Quad */
/* Lines */
gp0_cmd_line_f2 = 0x40,
gp0_cmd_line_g2 = 0x50,
/* Sprites + Tiles + Rects */
gp0_cmd_sprt_1 = 0x64,
gp0_cmd_sprt_8 = 0x74,
gp0_cmd_sprt_16 = 0x7C,
gp0_cmd_tile_1 = 0x60,
gp0_cmd_tile_8 = 0x68,
gp0_cmd_tile_16 = 0x70,
/* State setters (not drawing primitives; set render context). */
gp0_cmd_DrawModeSetting = 0xE1, /* TPage / draw-mode (semi-trans, dither, etc.) */
gp0_cmd_SetTextureWindow = 0xE2,
gp0_cmd_SetDrawArea_TopLeft = 0xE3,
gp0_cmd_SetDrawArea_BotRight = 0xE4,
gp0_cmd_SetDrawOffset = 0xE5,
gp0_cmd_SetMaskBit = 0xE6,
/* bitfield shifts / widths / masks ----
* Generic GP0/GP1 command byte (upper 8 bits of every word sent to either port). */
gp0_cmd_shift = 24,
gp0_cmd_width = 8,
gp0_cmd_mask = 0xFF,
/* Color word layout (lives in Poly_F3.color, Poly_G4.c0..c3, etc.):
* bits 31..24 = command byte
* bits 23..16 = BLUE
* bits 15..08 = GREEN
* bits 07..00 = RED (PSX GPU is BGR, NOT RGB) */
gp0_color_cmd_shift = 24, gp0_color_cmd_width = 8, gp0_color_cmd_mask = 0xFF,
gp0_color_blue_shift = 16, gp0_color_blue_width = 8, gp0_color_blue_mask = 0xFF,
gp0_color_green_shift = 8, gp0_color_green_width = 8, gp0_color_green_mask = 0xFF,
gp0_color_red_shift = 0, gp0_color_red_width = 8, gp0_color_red_mask = 0xFF,
};
/* ============================================================================
* Layer 1.5 (per-field encoders) + Layer 2 (composite) + Layer 3 (semantic GP0 word builders)
* ============================================================================
* Layer 1.5 encoders take one field's value, mask it to its own width, and shift it to its own position.
* Mirrors `enc_op` / `enc_rs` / `enc_rt` in mips.h and `enc_gte_sf` / `enc_gte_mx` in gte.h.
* Layer-2 composite encoders OR the per-field encoders together; layer-3 semantic macros delegate to the composites.
* No raw shifts or magic numbers in any macro body below this point.
* ============================================================================ */
/* ---- Layer 1.5: per-field encoders ---- */
#define enc_gp0_cmd(cmd) (((cmd) & gp0_cmd_mask) << gp0_cmd_shift)
#define enc_gp0_color_cmd(cmd) (((cmd) & gp0_color_cmd_mask) << gp0_color_cmd_shift)
#define enc_gp0_color_r(r) (((r) & gp0_color_red_mask) << gp0_color_red_shift)
#define enc_gp0_color_g(g) (((g) & gp0_color_green_mask) << gp0_color_green_shift)
#define enc_gp0_color_b(b) (((b) & gp0_color_blue_mask) << gp0_color_blue_shift)
/* ---- Layer 2: composite encoders ---- */
#define enc_color_word(cmd, r, g, b) (enc_gp0_color_cmd(cmd) | enc_gp0_color_r(r) | enc_gp0_color_g(g) | enc_gp0_color_b(b))
#define enc_gp0_cmd_word(cmd) (enc_gp0_cmd(cmd))
/* ---- Layer 3: semantic GP0 word builders ---- */
/* Pre-baked color+command words for all 8 polygon variants.
* Mirrors `load_word` / `add_ui` / `jump_reg` style in mips.h. */
#define gp0_word_poly_f3(r,g,b) enc_color_word(gp0_cmd_poly_f3, (r),(g),(b))
#define gp0_word_poly_ft3(r,g,b) enc_color_word(gp0_cmd_poly_ft3, (r),(g),(b))
#define gp0_word_poly_g3(r,g,b) enc_color_word(gp0_cmd_poly_g3, (r),(g),(b))
#define gp0_word_poly_gt3(r,g,b) enc_color_word(gp0_cmd_poly_gt3, (r),(g),(b))
#define gp0_word_poly_f4(r,g,b) enc_color_word(gp0_cmd_poly_f4, (r),(g),(b))
#define gp0_word_poly_ft4(r,g,b) enc_color_word(gp0_cmd_poly_ft4, (r),(g),(b))
#define gp0_word_poly_g4(r,g,b) enc_color_word(gp0_cmd_poly_g4, (r),(g),(b))
#define gp0_word_poly_gt4(r,g,b) enc_color_word(gp0_cmd_poly_gt4, (r),(g),(b))
/* Cache management — bare-cmd words (no color/range payload). */
#define gp0_word_clear_cache() enc_gp0_cmd_word(gp0_cmd_ClearCache)
#define gp0_word_fill_vram() enc_gp0_cmd_word(gp0_cmd_FillVram)
#define gp0_word_copy_vram() enc_gp0_cmd_word(gp0_cmd_CopyVram)
#define gp0_word_read_vram() enc_gp0_cmd_word(gp0_cmd_ReadVram)
/* NOP — bare-cmd word (no effect; used as DR_ENV padding). */
#define gp0_word_nop() enc_gp0_cmd_word(gp0_cmd_Nop)
/* ============================================================================
* GP1 command byte constants + Layer 1 (display-mode + range + draw-area bitfield shifts)
* ============================================================================
* GP1 status bits are read from HW_GP1;
* ctrl writes use GP1 commands packed into 32-bit words
* (cmd byte in the upper 8 bits via `enc_gp0_cmd(cmd)`).
* ============================================================================ */
enum {
gp1_cmd_Reset = 0x00,
gp1_cmd_ResetCmdBuffer = 0x01,
gp1_cmd_AcknowledgeIRQ = 0x02,
gp1_cmd_DisplayEnable = 0x03,
gp1_cmd_DMADirection = 0x04,
gp1_cmd_StartDisplayArea = 0x05,
gp1_cmd_HorizontalDisplayRange = 0x06,
gp1_cmd_VerticalDisplayRange = 0x07,
gp1_cmd_DisplayMode = 0x08,
/* Note: GP1 only has commands 0x00..0x08.
* The state-setter commands (SetTextureWindow, * SetDrawArea*, SetDrawOffset, SetMaskBit)
* live in the GP0 enum as * 0xE1..0xE6.
* DrawArea word builders are below as GP0s * macros (since they emit GP0 commands). */
/* ---- Display-mode payload flags (per PSX-SPX §"GP1 Display Mode").
* Bit positions match the encoder shifts below; values are the *payload* bits only (cmd byte is OR'd in by enc_gp1_disp_mode_word). */
gp1_disp_HRes_256 = 0x0,
gp1_disp_HRes_320 = 0x1,
gp1_disp_HRes_512 = 0x2,
gp1_disp_HRes_640 = 0x3,
gp1_disp_VRes_240 = 0x0,
gp1_disp_VRes_480 = 0x1,
gp1_disp_Color15 = 0x0,
gp1_disp_Color24 = 0x1,
gp1_disp_VInterlace = 0x1,
/* ---- Layer 1: GP1 display-mode + range + draw-area shifts/masks ---- */
gp1_disp_hres_shift = 0, gp1_disp_hres_width = 2, gp1_disp_hres_mask = 0x3,
gp1_disp_vres_shift = 2, gp1_disp_vres_width = 1, gp1_disp_vres_mask = 0x1,
gp1_disp_color_shift = 4, gp1_disp_color_width = 1, gp1_disp_color_mask = 0x1,
gp1_disp_interlace_shift = 5, gp1_disp_interlace_width = 1, gp1_disp_interlace_mask = 0x1,
/* GP1 horizontal display range: bits 0..11 = X2, bits 12..23 = X1 */
gp1_hrange_x1_shift = 12, gp1_hrange_x1_width = 12, gp1_hrange_x1_mask = 0xFFF,
gp1_hrange_x2_shift = 0, gp1_hrange_x2_width = 12, gp1_hrange_x2_mask = 0xFFF,
/* GP1 vertical display range: bits 0..9 = Y2, bits 10..19 = Y1 */
gp1_vrange_y1_shift = 10, gp1_vrange_y1_width = 10, gp1_vrange_y1_mask = 0x3FF,
gp1_vrange_y2_shift = 0, gp1_vrange_y2_width = 10, gp1_vrange_y2_mask = 0x3FF,
/* GP1 draw area (top-left or bottom-right): bits 0..9 = X, bits 10..19 = Y
* (10-bit signed — caller pre-signs and masks with the named mask) */
gp1_draw_x_shift = 0, gp1_draw_x_width = 10, gp1_draw_x_mask = 0x3FF,
gp1_draw_y_shift = 10, gp1_draw_y_width = 10, gp1_draw_y_mask = 0x3FF,
};
/* ---- Layer 1.5: GP1 per-field encoders ---- */
#define enc_gp1_disp_hres(h) (((h) & gp1_disp_hres_mask) << gp1_disp_hres_shift)
#define enc_gp1_disp_vres(v) (((v) & gp1_disp_vres_mask) << gp1_disp_vres_shift)
#define enc_gp1_disp_color(c) (((c) & gp1_disp_color_mask) << gp1_disp_color_shift)
#define enc_gp1_disp_interlace(i) (((i) & gp1_disp_interlace_mask) << gp1_disp_interlace_shift)
#define enc_gp1_hrange_x1(x1) (((x1) & gp1_hrange_x1_mask) << gp1_hrange_x1_shift)
#define enc_gp1_hrange_x2(x2) (((x2) & gp1_hrange_x2_mask) << gp1_hrange_x2_shift)
#define enc_gp1_vrange_y1(y1) (((y1) & gp1_vrange_y1_mask) << gp1_vrange_y1_shift)
#define enc_gp1_vrange_y2(y2) (((y2) & gp1_vrange_y2_mask) << gp1_vrange_y2_shift)
#define enc_gp1_draw_x(x) (((x) & gp1_draw_x_mask) << gp1_draw_x_shift)
#define enc_gp1_draw_y(y) (((y) & gp1_draw_y_mask) << gp1_draw_y_shift)
/* ---- Layer 2: GP1 composite encoders ---- */
#define enc_gp1_disp_mode_word(h, v, c, i) (enc_gp0_cmd(gp1_cmd_DisplayMode) | enc_gp1_disp_hres(h) | enc_gp1_disp_vres(v) | enc_gp1_disp_color(c) | enc_gp1_disp_interlace(i))
#define enc_gp1_hrange_word(x1, x2) (enc_gp0_cmd(gp1_cmd_HorizontalDisplayRange) | enc_gp1_hrange_x1(x1) | enc_gp1_hrange_x2(x2))
#define enc_gp1_vrange_word(y1, y2) (enc_gp0_cmd(gp1_cmd_VerticalDisplayRange) | enc_gp1_vrange_y1(y1) | enc_gp1_vrange_y2(y2))
/* ---- Layer 2: GP0 state-setter composite encoders ----
* GP0(0xE3) SetDrawArea top-left and GP0(0xE4) SetDrawArea bottom-right both use the same X/Y 10-bit signed payload as GP1 DisplayRange. */
#define enc_gp0_draw_area_tl_word(x, y) (enc_gp0_cmd(gp0_cmd_SetDrawArea_TopLeft) | enc_gp1_draw_x(x) | enc_gp1_draw_y(y))
#define enc_gp0_draw_area_br_word(x, y) (enc_gp0_cmd(gp0_cmd_SetDrawArea_BotRight) | enc_gp1_draw_x(x) | enc_gp1_draw_y(y))
/* ---- Layer 3: GP1 semantic word builders ---- */
#define gp1_word_Reset() enc_gp0_cmd_word(gp1_cmd_Reset)
#define gp1_word_ResetCmdBuffer() enc_gp0_cmd_word(gp1_cmd_ResetCmdBuffer)
#define gp1_word_AcknowledgeIRQ() enc_gp0_cmd_word(gp1_cmd_AcknowledgeIRQ)
#define gp1_word_StartDisplayArea() enc_gp0_cmd_word(gp1_cmd_StartDisplayArea)
#define gp1_word_display_enable(on) (enc_gp0_cmd(gp1_cmd_DisplayEnable) | ((on) & 1))
#define gp1_word_display_disable() gp1_word_display_enable(0)
#define gp1_word_display_mode_320x240_15bit_ntsc enc_gp1_disp_mode_word(gp1_disp_HRes_320, gp1_disp_VRes_240, gp1_disp_Color15, 0)
#define gp1_word_display_mode_640x480_24bit_ntsc_interlaced enc_gp1_disp_mode_word(gp1_disp_HRes_640, gp1_disp_VRes_480, gp1_disp_Color24, gp1_disp_VInterlace)
#define gp1_word_horizontal_range(x1, x2) enc_gp1_hrange_word((x1), (x2))
#define gp1_word_vertical_range(y1, y2) enc_gp1_vrange_word((y1), (y2))
/* ---- Layer 3: GP0 state-setter semantic word builders ---- */
/* DrawArea: top-left = (X, Y), bottom-right = (X, Y) — X/Y in 10-bit signed.
* Caller is responsible for sign-conversion before passing in. */
#define gp0_word_draw_area_top_left(x, y) enc_gp0_draw_area_tl_word((x), (y))
#define gp0_word_draw_area_bottom_right(x, y) enc_gp0_draw_area_br_word((x), (y))
/* ============================================================================
* Pre-baked GPU state words
* ============================================================================
* Common command words for boot-time GPU init and standard display configurations.
* ============================================================================ */
/* ---- Display enable (1-bit payload on DisplayEnable cmd) ---- */
#define gp1_word_display_enabled enc_gp0_cmd_word(gp1_cmd_DisplayEnable)
#define gp1_word_display_disabled (enc_gp0_cmd_word(gp1_cmd_DisplayEnable) | 1)
#define gp1_word_DisplayOn() gp1_word_display_enable(0)
#define gp1_word_DisplayOff() gp1_word_display_enable(1)
/* ---- DMA direction (2-bit payload on DMADirection cmd 0x04) ---- */
enum {
gp1_dma_dir_Off = 0,
gp1_dma_dir_FIFO = 1,
gp1_dma_dir_CPU_to_GPU = 2,
gp1_dma_dir_GPUREAD_to_CPU = 3,
};
#define gp1_word_dma_direction(dir) (enc_gp0_cmd(gp1_cmd_DMADirection) | ((dir) & 0x3))
#define gp1_word_dma_to_gpu() gp1_word_dma_direction(gp1_dma_dir_CPU_to_GPU)
#define gp1_word_dma_read_cpu() gp1_word_dma_direction(gp1_dma_dir_GPUREAD_to_CPU)
/* ---- Standard display ranges (NTSC + PAL pre-baked) ---- */
/* Horizontal range values are in video clock units (8 units/pixel); vertical range values are scanline numbers. */
enum {
/* NTSC horizontal range: X1=608, X2=3168 */
gp1_hrange_NTSC_x1 = 0x260,
gp1_hrange_NTSC_x2 = 0xC60,
/* PAL horizontal range (same as NTSC for most CRTs) */
gp1_hrange_PAL_x1 = 0x260,
gp1_hrange_PAL_x2 = 0xC60,
/* NTSC vertical range: Y1=24, Y2=264 */
gp1_vrange_NTSC_y1 = 24,
gp1_vrange_NTSC_y2 = 264,
/* PAL vertical range: Y1=24, Y2=504 */
gp1_vrange_PAL_y1 = 24,
gp1_vrange_PAL_y2 = 504,
};
#define gp1_word_horizontal_range_ntsc enc_gp1_hrange_word(gp1_hrange_NTSC_x1, gp1_hrange_NTSC_x2)
#define gp1_word_horizontal_range_pal enc_gp1_hrange_word(gp1_hrange_PAL_x1, gp1_hrange_PAL_x2)
#define gp1_word_vertical_range_ntsc enc_gp1_vrange_word(gp1_vrange_NTSC_y1, gp1_vrange_NTSC_y2)
#define gp1_word_vertical_range_pal enc_gp1_vrange_word(gp1_vrange_PAL_y1, gp1_vrange_PAL_y2)
/* ---- Draw-mode setting (TPage / draw-area allowance) ---- */
/* The "drawing enabled" word is the standard post-init state. */
enum {
/* Per psx-spx, the standard 0xE1 layout has dfe at bit 10. But libpsyx's PutDrawEnv
* uses bit 19 (in the "unused" 14-23 range) for dfe in the DR_ENV code[0] — and the
* PSX hardware honors bit 19 in the DR_ENV context (not bit 10). So we need a
* separate bit definition for the DR_ENV-specific DrawMode. */
gp0_DrawMode_DrawToDispBit = 10, // standard psx-spx bit 10 (dfe)
gp0_DrawMode_DR_ENV_DrawToDispBit = 19, // libpsyx DR_ENV code[0] (dfe in DR_ENV context)
gp0_DrawMode_DR_ENV_isbgBit = 19, // libpsyx uses bit 19 for isbg too
};
#define gp0_word_draw_mode_drawing_allowed (enc_gp0_cmd(gp0_cmd_DrawModeSetting) | (1 << gp0_DrawMode_DrawToDispBit))
/* DR_ENV-specific DrawMode variants (libpsyx SetDrawEnv layout).
* The DR_ENV is a 16-word packet emitted at boot by gp_screen_init's ac_put_draw_env_demo
* atom component. Within the DR_ENV, the 0xE1 command is reused in three different bit
* configurations:
* code[0] = `gp0_word_draw_mode_drawing_allowed` (dfe=1; standard post-init state)
* code[6] = `gp0_word_dr_env_bg_color_cmd(isbg, r, g, b)` (initial-bg-color path)
* code[7] = `gp0_word_dr_env_draw_mode(isbg)` (isbg-flag path)
* Bits 0-23 of the 0xE1 word are the payload; bits 24-31 are the cmd byte (0xE1). */
#define gp0_word_dr_env_bg_color_cmd(isbg, r, g, b) (enc_gp0_cmd(gp0_cmd_DrawModeSetting) | (1 << gp0_DrawMode_DrawToDispBit) | ((isbg) ? gp0_dr_env_isbg_bit : 0) | enc_gp0_color_r(r) | enc_gp0_color_g(g) | enc_gp0_color_b(b))
#define gp0_word_dr_env_draw_mode(isbg) (enc_gp0_cmd(gp0_cmd_DrawModeSetting) | (1 << gp0_DrawMode_DrawToDispBit) | ((isbg) ? gp0_dr_env_isbg_bit : 0))
/* State-setter bare-cmd words (no immediate payload; the GPU uses the current state machine already programmed). */
#define gp0_word_set_texture_window() enc_gp0_cmd_word(gp0_cmd_SetTextureWindow)
#define gp0_word_set_draw_offset() enc_gp0_cmd_word(gp0_cmd_SetDrawOffset)
#define gp0_word_set_mask_bit() enc_gp0_cmd_word(gp0_cmd_SetMaskBit)
/* DR_ENV code[5] Mask (0xE6 cmd + isbg bit). The isbg bit is set so the GPU knows the auto-clear path is active (paired with code[6] + code[7]). */
#define gp0_word_dr_env_mask() (gp0_word_set_mask_bit() | gp0_dr_env_isbg_bit)
/* DR_ENV pre-baked constants (libpsyx PutDrawEnv layout).
* DR_ENV is a 16-word packet: tag = (length << 24) | addr, where length = 15 (15 code words follow) and addr = 0 (chain to nothing). */
enum {
PolyTag_len_bits = 8,
PolyTag_addr_bits = 24,
gp0_dr_env_tag = (15 << 24) | 0x00FFFFFF,
gp0_dr_env_isbg_bit = (1 << gp0_DrawMode_DR_ENV_isbgBit),
};
/* ---- DrawArea at origin (0,0) and full screen (320x240) ---- */
#define gp0_word_draw_area_top_left_origin enc_gp0_draw_area_tl_word(0, 0)
#define gp0_word_draw_area_bottom_right_320x240 enc_gp0_draw_area_br_word(319, 239)
#define gp0_word_draw_area_bottom_right_640x480 enc_gp0_draw_area_br_word(639, 479)
#pragma endregion GPU Ports & Commands
#pragma region GPU Status
/* ============================================================================
* GPU status register bits
* ============================================================================
* Read from HW_GP1; the lower bits are DMA-block-size (variable-width).
* ============================================================================ */
enum {
gp1_Status_BitReady = 31,
gp1_Status_BitSendingDMA = 25,
gp1_Status_DMABlockSizeShift = 0,
};
#define gp1_status_is_ready() ((HW_GP1[0] >> gp1_Status_BitReady) & 1)
#define gp1_status_is_sending_dma() ((HW_GP1[0] >> gp1_Status_BitSendingDMA) & 1)
#pragma endregion GPU Status
#pragma region Primitives
/* ============================================================================
* Primitive structs (8 polygon variants + tag)
* ============================================================================
* Each struct follows the GPU-documented memory layout for the corresponding primitive command.
* The PolyTag is the OT-link header; the rest of the struct is the primitive's body.
*
* The current working layouts match the existing demo
* (floor_tri uses Poly_F3; cube_tri uses Poly_G4).
* They are NOT necessarily byte-identical to the PSX-SPX reference layout.
* The demo layout uses color+vertex interleaving that doesn't match the standard PSX SDK file format.
* For PSX-SDK file compatibility, the textured variants (FT*, GT*) would need layout adjustments.
* ============================================================================ */
/* ---------- RGB8 (3-byte packed color) ---------- */
typedef Struct_(RGB8) { B1 r; B1 g; B1 b; };
#define rgb8(r, g, b) (RGB8){ r, g, b }
#define rgb8(r,g,b) ((RGB8){r,g,b})
typedef B1 gp_Pixel16[1];
typedef B1 gp_Pixel24[3];
enum {
gp_b10_X = 0,
gp_b10_Y = 10,
gp_b16_X = 0,
gp_b16_Y = 16,
/* ---------- PolyTag (the OT-link header; 1 word) ---------- */
// enum {
// PolyTag_len_bits = 8,
// PolyTag_addr_bits = 24,
// };
typedef Struct_(PolyTag) {
union {
U4 code;
struct {
U4 addr: 24;
U4 len: 8;
};
};
};
typedef Struct_(gp_Vec2) { U2 y; U2 x; };
/* DSL cast convention: every cast uses `C_()`, every pointer qualifier is `R_` (restrict) or `V_` (volatile).
* No raw C-style casts. RHS values are assumed to be `U4` — caller passes a `U4` directly. */
#define set_len(tag,v) (C_(PolyTag_R,tag)->len = u4_(v))
#define set_addr(tag,v) (C_(PolyTag_R,tag)->addr = u4_(v))
/* `set_code` is no longer in the new PolyTag design — the code byte lives in the primitive body
* (e.g. `((Poly_F3*)(p))->code`), not in the tag.
* Use the typed primitive structs (Poly_F3, Poly_G4, etc.) and the `set_poly_*` setters,
* which set both the tag's length and the code. */
#define get_len(tag) C_(U4,C_(PolyTag_R,tag)->len)
#define get_addr(tag) C_(U4,C_(PolyTag_R,tag)->addr)
#if 1
void gp_screen_init(void) __asm__("gp_screen_init_asm");
#else
#define gp_screen_init() gp_screen_init_c11()
#endif
/* ---------- Poly_F3 (Flat Triangle; 5 words) ---------- */
typedef Struct_(Poly_F3) {
U4 tag;
RGB8 color;
B1 code;
union {
struct { V2_S2 p0; V2_S2 p1; V2_S2 p2; };
A3_V2_S2 points;
};
};
/* ---------- Poly_F4 (Flat Quad; 6 words) ---------- */
typedef Struct_(Poly_F4) {
U4 tag;
RGB8 color;
B1 code;
union {
struct { V2_S2 p0; V2_S2 p1; V2_S2 p2; V2_S2 p3; };
A4_V2_S2 points;
};
};
/* ---------- Poly_G3 (Gouraud Triangle; 7 words) ---------- */
typedef Struct_(Poly_G3) {
U4 tag; RGB8 c0; B1 code;
V2_S2 p0; RGB8 c1; B1 pad1;
V2_S2 p1; RGB8 c2; B1 pad2;
V2_S2 p2;
};
// TODO REVIEW:
/* ---------- Poly_G4 (Gouraud Quad; 9 words) ---------- */
typedef Struct_(Poly_G4) {
U4 tag; RGB8 c0; B1 code;
V2_S2 p0; RGB8 c1; B1 pad1;
V2_S2 p1; RGB8 c2; B1 pad2;
V2_S2 p2; RGB8 c3; B1 pad3;
V2_S2 p3;
};
/* --- GPU Command Semantics (GP0) --- */
/* ---------- Poly_FT3 (Flat Textured Triangle; placeholder layout) ---------- */
/* TODO(Ed): verify the textured-variant layout against PSX-SPX when needed. */
typedef Struct_(Poly_FT3) {
U4 tag;
RGB8 color;
B1 code;
U4 tpage;
U4 clut;
V2_S2 p0; U1 u0; U1 v0;
V2_S2 p1; U1 u1; U1 v1;
V2_S2 p2; U1 u2; U1 v2;
};
#define GPU_CMD_CLEAR_CACHE 0x01
#define GPU_CMD_VRAM_FILL 0x02
#define GPU_CMD_VRAM_COPY 0x80
#define GPU_CMD_VRAM_READ 0xC0
#define GPU_CMD_POLY_F3 0x20 /* Flat Triangle */
#define GPU_CMD_POLY_FT3 0x24 /* Flat Textured Triangle */
#define GPU_CMD_POLY_G3 0x30 /* Gouraud Triangle */
#define GPU_CMD_POLY_GT3 0x34 /* Gouraud Textured Triangle */
#define GPU_CMD_POLY_F4 0x28 /* Flat Quad */
#define GPU_CMD_POLY_FT4 0x2C /* Flat Textured Quad */
#define GPU_CMD_POLY_G4 0x38 /* Gouraud Quad */
#define GPU_CMD_POLY_GT4 0x3C /* Gouraud Textured Quad */
/* ---------- Poly_FT4 (Flat Textured Quad) ---------- */
typedef Struct_(Poly_FT4) {
U4 tag;
RGB8 color;
B1 code;
U4 tpage;
U4 clut;
V2_S2 p0; U1 u0; U1 v0;
V2_S2 p1; U1 u1; U1 v1;
V2_S2 p2; U1 u2; U1 v2;
V2_S2 p3; U1 u3; U1 v3;
};
/* --- Hardware MMIO Addresses --- */
/* ---------- Poly_GT3 (Gouraud Textured Triangle) ---------- */
typedef Struct_(Poly_GT3) {
U4 tag; RGB8 c0; B1 code;
V2_S2 p0; RGB8 c1; B1 pad1;
V2_S2 p1; RGB8 c2; B1 pad2;
V2_S2 p2;
U4 tpage;
U4 clut;
V2_S2 tp0; U1 u0; U1 v0;
V2_S2 tp1; U1 u1; U1 v1;
V2_S2 tp2; U1 u2; U1 v2;
};
#define HW_GP0_ADDR 0x1F801810 /* GPU Data Port */
#define HW_GP1_ADDR 0x1F801814 /* GPU Status/Control Port */
/* ---------- Poly_GT4 (Gouraud Textured Quad) ---------- */
typedef Struct_(Poly_GT4) {
U4 tag; RGB8 c0; B1 code;
V2_S2 p0; RGB8 c1; B1 pad1;
V2_S2 p1; RGB8 c2; B1 pad2;
V2_S2 p2; RGB8 c3; B1 pad3;
V2_S2 p3;
U4 tpage;
U4 clut;
V2_S2 tp0; U1 u0; U1 v0;
V2_S2 tp1; U1 u1; U1 v1;
V2_S2 tp2; U1 u2; U1 v2;
V2_S2 tp3; U1 u3; U1 v3;
};
/* ---------- Primitive setters (C-level) ----------
* DSL cast convention: every cast via C_(), every pointer via R_/V_. */
#define set_poly_f3(p) set_len(p, 4), C_(Poly_F3_R, p)->code = gp0_cmd_poly_f3
#define set_poly_ft3(p) set_len(p, 7), C_(Poly_FT3_R,p)->code = gp0_cmd_poly_ft3
#define set_poly_g3(p) set_len(p, 6), C_(Poly_G3_R, p)->code = gp0_cmd_poly_g3
#define set_poly_gt3(p) set_len(p, 9), C_(Poly_GT3_R,p)->code = gp0_cmd_poly_gt3
#define set_poly_f4(p) set_len(p, 5), C_(Poly_F4_R, p)->code = gp0_cmd_poly_f4
#define set_poly_ft4(p) set_len(p, 9), C_(Poly_FT4_R,p)->code = gp0_cmd_poly_ft4
#define set_poly_g4(p) set_len(p, 8), C_(Poly_G4_R, p)->code = gp0_cmd_poly_g4
#define set_poly_gt4(p) set_len(p, 12), C_(Poly_GT4_R,p)->code = gp0_cmd_poly_gt4
/* ---------- Ordering table ops ---------- */
#define orderingtbl_add_primitive(ot, p) set_addr(p, get_addr(ot)), set_addr(ot, p)
#define orderingtbl_add_primitives(ot, p0, p1) set_addr(p1, get_addr(ot)), set_addr(ot, p0)
#pragma endregion Primitives
#pragma region TPage
/* ============================================================================
* Texture Page (TPage) bit layout
* ============================================================================
* The TPage data word sent via GP0(0x2X) has:
* bits 0..3 = texture page X (4 bits, 64-px units, 0..16)
* bit 4 = texture page Y (1 bit, 64-px units, 0/1)
* bits 5..6 = semi-transparency (2 bits, 0..3)
* bits 7..8 = texture page colors (2 bits, 4bpp/8bpp/16bpp/2bpp-mixed)
* bit 9 = dither (1 bit, 0/1)
* bit 10 = drawing to display area (1 bit)
* bit 11 = texture disable (1 bit)
* bits 12..31 = reserved (zero)
* ============================================================================ */
enum {
/* ---- Layer 1: TPage bitfield shifts / widths / masks ---- */
gp0_tpage_x_shift = 0, gp0_tpage_x_width = 4, gp0_tpage_x_mask = 0xF,
gp0_tpage_y_shift = 4, gp0_tpage_y_width = 1, gp0_tpage_y_mask = 0x1,
gp0_tpage_semi_trans_shift = 5, gp0_tpage_semi_trans_width = 2, gp0_tpage_semi_trans_mask = 0x3,
gp0_tpage_color_depth_shift = 7, gp0_tpage_color_depth_width = 2, gp0_tpage_color_depth_mask = 0x3,
gp0_tpage_dither_shift = 9, gp0_tpage_dither_width = 1, gp0_tpage_dither_mask = 0x1,
gp0_tpage_draw_to_disp_shift = 10, gp0_tpage_draw_to_disp_width = 1, gp0_tpage_draw_to_disp_mask = 0x1,
gp0_tpage_tex_disable_shift = 11, gp0_tpage_tex_disable_width = 1, gp0_tpage_tex_disable_mask = 0x1,
/* TPage color-depth payload values (NOT bit positions — these go in
* the 2-bit field at gp0_tpage_color_depth_shift). */
gp0_tpage_color_4bpp = 0x0,
gp0_tpage_color_8bpp = 0x1,
gp0_tpage_color_16bpp = 0x2,
/* Default TPage value libpsyx's SetDefDrawEnv writes (matches the `li v1, 10; sh v1, 20(v0)` sequence at C11_only.elf:0x8001273C). */
gp0_tpage_default = 10,
/* TPage semi-transparency mode payload values (NOT bit positions). */
gp0_tpage_semi_trans_none = 0x0,
gp0_tpage_semi_trans_alpha = 0x1,
gp0_tpage_semi_trans_add = 0x2,
gp0_tpage_semi_trans_sub = 0x3,
};
/* ---- Layer 1.5: TPage per-field encoders. Mirrors enc_gte_sf/mx/v in gte.h. ---- */
#define enc_gp0_tpage_x(x) (((x) & gp0_tpage_x_mask) << gp0_tpage_x_shift)
#define enc_gp0_tpage_y(y) (((y) & gp0_tpage_y_mask) << gp0_tpage_y_shift)
#define enc_gp0_tpage_semi_trans(s) (((s) & gp0_tpage_semi_trans_mask) << gp0_tpage_semi_trans_shift)
#define enc_gp0_tpage_color_depth(c) (((c) & gp0_tpage_color_depth_mask) << gp0_tpage_color_depth_shift)
#define enc_gp0_tpage_dither(d) (((d) & gp0_tpage_dither_mask) << gp0_tpage_dither_shift)
#define enc_gp0_tpage_draw_to_disp(d) (((d) & gp0_tpage_draw_to_disp_mask) << gp0_tpage_draw_to_disp_shift)
#define enc_gp0_tpage_tex_disable(t) (((t) & gp0_tpage_tex_disable_mask) << gp0_tpage_tex_disable_shift)
/* ---- Layer 2: TPage composite encoder. Mirrors enc_gte_cmdw in gte.h ---- */
#define enc_gp0_tpage_word(x, y, semi_trans, color_depth, dither, draw_to_disp, tex_disable) \
(enc_gp0_tpage_x(x) \
| enc_gp0_tpage_y(y) \
| enc_gp0_tpage_semi_trans(semi_trans) \
| enc_gp0_tpage_color_depth(color_depth) \
| enc_gp0_tpage_dither(dither) \
| enc_gp0_tpage_draw_to_disp(draw_to_disp) \
| enc_gp0_tpage_tex_disable(tex_disable))
typedef Struct_(TexturePage) { U4 raw; };
/* ---- Layer 3: TPage semantic word builder ---- */
#define gp0_word_tpage(x, y, semi_trans, color_depth, dither, draw_to_disp, tex_disable) \
enc_gp0_tpage_word((x), (y), (semi_trans), (color_depth), (dither), (draw_to_disp), (tex_disable))
#pragma endregion TPage
#pragma region CLUT
/* ============================================================================
* CLUT (Color Look-Up Table) semantics
* ============================================================================
* CLUT is loaded into VRAM by sending a GP0 command whose payload is:
* bits 0..5 = Y in 16-px units (palette row)
* bits 6..14 = X in 16-px units (palette column)
* bits 15..23 = reserved (zero)
* bits 24..31 = command byte — 0x20 (4bpp load) or 0x25 (8bpp load)
* ============================================================================ */
enum {
/* ---- Layer 1: CLUT bitfield shifts / widths / masks ---- */
gp0_clut_y_shift = 0, gp0_clut_y_width = 6, gp0_clut_y_mask = 0x3F,
gp0_clut_x_shift = 6, gp0_clut_x_width = 9, gp0_clut_x_mask = 0x1FF,
/* CLUT-load cmd-byte variants — the upper byte of the GP0 word. */
gp0_clut_cmd_Load4bpp = 0x20,
gp0_clut_cmd_Load8bpp = 0x25,
};
/* ---- Layer 1.5: CLUT per-field encoders ---- */
#define enc_gp0_clut_x(x) (((x) & gp0_clut_x_mask) << gp0_clut_x_shift)
#define enc_gp0_clut_y(y) (((y) & gp0_clut_y_mask) << gp0_clut_y_shift)
/* ---- Layer 2: CLUT composite encoder ---- */
#define enc_gp0_clut_word(cmd, x, y) (enc_gp0_cmd(cmd) | enc_gp0_clut_x(x) | enc_gp0_clut_y(y))
/* ---- Layer 3: CLUT semantic word builders — one per depth variant,
* named cmd-byte (no opaque ternary). ---- */
#define gp0_word_clut_load_4bpp(x, y) enc_gp0_clut_word(gp0_clut_cmd_Load4bpp, (x), (y))
#define gp0_word_clut_load_8bpp(x, y) enc_gp0_clut_word(gp0_clut_cmd_Load8bpp, (x), (y))
#pragma endregion CLUT
#pragma region TIM File Format
/* ============================================================================
* TIM file format constants and headers
* ============================================================================
* TIM (Sony .TIM texture image) file structure:
* +0x00 U4 file_id (always 0x10 = TIM magic)
* +0x04 U4 version (always 0x00 for v1)
* +0x08 U4 flags (bits 0..2 = type, bit 3 = has_CLUT)
* +0x0C ... CLUT section (if flags & 0x8)
* +0x00 U4 clut_section_length
* +0x04 U2 clut_org_x
* +0x06 U2 clut_org_y
* +0x08 U2 num_colors
* +0x0A U2 depth_bpp
* +0x0C ... palette data
* ... ... Pixel section
* +0x00 U4 px_section_length
* +0x04 U2 px_width
* +0x06 U2 px_height
* +0x08 ... pixel data
*
* Future?: add `tim_load_to_vram(tim_ptr, vram_addr)` that emits the necessary GP0 commands.
* Stoppped for now at the struct + enum level.
* ============================================================================ */
enum {
tim_file_id_magic = 0x10,
tim_type_4bpp = 0x00,
tim_type_8bpp = 0x01,
tim_type_16bpp = 0x02,
tim_type_32bpp = 0x03,
tim_type_mixed = 0x04,
tim_flag_has_clut = 0x08,
};
typedef Struct_(TIM_Header) {
U4 file_id; /* always 0x10 = "TIM" magic */
U4 version; /* ignored; always 0 */
U4 flags; /* bits 0..2 = type, bit 3 = has_clut */
};
typedef Struct_(TIM_SectionHeader) {
U4 section_length; /* bytes in this section including this header */
U2 org_x; /* origin in VRAM */
U2 org_y;
U2 width; /* width in pixels */
U2 height; /* height in pixels */
};
#pragma endregion TIM File Format
#pragma region Tape-Side Macros
/* ============================================================================
* Tape-side GPU operations (NOT in this header)
* ============================================================================
*
* No `mac_gp0_send` or related macros live in gp.h.
* Rationale: the Lottes tape model uses OT-DMA for primitive submission, so atom bodies write to main RAM (the OT/primitive buffer)
* and to GTE state — never directly to the GPU ports at 0x1F801810 / 0x1F801814.
* See `mac_format_f3_color`, `mac_insert_ot_tag`, `mac_gte_store_f3` in lottes_tape.h for the patterns atom bodies actually use.
*
* If a feature need arises requires tape-side GPU port writes
* (e.g. DMA-kick to start GPU consumption of the OT, VBlank sync via GP1 status poll),
* the right home is `lottes_tape.h` alongside the rest of the `mac_*` family:
* 1. The caller pins a register to hold the IO base, e.g. register U4 r_io rgcc(R_T4) = IO_BASE_ADDR;
* The compiler emits `lui R_T4, IO_BASE_ADDR_HI16` outside the atom body (in the C prologue before tape_run).
* 2. The atom body uses `store_word(R_data, R_T4, GPIO_PORT0_OFFSET)` to write to GP0, and `store_word(R_data, R_T4, GPIO_PORT1_OFFSET)`
* to write to GP1. Both are preprocessor-encodable because R_T4 is a fixed register and the GPIO_PORT*_OFFSET constants
* fit in the `sw`'s 16-bit signed offset field. No placeholder-pun, no asm constraints, no hidden register choice.
* Same pattern as the old graphics_hello/hello_gp_routines.s `reg_io_offset`/`gcmd_push` convention.
*
* This mirrors the existing tape-side wave-context discipline:
* the caller binds the IO-base register via `rgcc()`, the macro assumes the binding is in effect,
* and the encoding falls out at preprocessor time.
* No additional GPU-domain macro layer required.
* ============================================================================ */
#pragma endregion Tape-Side Macros
+52
View File
@@ -0,0 +1,52 @@
/* ============================================================================
* duffle DSL — GPU Vendor Mnemonics (opt-in)
* ============================================================================
*
* Provides the PSYQ-style CamelCase aliases for the duffle GPU primitive setters and OT operations.
* The duffle snake_case names are primary; this header is for users who prefer the PSYQ SDK function names from the legacy C API.
*
* USAGE: #include "duffle/gp_vendor_sym.h" // after gp.h
*
* Mapping (vendor -> duffle):
* Primitive setters (PSYQ SDK-style):
* setPolyF3 -> set_poly_f3
* setPolyF4 -> set_poly_f4
* setPolyG3 -> set_poly_g3
* setPolyG4 -> set_poly_g4
* setPolyFT3 -> set_poly_ft3
* setPolyFT4 -> set_poly_ft4
* setPolyGT3 -> set_poly_gt3
* setPolyGT4 -> set_poly_gt4
*
* OT operations:
* AddPrim(ot, p) -> orderingtbl_add_primitive(ot, p)
*
* The gp0_cmd_* / gp1_cmd_* byte constants are already short and descriptive; no vendor alias is provided for them.
*
* The vendor mnemonics are NOT registered with the duffle word-count metadata (word_counts.metadata.h).
* They expand to the duffle macros which DO have word-count entries
* (the ones emitted by mac_format_f3_color / mac_gte_store_f3 / etc.). Verification: V13 (objdump byte-identical) holds.
* ============================================================================ */
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "gp.h"
#endif
#ifndef DUFFLE_GP_VENDOR_SYM_H
#define DUFFLE_GP_VENDOR_SYM_H
/* Primitive setters (PSYQ SDK-style) */
#define setPolyF3(p) set_poly_f3(p)
#define setPolyF4(p) set_poly_f4(p)
#define setPolyG3(p) set_poly_g3(p)
#define setPolyG4(p) set_poly_g4(p)
#define setPolyFT3(p) set_poly_ft3(p)
#define setPolyFT4(p) set_poly_ft4(p)
#define setPolyGT3(p) set_poly_gt3(p)
#define setPolyGT4(p) set_poly_gt4(p)
/* OT operations */
#define AddPrim(ot, p) orderingtbl_add_primitive((ot), (p))
#endif
+76
View File
@@ -0,0 +1,76 @@
#ifdef INTELLISENSE_DIRECTIVES
# include "gen/macs.h"
# include "gen/offsets.h"
# include "gte.h"
# include "gp.h"
# include "lottes_tape.h"
#endif
ATOM_FILE_DEBUGGER_LINE_MARKER(gte_atom_c);
#pragma region MACs (Mips Atom Components)
/* Words: 3; Loads 3 S2 indices from the face array */
FI_ Slice_MipsCode ac_load_tri_indices(U4 r_face_cusor, U4 r_i0, U4 r_i1, U4 r_i2) atom_dbg_skip MipsAtomComp_Proc_(ac_load_tri_indices, {
load_half_u(r_i0, r_face_cusor, 0 * S_(S2)),
load_half_u(r_i1, r_face_cusor, 1 * S_(S2)),
load_half_u(r_i2, r_face_cusor, 2 * S_(S2)),
})
/* Words: 3; Stores the 3 transformed (V2_S2 screen) vertices to the F3.
* PIPELINE: post-RTPT (SXY0=v0.screen, SXY1=v1.screen, SXY2=v2.screen). */
FI_ Slice_MipsCode ac_gte_store_f3(U4 r_primitive_cursor) atom_dbg_skip MipsAtomComp_Proc_(ac_gte_store_f3, {
gte_sw(C2_SXY0, r_primitive_cursor, O_(Poly_F3,p0)),
gte_sw(C2_SXY1, r_primitive_cursor, O_(Poly_F3,p1)),
gte_sw(C2_SXY2, r_primitive_cursor, O_(Poly_F3,p2)),
})
/* Words: 18; Translates indices to vertex addresses and pushes them to GTE */
I_ Slice_MipsCode ac_gte_load_tri_verts(U4 r_vert_base, U4 r_v0, U4 r_v1, U4 r_v2) atom_dbg_skip MipsAtomComp_Proc_(ac_gte_load_tri_verts, {
shift_lleft(R_AT, r_v0, v3s2_byteoff), add_u_self(R_AT, r_vert_base), load_word(R_V0, R_AT, O_(V3_S2,x)), load_word(R_V1, R_AT, O_(V3_S2,z)), gte_mv_to_data_r(R_V0, C2_VXY0), gte_mv_to_data_r(R_V1, C2_VZ0),
shift_lleft(R_AT, r_v1, v3s2_byteoff), add_u_self(R_AT, r_vert_base), load_word(R_V0, R_AT, O_(V3_S2,x)), load_word(R_V1, R_AT, O_(V3_S2,z)), gte_mv_to_data_r(R_V0, C2_VXY1), gte_mv_to_data_r(R_V1, C2_VZ1),
shift_lleft(R_AT, r_v2, v3s2_byteoff), add_u_self(R_AT, r_vert_base), load_word(R_V0, R_AT, O_(V3_S2,x)), load_word(R_V1, R_AT, O_(V3_S2,z)), gte_mv_to_data_r(R_V0, C2_VXY2), gte_mv_to_data_r(R_V1, C2_VZ2),
})
/* Words: 3; Stores the 3 transformed (V2_S2 screen) vertices of the
* G4 triangle portion to p0/p1/p2.
* PIPELINE: post-RTPT, pre-RTPS (SXY0=v0.screen, SXY1=v1.screen, SXY2=v2.screen).
* MUST be called BEFORE V3-RTPS, otherwise SXY0/1/2 get overwritten with v3
* (RTPS writes only to SXY2, but to keep the three registers aligned with v0/v1/v2 you must store before RTPS). */
FI_ Slice_MipsCode ac_gte_store_g4_p012(U4 r_primitive_cursor) atom_dbg_skip MipsAtomComp_Proc_(ac_gte_store_g4_p012, {
gte_sw(C2_SXY0, r_primitive_cursor, O_(Poly_G4,p0)),
gte_sw(C2_SXY1, r_primitive_cursor, O_(Poly_G4,p1)),
gte_sw(C2_SXY2, r_primitive_cursor, O_(Poly_G4,p2)),
})
/* Words: 1; Stores the V3 screen coord to the G4's p3 slot.
* PIPELINE: post-RTPS (SXY2 holds v3.screen because RTPS writes its single-vertex result to SXY2;
* SXY0 still holds v0.screen from the earlier RTPT.
*/
FI_ Slice_MipsCode ac_gte_store_g4_p3(U4 r_primitive_cursor) atom_dbg_skip MipsAtomComp_Proc_(ac_gte_store_g4_p3, { gte_sw(C2_SXY2, r_primitive_cursor, O_(Poly_G4,p3)) })
#pragma endregion MACs (Mips Atom Components)
#pragma region Bsked Atoms
typedef Struct_(Binds_SetGteWorld) {
M3_S2* transform;
};
internal MipsAtom_(set_gte_world) atom_info(
atom_bind(Binds_SetGteWorld)
, atom_reads(R_TapePtr)
){
/* Pop matrix address from tape into R_T3 ($11) */
load_word(R_T3, R_TapePtr, O_(Binds_SetGteWorld,transform)),
add_ui_self( R_TapePtr, S_(Binds_SetGteWorld)),
/* Load 3x3 Rotation + 3x1 Translation from R_T3 into GTE CONTROL Regs (ctc2) */
load_word(R_T0, R_T3, 0), load_word(R_T1, R_T3, 4),
gte_mv_to_ctrl_r(R_T0, gte_cr_RT11), gte_mv_to_ctrl_r(R_T1, gte_cr_RT12),
load_word(R_T0, R_T3, 8), load_word(R_T1, R_T3, 12), load_word(R_T2, R_T3, 16),
gte_mv_to_ctrl_r(R_T0, gte_cr_RT13), gte_mv_to_ctrl_r(R_T1, gte_cr_RT21), gte_mv_to_ctrl_r(R_T2, gte_cr_RT22),
load_word(R_T0, R_T3, 20), load_word(R_T1, R_T3, 24), load_word(R_T2, R_T3, 28),
gte_mv_to_ctrl_r(R_T0, gte_cr_TRX), gte_mv_to_ctrl_r(R_T1, gte_cr_TRY), gte_mv_to_ctrl_r(R_T2, gte_cr_TRZ),
mac_yield()
};
#pragma endregion Baked Atoms
+213 -249
View File
@@ -1,3 +1,26 @@
/* ============================================================================
* duffle DSL Suffix Conventions
* ============================================================================
*
* Every mnemonic in this header follows the same suffix grammar:
*
* Primitive commands: gp0_cmd_poly_f3 = 0x20 (byte opcode)
* Packed 32-bit cmd: gp0_word_poly_f3(r, g, b) (32-bit, shifted)
*
* Type ordering: domain?_(direction)?_action_target_modifier_type?
* Examples: add_ui (add + unsigned + immediate)
* add_s (add + signed, R-type implicit)
* shift_lleft (shift + logical + left)
* shift_aright (shift + arithmetic + right)
* call_reg(rs) (call + register, $ra implicit)
* gte_mv_to_data_r (gte + mv + to + data + register)
* gte_lw_v0_xy(base) (gte + lw + v0 + xy)
* load_upper_i (load-upper + immediate, unique verb)
*
* Vendor mnemonics (gte_mtc2, gte_mfc2, gte_lwc2, gte_swc2, etc.) are NOT in this header.
* They are in the opt-in `gte_vendor_sym.h` for users who prefer the textbook MIPS assembly mnemonics.
* ============================================================================ */
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "dsl.h"
@@ -10,76 +33,27 @@
* gte.h — Geometry Transformation Engine (COP2) for the PS1
* ============================================================================
*
* Hand-rolled DSL for emitting GTE/MIPS instruction words as raw `.word`
* constants from C. No GCC inline-assembly string syntax in the code body.
*
* PHILOSOPHY
* ----------
* 1. A 32-bit instruction word is composed from per-field encoders. Each
* encoder knows only its own bit range; the composite ORs them together.
* No magic numbers inside any encoder body — every shift and mask is a
* named constant from the bitfield-layout enum below.
*
* 2. Pure (compile-time) instructions — every GTE *command* (RTPS, RTPT,
* NCLIP, MVMVA, …) and every COP2 *transfer* (ctc2/cfc2) with a constant
* rs/rt/rd — are emitted as a single integer constant via
* `asm_inline(...)` from gcc_asm.h. The C compiler constant-folds
* these into `.word` directives in .rodata.
*
* 3. Runtime-base-register instructions (lwc2, swc2, lw, sw, …) cannot be
* a pure compile-time word because the `rs` field is chosen by the
* compiler at codegen. For these we use a "placeholder-pun" pattern:
* a fixed register number (R_T4 = $12) is baked into the rs field of
* the `.word` constant, and the macro declares a `"r"(arg)` input
* constraint plus a clobber on the same register. The compiler is
* therefore *forced* to bind `arg` to that exact register, and the
* constant is correct.
*
* USAGE
* -----
* // Pure command sequence — all bits compile-time:
* asm volatile(
* asm_inline( gte_cmd_rtpt , gte_cmd_nclip , gte_cmd_avsz3 )
* asm_clobber( clb_system )
* );
*
* // Runtime-base-register load — caller picks the base GPR:
* register V3_S2* p_in_12 __asm__("$12") = verts[0].ptr;
* gte_load_v0(p_in_12, R_T4); // R_T4 = 12 = $t4 = $12
*
* // Three independent bases for an RTPT pipeline:
* register V3_S2* p0 __asm__("$12") = verts[0].ptr;
* register V3_S2* p1 __asm__("$13") = verts[1].ptr;
* register V3_S2* p2 __asm__("$14") = verts[2].ptr;
* gte_load_v0(p0, R_T4);
* gte_load_v1(p1, R_T5);
* gte_load_v2(p2, R_T6);
* gte_rtpt();
* Hand-rolled DSL for emitting GTE/MIPS instruction words as raw `.word` constants from C.
* No GCC inline-assembly string syntax in the code body.
*
* STYLE NOTES
* -----------
* - Per-field encoders are named `enc_gte_<field>(value)` and each one
* self-masks its argument before shifting. Mirrors the `enc_op / enc_rs
* / enc_rt / ...` family in mips.h.
* - The composite `enc_gte_cmdw(sf, mx, v, cv, lm, cmd)` is a flat OR of
* the per-field encoders, plus the COP2/CO base.
* - Pre-baked shortcuts (`gte_cmd_rtpt`, `gte_cmd_rtps`, …) are defined
* for the common cases so call sites read like assembly source.
* - All register/field values are enums (not `#define`s) so they show up
* in debugger symbol tables and IDE autocomplete.
* - Per-field encoders are named `enc_gte_<field>(value)` and each one self-masks its argument before shifting.
* Mirrors the `enc_op / enc_rs / enc_rt / ...` family in mips.h.
* - The composite `enc_gte_cmdw(sf, mx, v, cv, lm, cmd)` is a flat OR of the per-field encoders, plus the COP2/CO base.
* - Pre-baked shortcuts (`gte_cmd_rtpt`, `gte_cmd_rtps`, …) are defined for the common cases so call sites read like assembly source.
* - All register/field values are enums (not `#define`s) so they show up in debugger symbol tables and IDE autocomplete.
*
* SEE ALSO
* --------
* - gcc_asm.h: the `.word` emitter (`asm_inline`, `asm_clobber`, clobbers)
* - mips.h: the MIPS encoder layer this builds on
* - mips.h: The MIPS encoder layer this builds on.
*/
/* C2 data registers */
/* --- GTE Data Registers (Coprocessor 2) ---
* Preprocessor-visible integer ids for the COP2 data register file.
* Each enum value is bound to a parallel `_Code` `#define` so the
* preprocessor can stringify the integer (for `reg_str`/`rgcc` paths).
* Each enum value is bound to a parallel `_Code` `#define` so the preprocessor can stringify the integer (for `reg_str`/`rgcc` paths).
* Same pattern as the GPR `_Code` set in mips.h. */
#define C2_VXY0_Code 0
#define C2_VZ0_Code 1
@@ -198,8 +172,8 @@ enum {
* \_____ GTE_PAYLOAD _____/ \__ GTE_CMD __/
*
* Shifts/masks below are the *bit positions* and *bit widths* of each
* configurable field, used by the ENC_GTE_CMD encoder. Mirrors the
* OPCODE_SHIFT / RS_SHIFT convention used in mips.h.
* configurable field, used by the ENC_GTE_CMD encoder.
* Mirrors the OPCODE_SHIFT / RS_SHIFT convention used in mips.h.
*/
gte_shift_sf = 19, gte_width_sf = 1, gte_mask_sf = 0x1,
@@ -212,22 +186,20 @@ enum {
/* --- GTE Control Register Indices (for ctc2/cfc2) ---
* Preprocessor-visible integer ids for the COP2 control register file.
* Each enum value is bound to a parallel `_Code` `#define` so the
* preprocessor can stringify the integer (for `reg_str`/`rgcc` paths).
* Same pattern as the GPR `_Code` set in mips.h. Note: indices 21-23
* are reserved/unused on real hardware, so there's a gap. */
* Each enum value is bound to a parallel `_Code` `#define` so the preprocessor can stringify the integer (for `reg_str`/`rgcc` paths).
* Same pattern as the GPR `_Code` set in mips.h. Note: indices 21-23 are reserved/unused on real hardware, so there's a gap. */
#define gte_cr_RT11_Code 0
#define gte_cr_RT12_Code 1
#define gte_cr_RT13_Code 2
#define gte_cr_RT21_Code 3
#define gte_cr_RT22_Code 4
#define gte_cr_RT23_Code 5
#define gte_cr_RT31_Code 6
#define gte_cr_RT32_Code 7
#define gte_cr_RT33_Code 8
#define gte_cr_TRX_Code 9
#define gte_cr_TRY_Code 10
#define gte_cr_TRZ_Code 11
#define gte_cr_RT12_Code 1 /* packed with RT13 in bits 16..31 */
#define gte_cr_RT13_Code 2 /* packed with RT22 in bits 16..31 */
#define gte_cr_RT21_Code 3 /* packed with RT31 in bits 16..31 */
#define gte_cr_RT22_Code 4 /* RT33 alone (low 16 bits used) */
// #define gte_cr_RT23_Code 5
// #define gte_cr_RT31_Code 6
// #define gte_cr_RT32_Code 7
// #define gte_cr_RT33_Code 8
#define gte_cr_TRX_Code 5 /* PSX SDK convention: C2 r5 = TRX (alone, 32-bit) */
#define gte_cr_TRY_Code 6 /* PSX SDK convention: C2 r6 = TRY (alone, 32-bit) */
#define gte_cr_TRZ_Code 7 /* PSX SDK convention: C2 r7 = TRZ (alone, 32-bit) */
#define gte_cr_L11_Code 12
#define gte_cr_L12_Code 13
#define gte_cr_L13_Code 14
@@ -243,13 +215,14 @@ enum {
#define gte_cr_RFC_Code 27
#define gte_cr_GFC_Code 28
#define gte_cr_BFC_Code 29
#define gte_cr_OFX_Code 30
#define gte_cr_OFY_Code 31
#define gte_cr_OFX_Code 24
#define gte_cr_OFY_Code 25
#define gte_cr_H_Code 26
enum {
gte_cr_RT11 = gte_cr_RT11_Code, gte_cr_RT12 = gte_cr_RT12_Code, gte_cr_RT13 = gte_cr_RT13_Code,
gte_cr_RT21 = gte_cr_RT21_Code, gte_cr_RT22 = gte_cr_RT22_Code, gte_cr_RT23 = gte_cr_RT23_Code,
gte_cr_RT31 = gte_cr_RT31_Code, gte_cr_RT32 = gte_cr_RT32_Code, gte_cr_RT33 = gte_cr_RT33_Code,
gte_cr_RT21 = gte_cr_RT21_Code, gte_cr_RT22 = gte_cr_RT22_Code, //gte_cr_RT23 = gte_cr_RT23_Code,
// gte_cr_RT31 = gte_cr_RT31_Code, gte_cr_RT32 = gte_cr_RT32_Code, gte_cr_RT33 = gte_cr_RT33_Code,
gte_cr_TRX = gte_cr_TRX_Code, gte_cr_TRY = gte_cr_TRY_Code, gte_cr_TRZ = gte_cr_TRZ_Code,
gte_cr_L11 = gte_cr_L11_Code, gte_cr_L12 = gte_cr_L12_Code, gte_cr_L13 = gte_cr_L13_Code,
gte_cr_L21 = gte_cr_L21_Code, gte_cr_L22 = gte_cr_L22_Code, gte_cr_L23 = gte_cr_L23_Code,
@@ -260,26 +233,59 @@ enum {
};
enum { _C2_OPS_ = 0
, op_lwc2 = 0x32 /* Load Word to Coprocessor 2 (GTE) */
, op_swc2 = 0x3A /* Store Word from Coprocessor 2 (GTE) */
};
/* COP2 (GTE) Transfer Format: ctc2 rt, rd or cfc2 rt, rd
/* COP2 transfer sub-opcodes (5-bit field in the `rs` slot of enc_gte_tx).
*
* Spans the 2x2 {From, To} × {Data, Control} register classes that the GTE exposes:
* bit 1 (0x02): register class — 0 = data, 1 = control
* bit 2 (0x04): direction — 0 = read, 1 = write
*
* The values 0x00 (sub_mfc2) and 0x04 (sub_mtc2) are the same 5-bit numbers as the general MIPS `cop_mf` / `cop_mt` defined in mips.h
* (which target the data register file on any coprocessor).
* They are re-aliased here so the four-way table reads like the spec mnemonics (MFC2 / CFC2 / MTC2 / CTC2)
* and so the encoding lives next to its only consumer (this header).
*
* Vendor mnemonic aliases (gte_mfc2 / gte_mtc2 / gte_cfc2 / gte_ctc2) live in gte_vendor_sym.h. */
enum { _C2_TX_SUBS_ = 0
, sub_mfc2 = 0x00 /* MFC2: Move From Coprocessor 2 data reg */
, sub_cfc2 = 0x02 /* CFC2: Copy From Coprocessor 2 ctrl reg */
, sub_mtc2 = 0x04 /* MTC2: Move To Coprocessor 2 data reg */
, sub_ctc2 = 0x06 /* CTC2: Copy To Coprocessor 2 ctrl reg */
};
/* COP2 (GTE) Transfer Format: mfc2 / cfc2 / mtc2 / ctc2 rt, rd
* Layout: [op_cop2:6][sub:5][rt:5][rd:5][0:11]
* - sub: cop_mf (0x00) for cfc2, cop_mt (0x04) for ctc2
* - rt: GPR source/dest
* - rd: COP2 control register index (0..31) */
* - sub: one of sub_mfc2 / sub_cfc2 / sub_mtc2 / sub_ctc2
* - rt: GPR source/dest
* - rd: COP2 register index (0..31):
* data class → C2_VXY0_Code..C2_LZCR_Code (gte_in_v0_xy..gte_math_accum2 aliases)
* ctrl class → gte_cr_RT11_Code..gte_cr_OFY_Code */
#define enc_gte_tx(sub, rt, rd) (enc_op(op_cop2) | enc_rs(sub) | enc_rt(rt) | enc_rd(rd))
// #define gte_mt(rt, rd) enc_gte_tx(cop_mt, (rt), (rd)) /* Move GPR (rt) to GTE Control Register (rd) */
// #define gte_mf(rt, rd) enc_gte_tx(cop_mf, (rt), (rd)) /* Move GTE Control Register (rd) to GPR (rt) */
/* Explicit GTE Data vs Control Register Transfers */
#define gte_mf(rt, rd) enc_gte_tx(0x00, (rt), (rd)) /* Move from GTE Data Reg (e.g. MAC0, OTZ) */
#define gte_cf(rt, rd) enc_gte_tx(0x02, (rt), (rd)) /* Move from GTE Control Reg */
#define gte_mt(rt, rd) enc_gte_tx(0x04, (rt), (rd)) /* Move to GTE Data Reg (e.g. VXY0) */
#define gte_ct(rt, rd) enc_gte_tx(0x06, (rt), (rd)) /* Move to GTE Control Reg (e.g. Matrices) */
// #define gte_mv_to_data_r(rt, rd) enc_gte_tx(cop_mt, (rt), (rd)) /* Move GPR (rt) to GTE Control Register (rd) */
// #define gte_mv_from_data_r(rt, rd) enc_gte_tx(cop_mf, (rt), (rd)) /* Move GTE Control Register (rd) to GPR (rt) */
/* GTE Data vs Control Register Transfers
*
* Each macro emits a single .word constant for one of MFC2/CFC2/MTC2/CTC2.
*
* `rd` is the C2 register index in the file the sub-opcode names:
* gte_mv_from_data_r / gte_mv_to_data_r → C2 data register file
* gte_mv_from_ctrl_r / gte_mv_to_ctrl_r → C2 ctrl register file
*
* Common pairs:
* gte_mv_from_data_r(R_T0, C2_MAC0) — read MAC0 into a GPR
* gte_mv_to_data_r (R_V0, C2_VXY0) — write GPR into VXY0
* gte_mv_to_ctrl_r (R_T0, gte_cr_RT11) — write GPR into rotation matrix
* gte_mv_from_ctrl_r(R_T0, gte_cr_OFX) — read screen-X offset */
#define gte_mv_from_data_r(rt, rd) enc_gte_tx(sub_mfc2, (rt), (rd)) /* Move From data reg */
#define gte_mv_from_ctrl_r(rt, rd) enc_gte_tx(sub_cfc2, (rt), (rd)) /* Copy From ctrl reg */
#define gte_mv_to_data_r(rt, rd) enc_gte_tx(sub_mtc2, (rt), (rd)) /* Move To data reg */
#define gte_mv_to_ctrl_r(rt, rd) enc_gte_tx(sub_ctc2, (rt), (rd)) /* Copy To ctrl reg */
/* COP2 Data Load (lwc2): `lwc2 rt, off(rs)`
* Layout: [op_lwc2:6][rs:5][rt:5][imm:16]
@@ -296,22 +302,21 @@ enum { _C2_OPS_ = 0
* `swc2` is redundant when we're already inside the `gte_` namespace.
* gte_lw rt, base, off → lwc2 rt, off(base)
* gte_sw rt, base, off → swc2 rt, off(base)
* For the typical user-facing vector-level load (xy + z as two
* instructions), use the higher-level `gte_load_vN` macros below. */
* For the typical user-facing vector-level load (xy + z as two instructions),
* use the higher-level `gte_load_vN` macros below. */
#define gte_lw(rt, base, off) enc_gte_lw(rt, base, off)
#define gte_sw(rt, base, off) enc_gte_sw(rt, base, off)
/* GTE Command Format (The math engine trigger)
/* GTE Command Format
* Opcode is always MIPS_OP_COP2, RS is always 1 (CO).
* The lower 25 bits are the GTE-specific command payload.
*
* The granular `enc_gte_<field>(x)` macros below mirror the `enc_op`/`enc_rs`
* pattern in mips.h: each one self-masks and shifts its own field, so a
* caller can build up a GTE command piece by piece (handy for state-driven
* MVMVA emitters that vary one field at a time).
* The granular `enc_gte_<field>(x)` macros below mirror the `enc_op`/`enc_rs` pattern in mips.h:
* Each one self-masks and shifts its own field, so a caller can build up a GTE command piece by piece
* (handy for state-driven MVMVA emitters that vary one field at a time).
*
* `ENC_GTE_CMD` is the all-in-one convenience for emitting a full command
* word in one go. It just ORs the per-field encoders together. */
* `ENC_GTE_CMD` is the all-in-one convenience for emitting a full command word in one go.
* It just ORs the per-field encoders together. */
#define gte_cmd_base (enc_op(op_cop2) | (1 << 25))
/* Per-field encoders. Each one does (value & mask) << shift on its own. */
@@ -335,41 +340,35 @@ enum { _C2_OPS_ = 0
/* GTE command words for the common cases.
*
* These are pure compile-time integer constants — the C compiler
* constant-folds them into `.word` directives in .rodata. Use them
* inside `asm_inline(...)` blocks (see `gte_rtpt` below for the
* canonical idiom).
* These are pure compile-time integer constants — the C compiler constant-folds them into `.word` directives in .rodata.
* Use them inside `asm_inline(...)` blocks (see `gte_rtpt` below for the idiom).
*
* Decomposition (per the `enc_gte_<field>` definitions above):
* gte_cmdw_<name> = gte_cmd_base | enc_gte_cmd(<cmd>)
* The SF/MX/V/CV/LM fields are all zero in the common cases (standard
* rotation-matrix, no scaling factor, V0 vector, translation vector,
* no clamp), so the only varying bits are the `cmd` field.
* The SF / MX / V / CV / LM fields are all zero in the common cases
* (standard rotation-matrix, no scaling factor, V0 vector, translation vector, no clamp),
* so the only varying bits are the `cmd` field.
*
* Naming follows the file's convention: `gte_cmd_*` is the raw
* 6-bit `cmd` field id, `gte_cmdw_*` is the fully-encoded 32-bit
* instruction word ready to drop into a `.word` directive.
* Naming convention:
* - `gte_cmd_*` : Raw 6-bit `cmd` field id
* - `gte_cmdw_* : 32-bit instruction word ready to drop into a `.word` directive.
*
* --------------------------------------------------------------------------
* PsyQ-compatibility note (RTPS/RTPT):
* The original Sony PsyQ `inline_n.h` ships RTPT as `cop2 0x0280030` and
* RTPS as `cop2 0x0180001`. Both have `0x20` set in the upper-reserved
* region (bit 21) AND `sf=1` (bit 19) — i.e. the "no division" flag.
* Per psx-spec these bits are reserved/must-be-zero, but the real GTE
* hardware and PCSX-Redux's GTE model both IGNORE them on these two
* commands (the perspective divide happens regardless of `sf`).
* The original Sony PsyQ `inline_n.h` ships RTPT as `cop2 0x0280030` and RTPS as `cop2 0x0180001`.
* Both have `0x20` set in the upper-reserved region (bit 21) AND `sf=1` (bit 19) — i.e. the "no division" flag.
* Per psx-spec these bits are reserved/must-be-zero,
* but the real GTE hardware and PCSX-Redux's GTE model both IGNORE them on these two commands
* (the perspective divide happens regardless of `sf`).
*
* If we emit a strictly-spec-compliant word (`sf=0`, reserved bits
* clear), PCSX-Redux's GTE checks those bits more strictly than the
* silicon does and RTPT silently no-ops — the floor's screen
* coordinates come out as raw projection-of-rotation (Z never
* divided), `nclip` ends up wrong, and the triangle is culled.
* If we emit a strictly-spec-compliant word (`sf=0`, reserved bits clear),
* PCSX-Redux's GTE checks those bits more strictly than the silicon does and RTPT silently no-ops —
* the floor's screen coordinates come out as raw projection-of-rotation (Z never divided),
* `nclip` ends up wrong, and the triangle is culled.
*
* So for RTPS and RTPT we OR-in the `0x28` "PsyQ compat" pattern to
* match the working bit pattern everyone has shipped for 25 years.
* NCLIP/OP/MVMVA stay spec-clean — their reserved bits really are
* zero in the original PsyQ source.
* So for RTPS and RTPT we OR-in the `0x28` "PsyQ compat" pattern to match the working bit pattern everyone has shipped for 25 years.
* NCLIP / OP / MVMVA stay spec-clean — their reserved bits really are zero in the original PsyQ source.
* --------------------------------------------------------------------------
*/
#define gte_cmdw_psyq_compat (1u << 21 | enc_gte_sf(gte_sf_integer))
@@ -378,8 +377,11 @@ enum { _C2_OPS_ = 0
#define gte_cmdw_rtpt (gte_cmd_base | enc_gte_cmd(gte_cmd_rtpt ) | gte_cmdw_psyq_compat)
#define gte_cmdw_nclip (gte_cmd_base | enc_gte_cmd(gte_cmd_nclip))
#define gte_cmdw_op (gte_cmd_base | enc_gte_cmd(gte_cmd_op ))
#define gte_cmdw_outer_product gte_cmdw_op /* "outer product" -- NOCASH/Sdk terminology */
#define gte_cmdw_wedge gte_cmdw_op /* "wedge product" -- geometric-algebra terminology */
#define gte_cmdw_mvmva (gte_cmd_base | enc_gte_cmd(gte_cmd_mvmva))
#define gte_cmdw_rotate_translate_perspective_single gte_cmdw_rtps
#define gte_cmdw_rotate_translate_perspective_triple gte_cmdw_rtpt
/* PsyQ compatibility bits for AVSZ3 (Bits 20, 22, 24 must be set) */
@@ -395,67 +397,58 @@ enum { _C2_OPS_ = 0
#define gte_cmd_avsz4 0x2E
#define gte_cmdw_avsz4 (gte_cmd_base | enc_gte_cmd(gte_cmd_avsz4) | gte_cmdw_psyq_avsz3_compat)
#define gte_cmdw_avg_sort_z4 gte_cmdw_avsz4
/**
* @brief Loads a single SVECTOR to GTE vector register V0
*
* @details Loads values from an SVECTOR struct to GTE data registers C2_VXY0
* (XY at offset 0) and C2_VZ0 (Z at offset 4) using `lwc2`.
*
* Uses string-style GCC inline asm with `%0` substitution because the
* base register `r0` is a runtime GPR chosen by the compiler — it cannot
* be encoded into a static `.word` constant.
* Uses string-style GCC inline asm with `%0` substitution because the base register `r0` is a runtime GPR chosen by the compiler.
* It cannot be encoded into a static `.word` constant.
*
* Usage:
* asm_gte_load_v0(svector_ptr);
* Usage: asm_gte_load_v0(svector_ptr);
*/
/* lwc2 encoding helpers parameterized on the base GPR.
*
* gte_lwc2_v0(base) → lwc2 $0, 0(base) ; C2_VXY0
* gte_lwc2_v0z(base) → lwc2 $1, 4(base) ; C2_VZ0
* gte_lwc2_v1(base) → lwc2 $2, 0(base) ; C2_VXY1
* gte_lwc2_v1z(base) → lwc2 $3, 4(base) ; C2_VZ1
* gte_lwc2_v2(base) → lwc2 $4, 0(base) ; C2_VXY2
* gte_lwc2_v2z(base) → lwc2 $5, 4(base) ; C2_VZ2
* gte_lw_v0_xy(base) → lwc2 $0, 0(base) ; C2_VXY0
* gte_lw_v0_z(base) → lwc2 $1, 4(base) ; C2_VZ0
* gte_lw_v1_xy(base) → lwc2 $2, 0(base) ; C2_VXY1
* gte_lw_v1_z(base) → lwc2 $3, 4(base) ; C2_VZ1
* gte_lw_v2_xy(base) → lwc2 $4, 0(base) ; C2_VXY2
* gte_lw_v2_z(base) → lwc2 $5, 4(base) ; C2_VZ2
*
* `base` is the GPR number to bake into the .word constant's `rs` field.
* These are pure compile-time integers; the C compiler constant-folds
* them into .word directives. */
* These are pure compile-time integers; the C compiler constant-folds them into .word directives. */
enum {
GTE_Z_Offset = 4
};
#define gte_lw_v0(base) enc_gte_lw(gte_in_v0_xy, (base), 0)
#define gte_lw_v0z(base) enc_gte_lw(gte_in_v0_z, (base), GTE_Z_Offset)
#define gte_lw_v1(base) enc_gte_lw(gte_in_v1_xy, (base), 0)
#define gte_lw_v1z(base) enc_gte_lw(gte_in_v1_z, (base), GTE_Z_Offset)
#define gte_lw_v2(base) enc_gte_lw(gte_in_v2_xy, (base), 0)
#define gte_lw_v2z(base) enc_gte_lw(gte_in_v2_z, (base), GTE_Z_Offset)
#define gte_lw_v0_xy(base) enc_gte_lw(gte_in_v0_xy, (base), 0)
#define gte_lw_v0_z(base) enc_gte_lw(gte_in_v0_z, (base), GTE_Z_Offset)
#define gte_lw_v1_xy(base) enc_gte_lw(gte_in_v1_xy, (base), 0)
#define gte_lw_v1_z(base) enc_gte_lw(gte_in_v1_z, (base), GTE_Z_Offset)
#define gte_lw_v2_xy(base) enc_gte_lw(gte_in_v2_xy, (base), 0)
#define gte_lw_v2_z(base) enc_gte_lw(gte_in_v2_z, (base), GTE_Z_Offset)
/* gte_load_vN(r_ptr, base) — placeholder-punned lwc2 loaders
*
* Emits `.word` constants encoding `lwc2 $N, off(<base>)` for the chosen
* GTE vector register, where `<base>` is the GPR number you pass in
* Emits `.word` constants encoding `lwc2 $N, off(<base>)` for the chosen GTE vector register, where `<base>` is the GPR number you pass in
* (typically one of R_T4..R_T9 for the standard "3-pointer" pattern).
*
* The caller MUST bind `r_ptr` to that same GPR via a register variable:
*
* register V3_S2* p_in_12 __asm__("$12") = my_ptr;
* gte_load_v0(p_in_12, R_T4); // R_T4 = 12, base is $12
*
* Then `"r"(r_ptr)` inside the asm binds to $12 (the only register
* `p_in_12` can live in), which is exactly the register the .word
* constants expect. A `"$12"` clobber would conflict with the
* register-variable binding ("asm specifier for variable conflicts
* with asm clobber list"), so we omit it. The other ABI-clobbers
* ($2/$8/$9/$31) stay because the GTE instructions don't touch
* caller-saved GPRs but the kernel does treat them as volatile.
* Then `"r"(r_ptr)` inside the asm binds to $12 (the only register `p_in_12` can live in),
* which is exactly the register the .word constants expect.
* A `"$12"` clobber would conflict with the register-variable binding ("asm specifier for variable conflicts with asm clobber list"), so we omit it.
* The other ABI-clobbers ($2/$8/$9/$31) stay because the GTE instructions don't touch caller-saved GPRs but the kernel does treat them as volatile.
*
* WHICH REGISTER TO PICK
* ----------------------
* Any caller-saved GPR is safe. Recommended default for an RTPT-style
* 3-pointer pipeline:
* Any caller-saved GPR is safe. Recommended default for an RTPT-style 3-pointer pipeline:
* gte_load_v0(p0, R_T4); // $12
* gte_load_v1(p1, R_T5); // $13
* gte_load_v2(p2, R_T6); // $14
@@ -468,32 +461,29 @@ enum {
* clobbers section : "$2", "$8", ..., "memory" (from asm_clobber)
* 3 colons total, GCC-legal. No string-syntax mnemonics in the .word body.
*
* The `asm_clobber(...)` helper from gcc_asm.h prepends the colon that
* starts the clobbers section. */
#define gte_load_v0(r_ptr, base) asm volatile( \
asm_words( gte_lw_v0(base), gte_lw_v0z(base) ) \
asm_rpins, r_use(r_ptr) \
* The `asm_clobber(...)` helper from gcc_asm.h prepends the colon that starts the clobbers section. */
#define gte_load_v0(r_ptr, base) asm volatile( \
asm_words( gte_lw_v0_xy(base), gte_lw_v0_z(base) ) \
asm_rpins, r_use(r_ptr) \
asm_clobber: rlit(R_V0), rlit(R_T0), rlit(R_T1), rlit(R_RA), clb_mem_drain \
)
#define gte_load_v1(r_ptr, base) asm volatile( \
asm_words( gte_lw_v1(base), gte_lw_v1z(base) ) \
asm_rpins, r_use(r_ptr) \
#define gte_load_v1(r_ptr, base) asm volatile( \
asm_words( gte_lw_v1_xy(base), gte_lw_v1_z(base) ) \
asm_rpins, r_use(r_ptr) \
asm_clobber: rlit(R_V0), rlit(R_T0), rlit(R_T1), rlit(R_RA), clb_mem_drain \
)
#define gte_load_v2(r_ptr, base) asm volatile( \
asm_words( gte_lw_v2(base), gte_lw_v2z(base) ) \
asm_rpins, r_use(r_ptr) \
#define gte_load_v2(r_ptr, base) asm volatile( \
asm_words( gte_lw_v2_xy(base), gte_lw_v2_z(base) ) \
asm_rpins, r_use(r_ptr) \
asm_clobber: rlit(R_V0), rlit(R_T0), rlit(R_T1), rlit(R_RA), clb_mem_drain \
)
/* gte_load_v0v1v2(p0, p1, p2, b0, b1, b2) — the canonical prelude to gte_cmd_rtpt.
*
* Loads all three GTE input vectors (6 words) from three separate pointers,
* one per GTE vector register, each loaded from its own base GPR. Caller
* must bind each `pN` to `bN` via a register variable.
/* gte_load_v0v1v2(p0, p1, p2, b0, b1, b2) — prelude to gte_cmd_rtpt.
*
* Loads all three GTE input vectors (6 words) from three separate pointers, one per GTE vector register,
* each loaded from its own base GPR. Caller must bind each `pN` to `bN` via a register variable.
* register V3_S2* p0 rgcc(R_T4) = verts[0].ptr; // → __asm__("$12")
* register V3_S2* p1 rgcc(R_T5) = verts[1].ptr; // → __asm__("$13")
* register V3_S2* p2 rgcc(R_T6) = verts[2].ptr; // → __asm__("$14")
@@ -502,9 +492,9 @@ enum {
*/
#define gte_load_v0v1v2(p0, p1, p2, b0, b1, b2) asm volatile( \
asm_words( \
gte_lw_v0(b0), gte_lw_v0z(b0), \
gte_lw_v1(b1), gte_lw_v1z(b1), \
gte_lw_v2(b2), gte_lw_v2z(b2) ) \
gte_lw_v0_xy(b0), gte_lw_v0_z(b0), \
gte_lw_v1_xy(b1), gte_lw_v1_z(b1), \
gte_lw_v2_xy(b2), gte_lw_v2_z(b2) ) \
asm_rpins \
, r_use(p0), r_use(p1), r_use(p2) \
asm_clobber: rlit(R_V0), rlit(R_T0), rlit(R_T1), rlit(R_RA), clb_mem_drain \
@@ -512,37 +502,28 @@ enum {
/**
* @brief Rotate, Translate and Perspective Triple (23 cycles)
*
* @details Performs rotation, translation and perspective calculation of three
* vertices at once. The equation performed is the same as gte_rtps() only
* repeated three times for each vertex. The result of the first vertex is
* stored in GTE data register C2_SXY0, the second vector in C2_SXY1 then
* C2_SXY2.
* @details Performs rotation, translation and perspective calculation of three vertices at once.
* The equation performed is the same as gte_rtps() only repeated three times for each vertex.
* The result of the first vertex is stored in GTE data register C2_SXY0, the second vector in C2_SXY1 then C2_SXY2.
*
* Encoder-style emission (no inline-asm strings in the code body):
* 1. Two `nop` words fill the COP2 pipeline latency — the GTE
* takes ~8 cycles per perspective divide, and the nops let any
* preceding lwc2/swc2 retire before RTPT starts reading its
* inputs from V0/V1/V2.
* 2. The RTPT command word itself is `gte_cmdw_rtpt` (see the
* pre-baked encoders above) — `0x0280030` decoded as
* `op_cop2` | CO(1) | cmd=RTPT, with all SF/MX/V/CV/LM fields
* zero (standard rotation, no scaling, V0 vector, translation
* vector, no clamp).
* 1. Two `nop` words fill the COP2 pipeline latency — the GTE takes ~8 cycles per perspective divide,
* and the nops let any preceding lwc2/swc2 retire before RTPT starts reading its inputs from V0/V1/V2.
* 2. The RTPT command word itself is `gte_cmdw_rtpt` (see the pre-baked encoders above) —
* `0x0280030` decoded as `op_cop2` | CO(1) | cmd=RTPT, with all SF/MX/V/CV/LM fields zero
* (standard rotation, no scaling, V0 vector, translation vector, no clamp).
*
* Clobbers the caller-saved GPRs via `clb_system` (per the kernel
* ABI) plus the standard "memory" barrier. Does not clobber any COP2
* data/control register — those have to be saved by the caller if
* they need to survive across the call (RTPT writes SXY0..2, SZ0..3,
* OTZ, MAC0..3, IR0..3, etc.).
* Clobbers the caller-saved GPRs via `clbr_volatile_gprs` (per the kernel ABI)
* plus the standard "memory" barrier. Does not clobber any COP2 data/control register —
* those have to be saved by the caller if they need to survive across the call (RTPT writes SXY0..2, SZ0..3, OTZ, MAC0..3, IR0..3, etc.).
*/
#define gte_rtpt() \
asm volatile( \
asm_words( nop, nop, gte_cmdw_rtpt ) \
asm_clobber: clb_system \
asm_clobber: clbr_volatile_gprs \
)
#define gte_rtpt_ori() \
#define gte_rtpt_asm_str() \
__asm__ volatile( \
"nop;" \
"nop;" \
@@ -550,37 +531,29 @@ enum {
/**
* @brief Normal clipping (8 cycles)
*
* @details Computes the sign of three screen coordinates (C2_SXY0-2) used for
* backface culling. If the value of C2_MAC0 is negative, the coordinates are
* inverted and thus the triangle is back facing.
* @details Computes the sign of three screen coordinates (C2_SXY0-2) used for backface culling.
* If the value of C2_MAC0 is negative, the coordinates are inverted and thus the triangle is back facing.
*
* The following equation is performed when executing this GTE command:
*
* MAC0 = SX0*SY1 + SX1*SY2 + SX2*SY0 - SX0*SY2 - SX1*SY0 - SX2*SY1
*
* Encoder-style emission (no inline-asm strings in the code body):
* 1. Two `nop` words fill the COP2 pipeline latency - the GTE
* pipeline takes a few cycles per op, and the nops let any
* preceding lwc2/swc2/RTPT retire before NCLIP starts reading
* its inputs from SXY0/SXY1/SXY2.
* 2. The NCLIP command word itself is `gte_cmdw_nclip` (see the
* pre-baked encoders above) - `0x01400006` decoded as
* `op_cop2` | CO(1) | cmd=NCLIP, with all SF/MX/V/CV/LM fields
* zero. NCLIP is spec-clean in the original PsyQ source
* (unlike RTPS/RTPT which carry the `gte_cmdw_psyq_compat`
* quirk), so `gte_cmdw_nclip` does NOT OR in any reserved bits.
* 1. Two `nop` words fill the COP2 pipeline latency
* - the GTE pipeline takes a few cycles per op, and the nops let any preceding
* lwc2/swc2/RTPT retire before NCLIP starts reading its inputs from SXY0/SXY1/SXY2.
* 2. The NCLIP command word itself is `gte_cmdw_nclip` (see the pre-baked encoders above)
* - `0x01400006` decoded as `op_cop2` | CO(1) | cmd=NCLIP, with all SF/MX/V/CV/LM fields zero.
* NCLIP is spec-clean in the original PsyQ source (unlike RTPS/RTPT which carry the `gte_cmdw_psyq_compat` quirk),
* so `gte_cmdw_nclip` does NOT OR in any reserved bits.
*
* Clobbers the caller-saved GPRs via `clb_system` (per the kernel
* ABI) plus the standard "memory" barrier. Does not clobber any COP2
* data/control register - those have to be saved by the caller if
* they need to survive across the call (NCLIP writes MAC0 only; it
* is purely a sign-of-double-product computation on SXY0..2).
* Clobbers the caller-saved GPRs via `clbr_volatile_gprs` (per the kernel ABI) plus the standard "memory" barrier.
* Does not clobber any COP2 data/control register.
* Those have to be saved by the caller if they need to survive across the call (NCLIP writes MAC0 only;
* it is purely a sign-of-double-product computation on SXY0..2).
*/
#define gte_nclip() \
asm volatile( \
asm_words( nop, nop, gte_cmdw_nclip ) \
asm_clobber: clb_system \
asm_clobber: clbr_volatile_gprs \
)
#define gte_stotz(r0) __asm__ volatile("swc2 $7, 0( %0 )" : : "r"(r0) : "memory")
@@ -601,17 +574,13 @@ enum {
"cop2 0x0158002D;")
/* asm_gte_matrix_set_rotation(r0)
* Loads the 3x3 rotation matrix at `r0` into the GTE's rotation-matrix control registers (RT11..RT22, indices 0..4) via ctc2.
*
* Loads the 3x3 rotation matrix at `r0` into the GTE's rotation-matrix
* control registers (RT11..RT22, indices 0..4) via ctc2.
*
* Memory layout at r0: five contiguous 32-bit words (offsets 0..16),
* each holding two packed 16-bit matrix elements. The first 1.5 rows
* of a standard PSX SDK MATRIX struct (where each row is laid out as
* Memory layout at r0: five contiguous 32-bit words (offsets 0..16), each holding two packed 16-bit matrix elements.
* The first 1.5 rows of a standard PSX SDK MATRIX struct (where each row is laid out as
* [RT_xx, RT_xy] | [RT_xz, pad] | ...).
*
* Generated MIPS (mirrors the source macro):
*
* lw $12, 0( %0 ) ; word 0
* lw $13, 4( %0 ) ; word 1
* ctc2 $12, $0 ; → C2_RT11
@@ -623,44 +592,39 @@ enum {
* ctc2 $13, $3 ; → C2_RT21
* ctc2 $14, $4 ; → C2_RT22
*
* Same contract as gte_load_v0: caller MUST bind `r0` to $12 via a
* register variable (`rgcc(R_T4)`) for the `lw $12, off(...)`
* instructions to read from the right base. The `"r"(r0)` constraint
* alone doesn't force a specific GPR — it just lets GCC pick one.
* The .word constants here bake R_T4/R_T5/R_T6 into the `rs` field
* of each lw, so the lw instructions will only do the right thing
* if $12/$13/$14 hold the matrix base at runtime.
* Same contract as gte_load_v0: caller MUST bind `r0` to $12 via a register variable (`rgcc(R_T4)`) for the `lw $12, off(...)`
* instructions to read from the right base. The `"r"(r0)` constraint alone doesn't force a specific GPR — it just lets GCC pick one.
* The .word constants here bake R_T4/R_T5/R_T6 into the `rs` field of each lw, so the lw instructions will
* only do the right thing if $12 / $13 / $14 hold the matrix base at runtime.
*
* M3_S2* m = ...;
* register M3_S2* m_in_12 rgcc(R_T4) = m;
* asm_gte_matrix_set_rotation(m_in_12);
*
* We clobber $12/$13/$14 (the ones we use as scratch inside the
* inline asm) plus the system clobbers; we don't clobber `r0` because
* the `rgcc` binding already says "this variable lives in $12".
* We clobber $12/$13/$14 (the ones we use as scratch inside the inline asm)
* plus the system clobbers; we don't clobber `r0` because the `rgcc` binding already says "this variable lives in $12".
*
* WARNING: Incomplete by design. The source macro only writes RT11..RT22
* (5 of 9 rotation elements); RT23 and the entire RT3x row are left
* untouched. Real libpsn00b SetRotMatrix writes all 9. Use only when the
* GTE's remaining rotation entries are already correct, or you will
* get stale-RT2x/RT3x artifacts in RTPS/RTPT/MVMVA output.
* WARNING: Incomplete by design. The source macro only writes RT11..RT22 (5 of 9 rotation elements);
* RT23 and the entire RT3x row are left untouched.
* Real libpsn00b SetRotMatrix writes all 9. Use only when the GTE's remaining rotation entries are already correct,
* or you will get stale-RT2x/RT3x artifacts in RTPS/RTPT/MVMVA output.
*/
#define asm_gte_matrix_set_rotation(r0) \
asm volatile( \
asm_words( \
load_word(R_T5, R_T4, 0) \
, load_word(R_T6, R_T4, 4) \
, gte_mt( R_T5, 0) \
, gte_mt( R_T6, 1) \
, gte_mv_to_data_r( R_T5, 0) \
, gte_mv_to_data_r( R_T6, 1) \
, load_word(R_T5, R_T4, 8) \
, load_word(R_T6, R_T4, 12) \
, load_word(R_T4, R_T4, 16) \
, gte_mt( R_T5, 2) \
, gte_mt( R_T6, 3) \
, gte_mt( R_T4, 4) \
, gte_mv_to_data_r( R_T5, 2) \
, gte_mv_to_data_r( R_T6, 3) \
, gte_mv_to_data_r( R_T4, 4) \
) \
, r_use(r0) \
asm_clobber: clb_system, rlit(R_T4), rlit(R_T5), rlit(R_T6) \
asm_clobber: clbr_volatile_gprs, rlit(R_T4), rlit(R_T5), rlit(R_T6) \
)
#pragma endregion ASM DSL
+42
View File
@@ -0,0 +1,42 @@
/* ============================================================================
* duffle DSL — GTE Vendor Mnemonics (opt-in)
* ============================================================================
*
* Provides the textbook MIPS assembly mnemonics for the GTE/COP2 instructions as thin aliases to the duffle macros in gte.h.
* The duffle names are primary; this header is for users who prefer the textbook mnemonics.
*
* USAGE: #include "duffle/gte_vendor_sym.h" // after gte.h
*
* Mapping (vendor -> duffle):
* Transfers (move GPR <-> GTE control/data register):
* gte_mfc2 -> gte_mv_from_data_r (move from coprocessor 2 data reg)
* gte_mtc2 -> gte_mv_to_data_r (move to coprocessor 2 data reg)
* gte_cfc2 -> gte_mv_from_ctrl_r (move from coprocessor 2 control reg)
* gte_ctc2 -> gte_mv_to_ctrl_r (move to coprocessor 2 control reg)
*
* Data load/store (load/store word to coprocessor 2 data register):
* gte_lwc2(rt, base, off) -> gte_lw(rt, base, off)
* gte_swc2(rt, base, off) -> gte_sw(rt, base, off)
* (the lower-level vector variants gte_lw_v0_xy etc. don't have
* vendor mnemonics; they're already gte_-prefixed and short)
* ============================================================================ */
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "gte.h"
#endif
#ifndef DUFFLE_GTE_VENDOR_SYM_H
#define DUFFLE_GTE_VENDOR_SYM_H
/* Transfers (move GPR <-> GTE control/data register) */
#define gte_mfc2(rt, rd) gte_mv_from_data_r((rt), (rd))
#define gte_mtc2(rt, rd) gte_mv_to_data_r((rt), (rd))
#define gte_cfc2(rt, rd) gte_mv_from_ctrl_r((rt), (rd))
#define gte_ctc2(rt, rd) gte_mv_to_ctrl_r((rt), (rd))
/* Data load/store (load/store word to coprocessor 2 data register) */
#define gte_lwc2(rt, base, off) gte_lw((rt), (base), (off))
#define gte_swc2(rt, base, off) gte_sw((rt), (base), (off))
#endif
+195 -211
View File
@@ -1,82 +1,208 @@
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "gen/macs.h"
# include "gen/offsets.h"
# include "dsl.h"
# include "gcc_asm.h"
# include "mips.h"
# include "gte.h"
# include "memory.h"
# include "atom_dsl.h"
# include "dsl.atom.h"
#endif
typedef U4 const MipsCode;
#define MipsAtom_(sym) MipsCode tmpl(code,sym) [] align_(4) =
#pragma region Tape Drive
/* ---------------------------------------------------------------------------
* TAPE DRIVE ABI & REGISTER ALIASES
* ---------------------------------------------------------------------------
* We map the MIPS temporary registers to a persistent global workspace.
* The C compiler is completely unaware of these bindings.
* ---------------------------------------------------------------------------*/
/* -----------------------------------------------------------------------------
* TAPE DRIVE ABI
* -----------------------------------------------------------------------------
* Note(Ed): One of the main purposes of this codebase is to help me
* learn this, as such the information below may be entirely realized
* or finalized conceptually.
* -----------------------------------------------------------------------------
* This ABI and its associated legos were directly inspired by researching
* the work of Timothy Lottes and Onat Türkçüoğlu; along with many others.
* It's the simplest bootstrap of a a directly executed chain of assemby
* arrays (Atoms) that terminate with a yield sequence to the next atom.
* These eventually lead to a terminal atom for the tape which is defined
* below as "tape_exit".
*
* This behaves as one of the simplest runtime harnesses ontop of a
* host-enviornment's execution engine to author and compose programs with.
* From here various conventions can be further applied.
* To make things easier to understand it may be better to focus on what this
* ABI does not have. It does not have have any branching within the tape but
* relative branches between atoms. Branching nearly is always downstream.
* Stack usage is non-existent. Push/Pop, FIFO, or Arena/Bump data structures
* are used by atoms explicitly. In it's current form withe C11 macro dsl,
* the user also has to do manual register allocation per atom.
*
* One of the remarkable things about utilizing this abi is its essentially
* interopable with CPUs, GPUs, FPGA, or, basically anything
* from the 5th generation consoles and onward.
* The ABI directly reflects how all computational hardware must be architected
* in order to execute digital logic effectively on current era tech.
* On the PS1 we don't have access to a few features like multi-threading,
* speculative execution, or L3 cache; but, we can set the foundation for legoing
* whats required baseline wise for eventually expanding the harness and core atoms
* to take those newer hardware features into account. For example, you can easily
* expand this to support wave-based execution model on a PS2 or PS3.
* Not having a stack or automatic register allocation means the user can't ignore
* excessive argument shuffle across workload or waves and thier phases.
* Crossing ABI boundaries to other runtimes that do has an obviouss penalties.
*
* Learning data-oreinted code becomes a natural progression. Your not fighting
* a stack-based procedural paradigm that wants to argument shuffle on the stack
* by lack of constraints on how the user may "call" a procedure. The user doesn't
* have to hammer down "rules" or patterns to know how to massage the compiler
* to get the asesmbly into its natural form. The form is obvious, and once
* the user gets to author their compoonents it becomes a game of tetris.
*
* Another feature is this ABI is very compatible with bootstrapping and developing
* simple toolchains built off of bit-packed annotated command streams the user can
* directly author, maintatain, and immediately execute. That being a color forth.
* This can make the tetris less of a chore with some helpful policy generation for
* allocation of registers, helping to choose resuable components, designing DSL on
* the fly, etc.
* -----------------------------------------------------------------------------
* TODO(Ed): We ned pretty ascii diagrams and proper guides, articles, etc.
* -----------------------------------------------------------------------------
* For now this thing is just functioning and I'm abusing C11 + a lua metaprogram
* to help establish a hybrid toolchain to ideate on a traditional text-based
* authoring UX for this paradigm.
* If pcsx-redux gets me viable hot-reload and persistent data storage beyond
* save-states (just copying ram to filesystem). I can author a color forth to
* mess around with, with an editor in-emulator or on the actual machine itself.
* Assembly is tedius, but I think this codebase most likely has some of the most,
* ergonomic you can come across..
* */
/* Register Allocation Info */
enum {
R_AtomJmp = R_T9,
R_TapePtr = R_T8, /* The Instruction Stream Pointer */
R_InCursor = R_T4, /* Input data cursor */
R_AtomJmp = R_T8 atom_reg, /* debug-visible; tape yield handshake scratch */
R_TapePtr = R_T9 atom_reg, /* The Instruction Stream Pointer */
/* Stringification codes for the GCC inline assembler clobber lists. */
#define R_AtomJmp_Code R_T8_Code
#define R_TapePtr_Code R_T9_Code
R_PrimCursor = R_T7, /* VRAM output cursor (primitive buffer) */
R_FaceCursor = R_T4, /* Input data cursor (indices/faces) */
R_VertBase = R_T5, /* Base address of the vertex array */
R_OtBase = R_T6, /* Base address of the Ordering Table */
// R_InCursor = R_T4,
// #define R_InCursor_Code R_T4_Code
/* Stringification codes for the GCC inline assembler clobber lists */
#define R_TapePtr_Code R_T8_Code
#define R_InCursor_Code R_T4_Code
// Reserved Registers (Callee-saved):
// - R_T9: Holds the Tape Ptr which we need to increment
// If we hit a wall with register allocations we can clobber V0 & V1 (return values), defering as opt-in by user.
// - R_RA: Not sure??
// Needed by ac_yield but can be used as atom scratch:
// - R_T8: Will be used as the atom jump register.
#define R_PrimCursor_Code R_T7_Code
#define R_FaceCursor_Code R_T4_Code
#define R_VertBase_Code R_T5_Code
#define R_OtBase_Code R_T6_Code
// All allocatable registers for mips atoms:
R_TScratchVolatile = R_AT, // This one is reserved for psuedo instructions, but you can technically use it.
R_TScratch0 = R_T0,
R_TScratch1 = R_T1,
R_TScratch2 = R_T2,
R_TScratch3 = R_T3,
R_TScratch4 = R_T4,
R_TScratch5 = R_T5,
R_TScratch6 = R_T6,
R_TScratch7 = R_T7,
R_TScratch8 = R_T8,
R_TScratch10 = R_V0, // Tend to be used with gte DMAs
R_TScratch11 = R_V1, // Tend to be used with gte DMAs
// Note(Ed): We can technically clobber these, but don't unless we hit a bottleneck.
// A 0-2
// S 0-7
};
/* The 'Exit' Atom */
MipsAtom_(tape_exit) { jump_reg(rret_addr), nop };
typedef U4 const MipsCode; // Underlying type to mips asm words.
typedef Slice_(MipsCode);
/* Generalized Tape Engine Runner */
FI_ void tape_run(Slice_U4 tape) { register U4* tp rgcc(R_TapePtr) = tape.ptr; asm volatile(
typedef U4 const MipsAtom; // Underlying type to an array of mips asm words that must terminate with an ac_yield.
#define MipsAtom_(sym) MipsCode sym [] align_(4) =
// Used for components with no args (e.g., ac_load_tri_indices) or identifier-args (hardcoded register names).
// MipsAtomComp_(ac_X) { body }
// expands to:
// MipsCode ac_X[] align_(4) = { body };
#define MipsAtomComp_(sym) MipsCode sym [] align_(4) =
// Used for components with value-args (e.g., ac_format_f3_color).
// FI_ Slice_MipsCode ac_X(args) MipsAtomComp_Proc_(ac_X, { body })
// expands to:
// FI_ Slice_MipsCode ac_X(args) { MipsCode ac_X[] align_(4) = { body }; return slice_from_array(MipsCode, ac_X); }
#define MipsAtomComp_Proc_(sym, ...) { MipsCode sym [] align_(4) = __VA_ARGS__; return slice_from_array(MipsCode, sym); }
/* Line-table anchor: gcc only adds a file to the .debug_line file table when the
file contains line-numbered content. Files containing only:
- `MipsAtomComp_` static-array declarations, or
- `MipsAtomComp_Proc_` (force-inline) function bodies whose line info gets
attributed to the call site at the include point are otherwise omitted from the file table,
which breaks the DWARF injection when it tries to resolve atom-component provenance paths.
Place `ATOM_FILE_LINE_MARKER();` once at file scope in any `.atom.c` that defines atoms.
The macro expands to a file-scope `internal U4 const` declaration keeps the file in the line table.
The constant is in `.rodata` and unreferenced; the linker may eliminate it.
The two-level concat + `__LINE__` suffix makes the identifier unique per call site
(the identifier embeds the source line, so duplicates across `#include`d files don't collide). */
#define ATOM_FILE_DEBUGGER_LINE_MARKER(file_name) internal U4 const tmpl(atom_file_debugger_line_marker,file_name) = 0
typedef Slice_(MipsAtom); typedef Slice_MipsAtom Tape;
/* The 'Exit' Atom */
atom_dbg_skip MipsAtom_(tape_exit) { jump_reg(rret_addr), nop };
// TODO(Ed): When we have a substantial workload/throughput, profile each of these to see impact at ABI boundaries.
/* Tape Runner (Default) */
FI_ void tape_run(Tape tape) { register U4* tape_ptr rgcc(R_TapePtr) = u4_r(tape.ptr); asm volatile(
asm_words(
add_ui( R_SP, R_SP, -MipsStackAlignment) /* Allocate stack space */
, store_word(R_RA, R_SP, 0) /* Safely backup $ra to the stack */
, load_word( R_AtomJmp, R_TapePtr, 0) /* Bootstrap the first jump */
, add_ui_1( R_TapePtr, S_(MipsCode)) /* Advance tape */
, jump_nreg( R_AtomJmp) /* jalr $t9 */
, nop /* Branch delay slot */
, load_word(R_RA, R_SP, 0) /* Restore $ra from stack */
, add_ui_1( R_SP, MipsStackAlignment) /* Deallocate stack space */
load_word( R_AtomJmp, R_TapePtr, 0) /* Bootstrap the first jump */
, add_ui_self(R_TapePtr, S_(MipsAtom)) /* Advance tape */
, call_reg( R_AtomJmp) /* jalr $t9 */
, nop /* Branch delay slot */
)
asm_rpins, r_use(tp)
asm_rpins, r_use(tape_ptr)
asm_clobber:
rlit(R_AT)
, rlit(R_V0), rlit(R_V1)
, rlit(R_T0), rlit(R_T1), rlit(R_T2), rlit(R_T3)
/* Tell GCC the tape engine owns and destroys the workspace registers */
, rlit(R_PrimCursor), rlit(R_FaceCursor), rlit(R_VertBase), rlit(R_OtBase)
, rlit(R_T9)
, clb_mem_drain
rlit(R_AT),
rlit(R_V0), rlit(R_V1), // We clobber these for GTE ACs (that don't expose register selection, might expose them in the future...)
rlit(R_T0), rlit(R_T1), rlit(R_T2), rlit(R_T3), rlit(R_T4),
rlit(R_T5), rlit(R_T6), rlit(R_T7), rlit(R_T8),
clb_mem_drain
); }
/* Tape Runner (Static and Arg Clobbers) */
FI_ void tape_run_a02_s07(Tape tape) { register U4* tape_ptr rgcc(R_TapePtr) = u4_r(tape.ptr); asm volatile(
asm_words(
load_word( R_AtomJmp, R_TapePtr, 0) /* Bootstrap the first jump */
, add_ui_self(R_TapePtr, S_(MipsAtom)) /* Advance tape */
, call_reg( R_AtomJmp) /* jalr $t9 */
, nop /* Branch delay slot */
)
asm_rpins, r_use(tape_ptr)
asm_clobber:
rlit(R_AT),
rlit(R_V0), rlit(R_V1), rlit(R_A0), rlit(R_A1), rlit(R_A2),
rlit(R_T0), rlit(R_T1), rlit(R_T2), rlit(R_T3), rlit(R_T4),
rlit(R_T5), rlit(R_T6), rlit(R_T7), rlit(R_T8),
rlit(R_S0), rlit(R_S1), rlit(R_S2), rlit(R_S3), rlit(R_S4),
rlit(R_S5), rlit(R_S6), rlit(R_S7),
clb_mem_drain
); }
// Procedural authoring of tapes:
typedef Relative_(FArena) Struct_(TapeBuilder) { U4 ptr; U4 capacity; U4 used; };
FI_ void tb_init(TapeBuilder* tb, FArena* arena) { tb->ptr = arena->start; tb->used = 0; }
FI_ TapeBuilder tb_make_old( FArena* arena) { return (TapeBuilder){ arena->start, 0 }; }
FI_ TapeBuilder tb_make(Slice mem) { return (TapeBuilder){ mem.ptr, mem.len, 0 }; }
#define tb_emit_(tb, atom) tb_emit(tb, tmpl(code,atom))
FI_ void tb_emit(TapeBuilder* tb, MipsCode* atom) { u4_r(tb->ptr)[tb->used] = u4_(atom); ++ tb->used; }
FI_ void tb_data(TapeBuilder* tb, U4 data) { u4_r(tb->ptr)[tb->used] = u4_(data); ++ tb->used; }
#define tb_emit_(atom) tb_emit(& tb, atom)
#define tb_data_(field, data) tb_data(& tb, u4_(data))
FI_ Slice_U4 tb_end (TapeBuilder* tb) { tb_emit(tb,code_tape_exit); return (Slice_U4){ C_(U4*,tb->ptr), tb->used }; }
FI_ Slice_U4 tb_slice(TapeBuilder tb) { return (Slice_U4){ C_(U4*,tb.ptr), tb.used }; }
#define tb_scope(tb) for(U4 tbs_once=0;tbs_once==0;++tbs_once,tb_emit(tb,code_tape_exit))
FI_ Tape tb_end (TapeBuilder* tb) { tb_emit(tb,tape_exit); return (Tape){ C_(U4*,tb->ptr), tb->used }; }
FI_ Tape tb_slice(TapeBuilder tb) { return (Tape){ C_(U4*,tb.ptr), tb.used }; }
#define tb_scope(tb) for(U4 tbs_once=0;tbs_once==0;++tbs_once,tb_emit(tb,tape_exit))
FI_ void tb_scope_run_end(TapeBuilder* tb) { tb_emit(tb,tape_exit); tape_run(tb_slice(tb[0])); }
#define tb_scope_run(tb) for(U4 tbs_once=0;tbs_once==0;++tbs_once,tb_scope_run_end(tb))
#pragma endregion Tape Drive
#pragma region Macro Mips Atom Components
@@ -85,56 +211,34 @@ FI_ Slice_U4 tb_slice(TapeBuilder tb) { return (Sli
* These do NOT yield. They are expanded inline inside Tape Atoms.
* ---------------------------------------------------------------------------*/
/* The 'Yield' sequence for Tape Atoms.
* Loads the next pointer from the tape, advances the tape, and jumps.
* Cost: ~ 4 cycles */
#define mac_yield() \
load_word(R_AtomJmp, R_TapePtr, 0) \
, add_ui_1( R_TapePtr, S_(MipsCode)) \
, jump_reg( R_AtomJmp) \
, nop
// The 'Yield' sequence for Tape Atoms (mac_yield).
// - mac_yield() is the safe default for atom-endings: 4 words, BD-slot of jr is mandatory nop.
// - mac_yield_load() + mac_yield_tail():
// - unconditional branch: mac_yield_load fills the branch's BD-slot (replaces a nop);
// - mac_yield_tail runs at the branch target (does NOT re-load R_AtomJmp).
/* Words: 3; Loads 3 S2 indices from the face array */
#define mac_load_tri_indices(rId_0, rId_1, rId_2) \
load_half_u(rId_0, R_FaceCursor, 0 * S_(S2)) \
, load_half_u(rId_1, R_FaceCursor, 1 * S_(S2)) \
, load_half_u(rId_2, R_FaceCursor, 2 * S_(S2))
atom_dbg_skip MipsAtomComp_(ac_yield) {
load_word(R_AtomJmp, R_TapePtr, 0),
add_ui_self( R_TapePtr, S_(MipsCode)),
jump_reg( R_AtomJmp), nop,
};
/* Words: 18; Translates indices to vertex addresses and pushes them to GTE
R_AT = rId_[#] << 3;
R_AT += R_VertBase;
R_V0 = R_AT[0];
gte_mt(R_V0, V.xy[#]);
gte_mt(R_V1, V.z [#]);
*/
#define mac_load_tri_verts(rId_0, rId_1, rId_2) \
shift_ll(R_AT, rId_0, 3), add_u(R_AT, R_AT, R_VertBase), load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4), gte_mt(R_V0, C2_VXY0), gte_mt(R_V1, C2_VZ0) \
, shift_ll(R_AT, rId_1, 3), add_u(R_AT, R_AT, R_VertBase), load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4), gte_mt(R_V0, C2_VXY1), gte_mt(R_V1, C2_VZ1) \
, shift_ll(R_AT, rId_2, 3), add_u(R_AT, R_AT, R_VertBase), load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4), gte_mt(R_V0, C2_VXY2), gte_mt(R_V1, C2_VZ2)
atom_dbg_skip MipsAtomComp_(ac_yield_load) {
load_word(R_AtomJmp, R_TapePtr, 0),
};
//TODO(Ed): Add more type annotation
/* Words: 11; Correctly inserts a primitive into the Ordering Table linked list */
#define mac_insert_ot_tag(r_otz, prim_length) \
shift_ll( R_T1, r_otz, 2) \
, add_u( R_T1, R_T1, R_OtBase) /* T1 = &OrderingTable[OTZ] */ \
, load_word( R_AT, R_T1, 0) /* AT = old_ot_head */ \
, load_ui( R_V0, prim_length) /* V0 = Length << 24 */ \
, shift_ll( R_AT, R_AT, 8) /* Strip upper 8 bits from old_ot */ \
, shift_lr( R_AT, R_AT, 8) \
, or_u( R_AT, R_AT, R_V0) /* Merge length */ \
, store_word(R_AT, R_PrimCursor, 0) /* prim->tag = old_ot_head */ \
, shift_ll( R_AT, R_PrimCursor, 8) /* AT = PrimCur & 0x00FFFFFF */ \
, shift_lr( R_AT, R_AT, 8) \
, store_word(R_AT, R_T1, 0) /* OrderingTable[OTZ] = PrimCur */
atom_dbg_skip MipsAtomComp_(ac_yield_tail) {
add_ui_self(R_TapePtr, S_(MipsCode)),
jump_reg( R_AtomJmp), nop,
};
#pragma endregion Macro Atom Components
#pragma region Mips Atom Builder
// This allows for runtime procedural authoring of mips atoms.
// This helps with runtime procedural authoring of mips atoms.
typedef Struct_(FMipsAtom512) { U4 data[512]; U4 used; };
typedef Slice_(MipsCode); typedef Slice_MipsCode MipsAtom;
// FArena Related
typedef Relative_(FArena) Struct_(MipsAtomBuilder) { U4 start; U4 capacity; U4 used; };
// Whatever the builder is writting to should most likely coresspond
@@ -149,134 +253,14 @@ FI_ void atombuilder_unroll(MipsAtomBuilder_R ab, Slice_MipsCode_R code) {
// When done authoring, utilize this to cap-off the atom
FI_ void atombuilder_end(MipsAtomBuilder_R ab) {
LP_ MipsAtom_(yield) { mac_yield() };
mem_copy(ab->start, u4_(code_yield), S_(code_yield));
mem_bump(ab->start, ab->capacity, & ab->used, S_(code_yield));
mem_copy(ab->start, u4_(ac_yield), S_(ac_yield));
mem_bump(ab->start, ab->capacity, & ab->used, S_(ac_yield));
}
#define mipsatom_from_builder(ab) (MipsAtom){ab.start, ab.used}
#define mipsatom_from_builder(ab) (Slice_MipsCode){ab.start, ab.used}
#pragma endregion Mips Atom Builder
#pragma region Baked Mips Atoms
// These atoms are resolved at compile time and are (usually) statically linked readonly data.
enum {
bios_flushcache = 0x44,
bios_table_addr = 0xA0,
};
/* Flushes the Instruction Cache (PSX A-function 0x44 via BIOS stub at 0xA0).
*
* Sequence (per MIPS ABI; arguments in arg registers, RA pushed to stack):
* 1. sp -= 8; sw $ra, 4($sp) ; save RA
* 2. $a0 = bios_flushcache (arg0)
* 3. $t0 = bios_table_addr ; t0 = &BIOS A-function table
* 4. jalr $t0, $ra ; call BIOS(flushcache)
* nop ; branch delay slot
* 5. lw $ra, 4($sp); jr $ra ; restore & return
* 6. sp += 8
*/
// TODO(Ed): Annotate magic offsets
internal MipsAtom_(mips_flush_icache) {
add_ui(rstack_ptr, rstack_ptr, -8) /* sp -= 8 */
, store_word(rret_addr, rstack_ptr, 4) /* sw $ra, 4($sp) */
, add_ui(rret_0, rdiscard, bios_flushcache) /* addiu $a0, $0, 0x44 */
, add_ui(rtmp_0, rdiscard, bios_table_addr) /* addiu $t0, $0, 0xA0 */
, jump_link(rtmp_0, rret_addr) /* jalr $t0, $ra */
, nop /* BD slot */
, load_word(rret_addr, rstack_ptr, 4) /* lw $ra, 4($sp) */
, jump_reg(rret_addr) /* jr $ra */
, add_ui(rstack_ptr, rstack_ptr, 8) /* sp += 8 (BD) */
, mac_yield()
};
typedef Struct_(Binds_SetGteWorld) {
U4 transform;
};
// TODO(Ed): Bugged, fix
internal MipsAtom_(set_gte_world) {
/* Pop matrix address from tape into R_T3 ($11) */
load_word(R_T3, R_TapePtr, O_(Binds_SetGteWorld,transform)),
add_ui_1( R_TapePtr, S_(Binds_SetGteWorld)),
// TODO(Ed): Annotate magic offsets.
/* Load 3x3 Rotation + 3x1 Translation from R_T3 into GTE CONTROL Regs (ctc2) */
load_word(R_T0, R_T3, 0), load_word(R_T1, R_T3, 4),
gte_ct( R_T0, gte_cr_RT11), gte_ct( R_T1, gte_cr_RT12),
load_word(R_T0, R_T3, 8), load_word(R_T1, R_T3, 12), load_word(R_T2, R_T3, 16),
gte_ct( R_T0, gte_cr_RT13), gte_ct( R_T1, gte_cr_RT21), gte_ct( R_T2, gte_cr_RT22),
load_word(R_T0, R_T3, 20), load_word(R_T1, R_T3, 24), load_word(R_T2, R_T3, 28),
gte_ct( R_T0, gte_cr_TRX), gte_ct( R_T1, gte_cr_TRY), gte_ct( R_T2, gte_cr_TRZ),
mac_yield()
};
// TODO(Ed): I'm not sure yet if the bindings are redundant with the floortri atom yet.
/* DIAGNOSTIC 1: Pure tape loop test */
internal MipsAtom_(diag_yield) { mac_yield() };
// TODO(Ed): Reduce magic numbers/offsets
/* DIAGNOSTIC 2: Pure memory test (No GTE). Draws a fixed cyan triangle. */
internal MipsAtom_(diag_color) {
store_word(R_0, R_T7, 0),
load_ui( R_AT, 0x20FF), /* High: MipsCode 0x20 + Color B:FF */
or_i( R_AT, R_AT, 0xFF00), /* Low: Color G:FF, R:00 (Cyan) */
store_word(R_AT, R_T7, 4),
/* Fake coordinates - Swapped winding order to prevent GPU culling! */
load_ui(R_AT, 0x0010), or_i(R_AT, R_AT, 0x0010), store_word(R_AT, R_T7, 8), /* (16, 16) */
load_ui(R_AT, 0x0050), or_i(R_AT, R_AT, 0x0010), store_word(R_AT, R_T7, 12), /* (80, 16) */
load_ui(R_AT, 0x0010), or_i(R_AT, R_AT, 0x0050), store_word(R_AT, R_T7, 16), /* (16, 80) */
add_ui( R_T1, R_0, 10),
shift_ll(R_T1, R_T1, 2),
add_u( R_T1, R_T1, R_T6),
load_word( R_AT, R_T1, 0),
load_ui( R_V0, 0x0400), // <--- Fills load delay slot!
store_word(R_AT, R_T7, 0),
shift_ll( R_AT, R_T7, 8), shift_lr(R_AT, R_AT, 8),
or_u( R_AT, R_AT, R_V0),
store_word(R_AT, R_T1, 0),
add_ui(R_T7, R_T7, 20),
mac_yield()
};
// TODO(Ed): Reduce magic numbers/offsets
/* DIAGNOSTIC 3: Pure GTE test (No Memory Writes) */
internal MipsAtom_(diag_gte) {
/* Load 3 indices */
load_half_u(R_T0, R_T4, 0),
load_half_u(R_T1, R_T4, 2),
load_half_u(R_T2, R_T4, 4),
/* Load Vertices into GTE */
shift_ll( R_AT, R_T0, 3), add_u( R_AT, R_AT, R_T5),
load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4),
gte_mt( R_V0, C2_VXY0), gte_mt( R_V1, C2_VZ0),
shift_ll( R_AT, R_T1, 3), add_u( R_AT, R_AT, R_T5),
load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4),
gte_mt( R_V0, C2_VXY1), gte_mt( R_V1, C2_VZ1),
shift_ll( R_AT, R_T2, 3), add_u( R_AT, R_AT, R_T5),
load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4),
gte_mt( R_V0, C2_VXY2), gte_mt( R_V1, C2_VZ2),
/* Run Math */
nop, nop, gte_cmdw_rtpt,
nop, nop, gte_cmdw_nclip,
nop, nop,
/* Advance Face Cursor and Yield */
add_ui(R_T4, R_T4, 8),
mac_yield()
};
#pragma endregion Baked Mips Atoms
+29
View File
@@ -0,0 +1,29 @@
#ifdef INTELLISENSE_DIRECTIVES
# include "gen/macs.h"
# include "gen/offsets.h"
# include "math.h"
# include "lottes_tape.h"
#endif
ATOM_FILE_DEBUGGER_LINE_MARKER(math_atom_c);
#pragma region MACs (Mips Atom Component)
FI_ Slice_MipsCode ac_load_v2s2(U4 rs_x, U4 rs_y, U4 r_base, U4 offset) atom_dbg_skip MipsAtomComp_Proc_(ac_load_v2s2, {
load_half( rs_x, r_base, O_(V3_S2,x)),
load_half( rs_y, r_base, O_(V3_S2,y)),
})
FI_ Slice_MipsCode ac_store_v2s2(U4 rt_x, U4 rt_y, U4 base, U4 offset) atom_dbg_skip MipsAtomComp_Proc_(ac_store_v2s2, {
store_half(rt_x, base, offset + O_(V2_S2,x)),
store_half(rt_y, base, offset + O_(V2_S2,y)),
})
FI_ Slice_MipsCode ac_store_rects2(U4 rt_x, U4 rt_y, U4 rt_width, U4 rt_height, U4 base, U4 offset) atom_dbg_skip MipsAtomComp_Proc_(ac_store_rects2, {
store_half(rt_x, base, offset + O_(Rect_S2,x)),
store_half(rt_y, base, offset + O_(Rect_S2,y)),
store_half(rt_width, base, offset + O_(Rect_S2,width)),
store_half(rt_height, base, offset + O_(Rect_S2,height)),
})
#pragma endregion MACs (Mips Atom Component)
+12 -10
View File
@@ -7,6 +7,11 @@
#define max(A, B) (((A) > (B)) ? (A) : (B))
#define clamp_bot(X, B) max(X, B)
enum {
v3s2_byteoff = 3, // log2(8), used with shift_left_logical op for index via byte offset.
};
typedef Array_(U1, 2);
typedef Array_(U4, 2);
typedef Array_(S2, 2);
typedef Array_(S2, 3);
@@ -18,6 +23,7 @@ typedef S2 A3x3_S2[3][3];
typedef Struct_(Extent2_S2) { S2 width; S2 height; };
typedef Struct_(Extent2_S4) { S4 width; S4 height; };
typedef Struct_(V2_U1) { U1 x; U1 y; };
typedef Struct_(V2_S2) { S2 x; S2 y; };
typedef Struct_(V2_S4) { S4 x; S4 y; };
typedef Struct_(V3_S2) { S2 x; S2 y; S2 z; S2 pad; };
@@ -28,11 +34,12 @@ typedef Struct_(V4_S4) { S4 x; S4 y; S4 z; S4 w; };
typedef Struct_(R2_S2) { V2_S2 p0; V2_S2 p1; };
typedef Struct_(R2_S4) { V2_S4 p0; V2_S4 p1; };
typedef Struct_(Rect_S2) { S2 x; S2 y; S2 width; S2 height; };
typedef Struct_(Rect_S4) { S4 x; S4 y; S4 width; S4 height; };
typedef Struct_(Rect_S2) { S2 x; S2 y; S2 width; S2 height; };
typedef Struct_(Rect_S4) { S4 x; S4 y; S4 width; S4 height; };
typedef Struct_(M3_S2) { A3x3_S2 m; A3_S4 t; };
typedef Struct_(M3_S2) { A3x3_S2 m; A3_S4 t; };
typedef Array_(V2_S2, 2);
typedef Array_(V2_S2, 3);
typedef Array_(V2_S2, 4);
@@ -54,10 +61,5 @@ FI_ void add_a3s4_fp(A3_S4_R out_a, A3_S4 b) {
(out_a[0])[2] += b[2] >> 1;
}
FI_ void add_v3s4(V3_S4_R out_a, V3_S4 b) {
add_a3s4(pcast(A3_S4_R, out_a), pcast(A3_S4, b));
}
FI_ void add_v3s4_fp(V3_S4_R out_a, V3_S4 b) {
add_a3s4_fp(pcast(A3_S4_R, out_a), pcast(A3_S4, b));
}
FI_ void add_v3s4 (V3_S4_R out_a, V3_S4 b) { add_a3s4 (pcast(A3_S4_R, out_a), pcast(A3_S4, b)); }
FI_ void add_v3s4_fp(V3_S4_R out_a, V3_S4 b) { add_a3s4_fp(pcast(A3_S4_R, out_a), pcast(A3_S4, b)); }
+4 -3
View File
@@ -67,12 +67,13 @@ typedef Slice_(B1);
#define slice_end(slice) ((slice).ptr + (slice).len)
#define S_slice(s) ((s).len * S_((s).ptr[0]))
#define slice_ut(ptr,len) slice_ut_(u4_(ptr), u4_(len))
#define slice_ut_arr(a) slice_ut_(u4_(a), S_(a))
#define slice_ut(ptr,len) slice_ut_(u4_(ptr), u4_(len))
#define slice_ut_arr(a) slice_ut_(u4_(a), S_(a))
#define slice_to_ut(s) slice_ut_(u4_((s).ptr), S_slice(s))
#define slice_iter(container, iter) (T_((container).ptr) iter = (container).ptr; iter != slice_end(container); ++ iter)
#define slice_arg_from_array(type, ...) & (tmpl(Slice,type)) { .ptr = array_decl(type,__VA_ARGS__), .len = array_len( array_decl(type,__VA_ARGS__)) }
#define slice_from_array(type, array) (tmpl(Slice,type)) { .ptr = array, .len = S_(array) }
FI_ void slice_zero_(Slice s) { slice_assert(s); mem_zero(s.ptr, s.len); }
#define slice_zero(s) slice_zero_(slice_to_ut(s))
@@ -102,7 +103,7 @@ FI_ void farena_init(FArena_R arena, Slice mem) { assert(arena != nullptr);
arena->used = 0;
}
FI_ FArena farena_make(Slice mem) { FArena a; farena_init(& a, mem); return a; }
I_ Slice farena_push(FArena_R arena, U4 amount, Opt_farena o) {
I_ Slice farena_push(FArena_R arena, U4 amount, Opt_farena o) {
if (amount == 0) { return (Slice){}; }
U4 desired = amount * (o.type_width == 0 ? 1 : o.type_width);
U4 to_commit = align_pow2(desired, o.alignment ? o.alignment : MEM_ALIGNMENT_DEFAULT);
+38
View File
@@ -0,0 +1,38 @@
#ifdef INTELLISENSE_DIRECTIVES
# include "gen/macs.h"
# include "gen/offsets.h"
# include "lottes_tape.h"
#endif
ATOM_FILE_DEBUGGER_LINE_MARKER(mips_atom_c);
#pragma region Baked Atoms
enum {
bios_flushcache = 0x44,
bios_table_addr = 0xA0,
};
/* Flushes the Instruction Cache (PSX A-function 0x44 via BIOS stub at 0xA0).
* Sequence (per MIPS ABI; arguments in arg registers, RA pushed to stack):
* 1. sp -= 8; sw $ra, 4($sp) ; save RA
* 2. $a0 = bios_flushcache (arg0)
* 3. $t0 = bios_table_addr ; t0 = &BIOS A-function table
* 4. jalr $t0, $ra ; call BIOS(flushcache)
* nop ; branch delay slot
* 5. lw $ra, 4($sp); jr $ra ; restore & return
* 6. sp += 8
*/
internal MipsAtom_(mips_flush_icache) {
add_ui(rstack_ptr, rstack_ptr, -MipsStackAlignment), // sp -= 8
store_word(rret_addr, rstack_ptr, S_(U4)), // sw $ra, 4($sp)
add_ui(rret_0, rdiscard, bios_flushcache), // addiu $a0, $0, 0x44
add_ui(rtmp_0, rdiscard, bios_table_addr), // addiu $t0, $0, 0xA0
jump_link(rtmp_0, rret_addr), nop, // jalr $t0, $ra, BD slot
load_word(rret_addr, rstack_ptr, S_(U4)), // lw $ra, 4($sp)
jump_reg(rret_addr), // jr $ra
add_ui(rstack_ptr, rstack_ptr, MipsStackAlignment), // sp += 8 (BD)
mac_yield(),
};
#pragma endregion Baked Atoms
+165 -89
View File
@@ -1,3 +1,69 @@
/* ============================================================================
* duffle DSL Suffix Conventions
* ============================================================================
* Every mnemonic in this header follows the same suffix grammar:
* _i: Immediate value (16-bit constant operand).
* Combine with _u or _s (single-letter modifier + type combined): add_ui, add_si.
* Examples: add_ui, add_si, and_i, or_i, xor_i, load_upper_i. and_i is sign-agnostic (andi zero-extends).
* load_upper_i is a unique verb; _i is the immediate marker, not a modifier+type combination.
* _u: Unsigned (no-overflow, no-sign-extension).
* R-type arithmetic examples: add_u, sub_u, mult_u, div_u. I-type (combined with _i): add_ui.
* _s: Signed (overflow-traps, sign-extends).
* R-type: add_s, sub_s, mult_s, div_s, set_lt_s. I-type (combined with _i): add_si.
*
* --- Shift family (R-type): verb-modifier-direction ---
* The shift macros use `shift_<modifier><direction>`.
* Modifier is the single letter `l` (logical) or `a` (arithmetic).
* Direction is the word `left` or `right`. Combined: `_lleft`, `_lright`, `_aright`.
* Examples: shift_lleft( rd, rt, shamt) (= sll)
* shift_lright(rd, rt, shamt) (= srl)
* shift_aright(rd, rt, shamt) (= sra)
* (no `_aleft`; MIPS has no `sla` — arithmetic-left is bit-identical to logical-left, so use shift_lleft for that case)
*
* --- Jump/Call family ---
* Simple jumps keep the original short names: jump (j), jump_reg (jr), jump_link (jalr rs, rd).
* The jump-and-link-to variants (jal, jalr rs with default $ra) get the `call_` verb instead:
* call_addr (jal), call_reg (jalr rs, default $ra).
* Examples: jump(off) (= j)
* jump_reg(rs) (= jr)
* jump_link(rs, rd) (= jalr rs, rd)
* call_reg(rs) (= jalr rs, default $ra)
* call_addr(off) (= jal)
*
* _r: Register marker — used only when the register type needs disambiguation (e.g., GTE data register vs control register).
* NOT used in plain R-type arithmetic (the R-type is implicit). Examples: gte_mv_to_data_r, gte_mv_to_ctrl_r.
* _self: Destination equals one source operand.
* Examples: add_ui_self (I-type, to self), add_u_self (R-type, to self).
* _mv_to_: Direction: data flows into X.
* Example: gte_mv_to_data_r, gte_mv_to_ctrl_r.
* _mv_from_: Direction: data flows out of X.
* Example: gte_mv_from_data_r, gte_mv_from_ctrl_r.
* _str: String-form — emits inline-asm string instead of `.word`.
* Example: gte_rtpt_asm_str.
* _2w / _1w: Word count of the emitted sequence.
* Example: load_imm_2w.
*
* _cop2: RESERVED — DO NOT USE in macro names. The `gte_` namespace prefix already implies coprocessor 2. Use `c2` only in:
* (a) integer opcode enums (op_lwc2 = 0x32, op_swc2 = 0x3A)
* (b) vendor-mnemonic macro aliases (gte_mtc2, gte_mfc2)
*
* Primitive commands: gp0_cmd_poly_f3 = 0x20 (byte opcode)
* Packed 32-bit cmd: gp0_word_poly_f3(r, g, b) (32-bit, shifted)
*
* Type ordering: domain?_(direction)?_action_target_modifier_type?
* Examples: add_ui (add + unsigned + immediate)
* add_s (add + signed, R-type implicit)
* shift_lleft (shift + logical + left)
* shift_aright (shift + arithmetic + right)
* call_reg(rs) (call + register, $ra implicit)
* gte_mv_to_data_r (gte + mv + to + data + register)
* gte_lw_v0_xy(base) (gte + lw + v0 + xy)
* load_upper_i (load-upper + immediate, unique verb)
*
* Vendor mnemonics (sll, srl, sra, jr, j, jal, jalr) are NOT in this header.
* They live in the opt-in `mips_vendor_sym.h` for users who prefer the textbook MIPS assembly mnemonics.
* ============================================================================ */
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "dsl.h"
@@ -11,19 +77,17 @@ enum {
/* ============================================================================
* REGISTER INTEGER IDS (preprocessor-visible)
* ============================================================================
* Every R_* enum below has a parallel R_*_Code `#define` so that the
* preprocessor can stringify the integer (e.g. for asm clobber lists and
* register-variable declarations via `rgcc(R_X)`). The enum value is
* bound to the `#define` so the two forms cannot drift apart.
* Every R_* enum below has a parallel R_*_Code `#define` so that the preprocessor can stringify the integer
* (e.g. for asm clobber lists and register-variable declarations via `rgcc(R_X)`).
* The enum value is bound to the `#define` so the two forms cannot drift apart.
*
* Only registers that get stringified need a `_Code` form; the rest are
* plain enum values. If you need to add a new one, follow the pattern:
* Only registers that get stringified need a `_Code` form; the rest are plain enum values.
* If you need to add a new one, follow the pattern:
* #define R_T7_Code 15
* R_T7 = R_T7_Code, // in the enum
* R_T7 = R_T7_Code, // in the enum
*
* User code should always reference the enum form (`R_T4`) at arithmetic
* sites and let `rlit(R_T4_Code)` / `rgcc(R_T4)` handle the stringify
* cases — never write the bare number `12`.
* User code should always reference the enum form (`R_T4`) at arithmetic sites and let
* `rlit(R_T4_Code)` / `rgcc(R_T4)` handle the stringify cases — never write the bare number `12`.
* ============================================================================ */
#define R_0_Code 0
#define R_AT_Code 1
@@ -138,7 +202,6 @@ enum {
/* 2F: N/A */
// , op_lwc0
// , op_load_addr = op_la
// , op_load_imm = op_li
, op_jump = op_j
@@ -240,13 +303,15 @@ enum { _BitOffsets = 0
* Argument order matches the MIPS assembly syntax:
* dest-first, then source operands, then immediate last.
*
* load_word(rt, base, off) → lw rt, off(base)
* store_word(rt, base, off) → sw rt, off(base)
* add_ui(rt, rs, imm) → addiu rt, rs, imm
* shift_ll(rd, rt, shamt) → sll rd, rt, shamt
* jump_reg(rs) → jr rs
* jump_link(rs, rd) → jalr rs (link in rd, default $ra)
* nop sll $0, $0, 0
* load_word(rt, base, off) → lw rt, off(base)
* store_word(rt, base, off) → sw rt, off(base)
* add_ui(rt, rs, imm) → addiu rt, rs, imm
* shift_lleft(rd, rt, shamt) → sll rd, rt, shamt
* shift_lright(rd, rt, shamt) → srl rd, rt, shamt
* shift_aright(rd, rt, shamt) → sra rd, rt, shamt
* jump_reg(rs)jr rs
* jump_link(rs, rd) → jalr rs (link in rd, default $ra)
* nop → sll $0, $0, 0
*/
#define load_word(rt, base, off) enc_i(op_lw, (base), (rt), (off))
#define load_byte(rt, base, off) enc_i(op_lb, (base), (rt), (off))
@@ -255,29 +320,37 @@ enum { _BitOffsets = 0
#define load_half_u(rt, base, off) enc_i(op_lhu, (base), (rt), (off))
#define store_word(rt, base, off) enc_i(op_sw, (base), (rt), (off))
#define add_ui(rt, rs, imm) enc_i(op_addiu, (rs), (rt), (imm))
#define and_si(rt, rs, imm) enc_i(op_andi, (rs), (rt), (imm))
#define and_i(rt, rs, imm) enc_i(op_andi, (rs), (rt), (imm))
// #define and_si and_i
#define or_i(rt, rs, imm) enc_i(op_ori, (rs), (rt), (imm))
#define xor_i(rt, rs, imm) enc_i(op_xori, (rs), (rt), (imm))
#define load_ui(rt, imm) enc_i(op_lui, R_0, (rt), (imm))
#define load_upper_i(rt, imm) enc_i(op_lui, R_0, (rt), (imm))
#define load_u1 load_byte_u
#define load_u2 load_half_u
#define load_u4 load_word
// Ergonomic add to the same register.
#define add_ui_1(rt_rs, imm) enc_i(op_addiu, (rt_rs), (rt_rs), (imm))
#define or_i_self(rt_rs, imm) enc_i(op_ori, (rt_rs), (rt_rs), (imm))
#define add_ui_self(rt_rs, imm) enc_i(op_addiu, (rt_rs), (rt_rs), (imm))
/* Logic Opcodes */
#define and_u(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_and)
#define or_u(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_or)
#define xor_u(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_xor)
#define nor_u(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_nor)
#define and_u(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_and)
#define or_u(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_or)
#define xor_u(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_xor)
#define nor_u(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_nor)
/* Shift family (R-type). shift_ll/lr/ra: `sll rd, rt, shamt` */
#define shift_ll(rd, rt, shamt) enc_r(op_special, R_0, (rt), (rd), (shamt), fc_sll)
#define shift_lr(rd, rt, shamt) enc_r(op_special, R_0, (rt), (rd), (shamt), fc_srl)
#define shift_ra(rd, rt, shamt) enc_r(op_special, R_0, (rt), (rd), (shamt), fc_sra)
#define or_u_self(rd_rs, rt) enc_r(op_special, (rd_rs), (rt), (rd_rs), 0, fc_or)
/* Shift family (R-type). shift_lleft/lright/aright: `sll/srl/sra rd, rt, shamt` */
#define shift_lleft(rd, rt, shamt) enc_r(op_special, R_0, (rt), (rd), (shamt), fc_sll)
#define shift_lright(rd, rt, shamt) enc_r(op_special, R_0, (rt), (rd), (shamt), fc_srl)
#define shift_aright(rd, rt, shamt) enc_r(op_special, R_0, (rt), (rd), (shamt), fc_sra)
#define shift_lleft_self(rd_rt, shamt) enc_r(op_special, R_0, (rd_rt), (rd_rt), (shamt), fc_sll)
#define mask_upper(rd, rt, shamt) shift_lleft(rd, rt, shamt), shift_lright(rd, rt, shamt)
/* jr rs — jump to address in rs. */
#define jump_reg(rs) enc_r(op_special, (rs), R_0, R_0, 0, fc_jr)
@@ -286,14 +359,32 @@ enum { _BitOffsets = 0
* Layout: [op_special][rs:5][rt=0:5][rd:5][shamt=0:5][fc_jalr=0x09] */
#define jump_link(rs, rd) enc_r(op_special, (rs), R_0, (rd), 0, fc_jalr)
/* jalr rs — link in $ra and jump to address in rs (most common form). */
#define jump_nreg(rs) jump_link((rs), R_RA)
/* call_reg rs — jump-and-link to register-held address; link in $ra. */
#define call_reg(rs) jump_link((rs), R_RA)
/* j target — absolute jump within the current 256MB region. */
/* j target — absolute jump within the current 256MB region.
* WARNING: `jump(off)` CANNOT BE USED for within-atom jumps in the current pipeline.
* The MIPS j opcode encodes `(target_addr >> 2)` in its 26-bit immediate field; an ABSOLUTE byte address, not a relative word offset.
* The metaprogram computes `off` as a relative word offset (`target_word_idx - branch_word_idx - 1`), which the assembler/linker does NOT resolve.
*
* `jump(off)` is only safe when the BUILD PIPELINE owns the absolute position of the emitted code — i.e. when: s
* - the build emits a symbol-relative `.word` expression that the linker resolvess via `R_MIPS_26`, OR
* - the code is hand-assembled with explicit absolute targets, OR a custom post-build patcher resolves the 26-bit field.
*/
#define jump(off) enc_i(op_j, R_0, R_0, (off))
/* jal target — absolute call within the current 256MB region. */
#define jump_nlink(off) enc_i(op_jal, R_0, R_0, (off))
/* jump_rel off — unconditional relative jump (the within-atom-safe `jump`).
* MIPS I R3000A has no "branch always" opcode. The idiom for an unconditional relative jump is `beq $0, $0, off`.
*/
#define jump_rel(off) branch_equal(R_0, R_0, (off))
/* call_addr off — jump-and-link to immediate address.
*
* Same WARNING as `jump(off)` above: the jal opcode also encodes an absolute 26-bit target.
* For within-atom calls, the current pipeline has no equivalent always-taken call-and-link idiom.
* Workaround: `branch_link` (always-taken branch + explicit `la $ra, next_word_addr; jr $ra`), or just use `call_reg($tmp)` after loading the target into a register.
*/
#define call_addr(off) enc_i(op_jal, R_0, R_0, (off))
/* --- Store family (mirrors the load family) --- */
#define store_byte(rt, base, off) enc_i(op_sb, (base), (rt), (off))
@@ -307,12 +398,9 @@ enum { _BitOffsets = 0
* mult_s / mult_u → mult / multu (writes HI/LO; result in LO)
* div_s / div_u → div / divu (LO = quot, HI = rem)
*
* NOTE: dsl.h defines `add_s`/`sub_s`/`mut_s`/`gt_s`/etc. as
* _Generic-based signed integer-arithmetic helpers for U1/U2/U4. Those
* live in a different conceptual layer (generic arithmetic on DSL
* types) and would collide with the instruction encoders here. The
* `#undef` below lets the gas-style names below win; if a file needs
* both, the dsl.h versions can be reached via their long forms
* NOTE: dsl.h defines `add_s`/`sub_s`/`mut_s`/`gt_s`/etc. as _Generic-based signed integer-arithmetic helpers for U1/U2/U4.
* Those live in a different conceptual layer (generic arithmetic on DSL types) and would collide with the instruction encoders here.
* The `#undef` below lets the gas-style names below win; if a file needs both, the dsl.h versions can be reached via their long forms
* (e.g. `def_signed_op`-style or the underlying `add_s1/s2/s4`). */
#undef add_s
#undef sub_s
@@ -325,15 +413,17 @@ enum { _BitOffsets = 0
#define div_s(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_div)
#define div_u(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_divu)
#define add_u_self(rd_rs, rt) add_u(rd_rs, rd_rs, rt)
/* --- Arithmetic I-type (immediate) --- */
#define add_si(rt, rs, imm) enc_i(op_addi, (rs), (rt), (imm))
/* add_ui already exists above as add_ui */
/* --- Set on less than (R-type and I-type) --- */
#define slt_s(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_slt)
#define slt_u(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_sltu)
#define slt_si(rt, rs, imm) enc_i(op_slti, (rs), (rt), (imm))
#define slt_ui(rt, rs, imm) enc_i(op_sltiu, (rs), (rt), (imm))
#define set_lt_s(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_slt)
#define set_lt_u(rd, rs, rt) enc_r(op_special, (rs), (rt), (rd), 0, fc_sltu)
#define set_lt_si(rt, rs, imm) enc_i(op_slti, (rs), (rt), (imm))
#define set_lt_ui(rt, rs, imm) enc_i(op_sltiu, (rs), (rt), (imm))
/* --- Move from/to HI/LO (mult/div results) --- */
#define mov_from_high(rd) enc_r(op_special, R_0, R_0, (rd), 0, fc_mfhi)
@@ -342,7 +432,7 @@ enum { _BitOffsets = 0
#define mov_to_low(rs) enc_r(op_special, (rs), R_0, R_0, 0, fc_mtlo)
/* --- Atomic branches (no pseudos like bgt/bge; compose with slt_* + branch_ne) ---
* branch_equal rs, rt, off → beq rs, rt, off
* branch_equal rs, rt, off → beq rs, rt, off
* branch_ne rs, rt, off → bne rs, rt, off
* branch_lt_zero rs, off → bltz rs, off
* branch_gt_zero rs, off → bgtz rs, off
@@ -362,32 +452,26 @@ enum { _BitOffsets = 0
#define breakpoint() enc_r(op_special, R_0, R_0, R_0, 0, fc_break)
/* --- Shift-amount alias (matches the gas convention `\p3 = shamt`) --- */
#define shift_amount(rd, rt, n) shift_ll(rd, rt, n)
#define shift_amount(rd, rt, n) shift_lleft(rd, rt, n)
/* nop — canonical sll $0, $0, 0 */
#define nop shift_ll(rdiscard, rdiscard, 0)
/* nop — sll $0, $0, 0 */
#define nop shift_lleft(rdiscard, rdiscard, 0)
#define nop2 nop, nop
#define load_imm_1w(rt, imm) add_ui((rt), R_0, (imm))
#define load_imm_1w_s0(rt, imm) add_si((rt)), R_0, (imm))
/* load_imm_2w — unconditional 2-word `li` form: `lui` + (ori | addi).
*
* Granular companion to `load_imm`: skips the compile-time range checks
* and always emits 2 .words. Use this when:
* Granular companion to `load_imm`: skips the compile-time range checks and always emits 2 .words. Use this when:
* - you know `imm` is > 0xFFFF (otherwise you're wasting a word), OR
* - `imm` is not a compile-time constant and you want predictable
* 2-word emission without the `__builtin_constant_p` branches.
* - `imm` is not a compile-time constant and you want predictable 2-word emission without the `__builtin_constant_p` branches.
*
* The lo16 strategy is still chosen at expansion time on the lo half:
* lo16 in 0x0000..0x7FFF → addi (sign-ext is harmless, the lui
* already cleared bits 15..0)
* lo16 in 0x8000..0xFFFF → ori (zero-extends to preserve the
* intended bit pattern)
*
* For situations where you need to bypass even this choice (e.g. to
* force a specific encoding for a known discontiguous high/low pair),
* see `load_imm_2w_ori` and `load_imm_2w_addi` below.
* lo16 in 0x0000..0x7FFF → addi (sign-ext is harmless, the lui already cleared bits 15..0)
* lo16 in 0x8000..0xFFFF ori (zero-extends to preserve the intended bit pattern)
*
* For situations where you need to bypass even this choice (e.g. to force a specific encoding for a known discontiguous high/low pair),
* see `load_imm_2w_ori_forced` and `load_imm_2w_addi_forced` below.
* Statement-level (not expression-level): emits its own `asm volatile(...)`.
*/
#define load_imm_2w(rt, imm) do { \
@@ -407,9 +491,9 @@ enum { _BitOffsets = 0
} \
} while (0)
/* load_imm_2w_ori — force the `lui` + `ori` form regardless of lo16 sign.
/* load_imm_2w_ori_forced — force the `lui` + `ori` form regardless of lo16 sign.
* Use when you specifically need zero-extension in the lo half. */
#define load_imm_2w_ori(rt, imm) do { \
#define load_imm_2w_ori_forced(rt, imm) do { \
asm volatile( \
asm_words(load_ui((rt), u4_lo(imm)), \
or_i((rt), (rt), C_(U2,u4_hi(imm))) ) \
@@ -417,11 +501,10 @@ enum { _BitOffsets = 0
); \
} while (0)
/* load_imm_2w_addi — force the `lui` + `addi` form regardless of lo16 sign.
* Use when you know sign-extension is fine (e.g. lo16 is treated as
* signed downstream) and you want a smaller effective instruction
* (the assembler/MIPS hardware will sign-extend the imm16). */
#define load_imm_2w_addi(rt, imm) do { \
/* load_imm_2w_addi_forced — force the `lui` + `addi` form regardless of lo16 sign.
* Use when you know sign-extension is fine (e.g. lo16 is treated as signed downstream)
* and you want a smaller effective instruction (the assembler/MIPS hardware will sign-extend the imm16). */
#define load_imm_2w_addi_forced(rt, imm) do { \
/*U4 _li2a_imm_ = (U4)(imm);*/ \
asm volatile(asm_words( \
lui_op((rt), u4_lo(imm)), \
@@ -432,23 +515,17 @@ enum { _BitOffsets = 0
/* load_imm rt, imm — true `li` semantics (assembler `li` pseudo)
*
* Dispatches at compile time on the immediate's range, picking the
* smallest single-instruction form when possible:
*
* imm in 0 .. 0x7FFF addi rt, $0, imm (1 word)
* imm in 0x8000 .. 0xFFFF → ori rt, $0, imm (1 word; sign-bit must be zeroed)
* imm in 0x10000 .. 0xFFFFFFFF → lui + (ori | addi) (2 words)
*
* Statement-level (not expression-level): the macro emits its own
* `asm volatile(...)` block with 1 or 2 .word constants. Callers can
* group multiple `load_imm` calls in a single volatile by using the
* lower-level encoders directly:
* Dispatches at compile time on the immediate's range, picking the smallest single-instruction form when possible:
* imm in 0 .. 0x7FFF → addi rt, $0, imm (1 word)
* imm in 0x8000 .. 0xFFFF → ori rt, $0, imm (1 word; sign-bit must be zeroed)
* imm in 0x10000 .. 0xFFFFFFFF → lui + (ori | addi) (2 words)
*
* Statement-level (not expression-level): the macro emits its own `asm volatile(...)` block with 1 or 2 .word constants.
* Callers can group multiple `load_imm` calls in a single volatile by using the lower-level encoders directly:
* load_imm(R_T4, 0x12345678); // emits 2 .words
*
* Falls back to a 2-word form if `imm` is not a compile-time constant,
* but that path is unusual (load_imm is most useful with literal
* addresses and magic numbers). */
* Falls back to a 2-word form if `imm` is not a compile-time constant, but that path is unusual
* (load_imm is most useful with literal addresses and magic numbers). */
#define load_imm(rt, imm) do { \
if (cexpr_(imm) && ((imm) <= 0x7FFFU)) { \
/* Small positive: addi rt, $0, imm */ \
@@ -488,13 +565,12 @@ enum { _BitOffsets = 0
/* Standard clobber list for pure-MIPS asm volatile blocks: caller-saved
* GPRs that the kernel treats as volatile (v0/v1/t0/t1/ra) plus the
* "memory" barrier. The register ids are passed through `rlit` so
* the R_*_Code `#define`s are stringified into "$N" at expansion time. */
#define clb_system rlit(R_V0), rlit(R_T0), rlit(R_T1), rlit(R_RA), clb_mem_drain
* GPRs that the kernel treats as volatile (v0/v1/t0/t1/ra) plus the "memory" barrier.
* The register ids are passed through `rlit` so the R_*_Code `#define`s are stringified into "$N" at expansion time. */
#define clbr_volatile_gprs rlit(R_V0), rlit(R_T0), rlit(R_T1), rlit(R_RA), clb_mem_drain
#define asm_mips_flush_icache() asm volatile( asm_words( \
add_ui(rstack_ptr, rstack_ptr, -8) \
add_ui(rstack_ptr, rstack_ptr, -MipsStackAlignment) \
, store_word(rret_addr, rstack_ptr, 4) \
, add_ui(rret_0, rdiscard, bios_flushcache) \
, add_ui(rtmp_0, rdiscard, bios_table_addr) \
@@ -502,5 +578,5 @@ enum { _BitOffsets = 0
, nop \
, load_word(rret_addr, rstack_ptr, 4) \
, jump_reg(rret_addr) \
, add_ui(rstack_ptr, rstack_ptr, 8) \
) asm_clobber: clb_system )
, add_ui(rstack_ptr, rstack_ptr, MipsStackAlignment) \
) asm_clobber: clbr_volatile_gprs )
+44
View File
@@ -0,0 +1,44 @@
/* ============================================================================
* duffle DSL — MIPS Vendor Mnemonics (opt-in)
* ============================================================================
*
* Provides the textbook MIPS assembly mnemonics as thin aliases to the duffle macros in mips.h.
* The duffle names are primary; this header is for users who prefer the textbook mnemonics.
*
* USAGE: #include "duffle/mips_vendor_sym.h" // after mips.h
*
* Mapping (vendor -> duffle):
* Shift family:
* sll -> shift_lleft (shift left logical)
* srl -> shift_lright (shift right logical)
* sra -> shift_aright (shift right arithmetic)
* (no sllv/srlv/srav; the shift macros take a literal shamt)
*
* Jump family (1-arg / implicit-rd forms):
* jr -> jump_reg (jump register)
* j -> jump (jump to immediate address)
* jal -> call_addr (jump-and-link to immediate address)
* jalr -> call_reg (jump-and-link to register, default $ra)
* (for the 2-arg `jalr rs, rd`, use `jump_link(rs, rd)` directly)
* ============================================================================ */
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "mips.h"
#endif
#ifndef DUFFLE_MIPS_VENDOR_SYM_H
#define DUFFLE_MIPS_VENDOR_SYM_H
/* Shift family */
#define sll shift_lleft
#define srl shift_lright
#define sra shift_aright
/* Jump family (1-arg / implicit-$ra forms) */
#define jr jump_reg
#define j jump
#define jal call_addr
#define jalr call_reg
#endif
+184
View File
@@ -0,0 +1,184 @@
#ifdef INTELLISENSE_DIRECTIVES
# include "gen/macs.h"
# include "gen/offsets.h"
# include "mips.h"
# include "dsl.atom.h"
# include "lottes_tape.h"
# include "pad.h"
#endif
ATOM_FILE_DEBUGGER_LINE_MARKER(pad_atom_c);
#pragma region Baked Atoms
/* ----- pad_bios_snapshot -----
* Per-frame snapshot of one BIOS pad buffer into PadState.
* Decoder (branch ladder on raw[0] status + raw[1] id):
* 1. raw[0] == 0xFF -> Disconnected (buttons=0, axes=0x80)
* 2. raw[0]==0 && raw[1]==0 -> Pending (buttons=0, axes=0x80)
* 3. raw[1] == 0x41 -> Digital (buttons normalized; axes=0x80)
* 4. raw[1] == 0x53 -> AnalogStick (buttons normalized; axes from raw[4..7])
* 5. raw[1] in 0x7x -> AnalogPad (buttons normalized; axes from raw[4..7])
* 6. else -> Unsupported (buttons=0, axes=0x80)
*
* Buttons normalization: byte_swap16((~raw_buttons) & 0xFFFF).
* raw_buttons = load_half_u(raw, 2) = raw[2] | (raw[3] << 8).
* byte_swap16(x) = (x >> 8) | (x << 8); nor(x, R_0) = ~x. store_half truncates to 16 bits so the upper-16 mask is implicit in the store.
*
* Register use (atom-local; no wave-context touched):
* R_T0 = raw base (kept throughout; axes loads read raw[4..7] from R_T0)
* R_T1 = state base (kept throughout; all stores go through R_T1)
* R_T2 = raw[0] status (alive across the disc/pending/id dispatch, then dead)
* R_T3 = raw[1] id (alive across the id dispatch, then dead)
* R_T4 = scratch (shifts, compares, immediate loads, store values)
* R_T5 = scratch (parallel lui+ori for the 0x80808080 axes constant + byte-swap target)
*/
enum {
R_PadRaw = R_T0 atom_reg atom_type(U1),
R_PadState = R_T1 atom_reg,
R_RawStatus = R_T2 atom_reg,
R_RawId = R_T3 atom_reg,
};
typedef Struct_(Binds_PadBiosSnapshot) {
PadBiosRaw* raw;
PadState* state;
};
internal MipsAtom_(pad_bios_snapshot) atom_info(atom_bind(Binds_PadBiosSnapshot)
, atom_reads( R_PadRaw, R_PadState, R_RawStatus, R_RawId, R_T4, R_T5, R_TapePtr)
, atom_writes(R_PadRaw, R_PadState, R_RawStatus, R_RawId, R_T4, R_T5, R_TapePtr)
) {
/* === Bind consumption: T0 = raw, T1 = state, advance R_TapePtr by 8. */
load_word(R_PadRaw, R_TapePtr, O_(Binds_PadBiosSnapshot,raw)),
load_word(R_PadState, R_TapePtr, O_(Binds_PadBiosSnapshot,state)),
add_ui_self( R_TapePtr, S_(Binds_PadBiosSnapshot)),
/* === Read raw[0] (status) + raw[1] (id) */
load_byte_u(R_RawStatus, R_PadRaw, 0),
load_byte_u(R_RawId, R_PadRaw, 1),
atom_label(snap_root) /* === Case 1: Disconnected (status == 0xFF). */
add_ui(R_T4, R_0, 0xFF), branch_ne(R_RawStatus, R_T4, atom_offset(snap_root, skip_disconnected)),
/* BD-slot: pre-compute PadStatus_Disconnected. Branch reads R_T4=0xFF in EX before this WB completes.
* If branch NOT taken (fall through to pending/id_dispatch), R_T4 is overwritten by the next case body's add_ui — harmless. */
atom_label(disconnected) /* === Disconnected body. */
/* R_T4 = PadStatus_Disconnected from snap_root BD-slot. */
store_word(R_T4, R_PadState, O_(PadState,status)),
store_half(R_0, R_PadState, O_(PadState,buttons)),
/* axes = 0x80808080 (centered) — single sw writes the 4-byte axes block at offset 8 (left_x, left_y, right_x, right_y). */
load_upper_i(R_T4, 0x8080), or_i_self(R_T4, 0x8080),
store_word( R_T4, R_PadState, O_(PadState,left_x)),
store_byte( R_RawId, R_PadState, O_(PadState,id)),
jump_rel(atom_offset(disconnected, snap_end)),
/* BD-slot: load next atom's entry point (replaces the nop).
* The unconditional branch always jumps to snap_end, where mac_yield_tail()
* transfers control to R_AtomJmp without re-loading it. */
mac_yield_load(),
atom_label(skip_disconnected)
/* === Case 2: Pending (status == 0 && id == 0)
* Combined check: if (status | id) != 0 then skip to id_dispatch.
* Falls through to the Pending case only when both are zero. */
or_u_self(R_RawStatus, R_RawId), branch_ne(R_RawStatus, R_0, atom_offset(case_2, id_dispatch)),
/* BD-slot: pre-compute PadStatus_Pending. Branch reads R_RawStatus in EX before this WB completes.
* If branch NOT taken (fall through to id_dispatch), R_T4 is overwritten by the digital/analog body add_ui — harmless. */
atom_label(pending) /* === Pending body */
/* R_T4 = PadStatus_Pending from case_2 BD-slot. */
store_word(R_T4, R_PadState, O_(PadState,status)),
store_half(R_0, R_PadState, O_(PadState,buttons)),
/* axes = 0x80808080 (centered) — single sw writes the 4-byte axes block at offset 8 (left_x, left_y, right_x, right_y). */
load_upper_i(R_T4, 0x8080), or_i_self(R_T4, 0x8080),
store_word( R_T4, R_PadState, O_(PadState,left_x)),
store_byte( R_RawId, R_PadState, O_(PadState,id)),
jump_rel(atom_offset(pending, snap_end)),
mac_yield_load(),
atom_label(id_dispatch) /* === Case 3-6: ID dispatch */
add_ui(R_T4, R_0, 0x41), branch_ne(R_RawId, R_T4, atom_offset(id_dispatch, try_analog_stick)),
/* BD-slot: pre-compute PadStatus_Digital. Branch reads R_RawId in EX before this WB completes.
* If branch NOT taken (fall through to try_analog_stick), R_T4 is overwritten by the analog body add_ui. */
/* === Digital body (status, buttons normalize, axes=0x80, id, branch. */
/* R_T4 = PadStatus_Digital from id_dispatch BD-slot. */
store_word( R_T4, R_PadState, O_(PadState,status)),
load_half_u(R_T4, R_PadRaw, 2 * S_(U1)),
/* Fill R_T4's load-delay slot with the 0x80808080 axes constant into R_T5
* (R_T5 is dead on this path; it's only consumed at the analog_pad range check). */
load_upper_i(R_T5, 0x8080), or_i_self(R_T5, 0x8080),
nor_u( R_T4, R_T4, R_0), /* raw_buttons is already in host bit order; no swap needed */
store_half( R_T4, R_PadState, O_(PadState,buttons)),
/* axes = 0x80808080 (centered) — single sw writes the 4-byte axes block at offset 8 (left_x, left_y, right_x, right_y). */
store_word( R_T5, R_PadState, O_(PadState,left_x)),
add_ui( R_T4, R_0, 0x41),
store_byte( R_T4, R_PadState, O_(PadState,id)),
jump_rel(atom_offset(id_dispatch, snap_end)),
mac_yield_load(),
atom_label(try_analog_stick) /* === Case 4: AnalogStick (id == 0x53)*/
add_ui(R_T4, R_0, 0x53), branch_ne(R_RawId, R_T4, atom_offset(try_analog_stick, try_analog_pad)),
/* BD-slot: pre-compute PadStatus_AnalogStick. Branch reads R_RawId in EX before this WB completes.
* If branch NOT taken (fall through to try_analog_pad), R_T4 is overwritten by the analog_pad body add_ui. */
atom_label(analog_stick) /* === AnalogStick body
* Axes are loaded as two halfwords: raw[6..7] → left_xy (sh at offset 8), raw[4..5] → right_xy (sh at offset 10).
* R_T5 holds left_xy / id-value in turn (it's dead on this path — only consumed at the analog_pad range check). */
/* R_T4 = PadStatus_AnalogStick from try_analog_stick BD-slot. */
store_word( R_T4, R_PadState, O_(PadState,status)),
load_half_u( R_T4, R_PadRaw, 2 * S_(U1)), /* R_T4 = raw_buttons */
load_half_u( R_T5, R_PadRaw, 6 * S_(U1)), /* R_T5 = left_xy; fills R_T4's load-delay slot (doesn't read R_T4) */
nor_u( R_T4, R_T4, R_0), /* R_T4 = ~raw_buttons */
store_half( R_T4, R_PadState, O_(PadState,buttons)),
load_half_u( R_T4, R_PadRaw, 4 * S_(U1)), /* R_T4 = right_xy; fills R_T5's load-delay slot */
store_half( R_T5, R_PadState, O_(PadState,left_x)), /* R_T5 settled, store left_xy */
store_half( R_T4, R_PadState, O_(PadState,right_x)),
add_ui( R_T5, R_0, 0x53), /* R_T5 = id value (clobbers left_xy, already stored) */
store_byte( R_T5, R_PadState, O_(PadState,id)),
jump_rel(atom_offset(analog_stick, snap_end)),
mac_yield_load(),
atom_label(try_analog_pad) /* === Case 5-6: AnalogPad (id & 0xF0 == 0x70) */
and_i( R_T4, R_RawId, 0xF0),
add_ui( R_T5, R_0, 0x70),
branch_ne(R_T4, R_T5, atom_offset(try_analog_pad, try_unsupported)),
/* BD-slot: pre-compute PadStatus_AnalogPad. Branch reads R_T4 in EX before this WB completes.
* If branch NOT taken (fall through to try_unsupported), R_T4 is overwritten by the unsupported body add_ui. */
atom_label(analog_pad) /* === AnalogPad body
* Same shape as AnalogStick with AnalogPad status. R_T5 holds left_xy (it's dead on this path). */
/* R_T4 = PadStatus_AnalogPad from try_analog_pad BD-slot. */
store_word( R_T4, R_PadState, O_(PadState,status)),
load_half_u(R_T4, R_PadRaw, 2 * S_(U1)), /* R_T4 = raw_buttons */
load_half_u(R_T5, R_PadRaw, 6 * S_(U1)), /* R_T5 = left_xy; fills R_T4's load-delay slot */
nor_u( R_T4, R_T4, R_0), /* R_T4 = ~raw_buttons */
store_half( R_T4, R_PadState, O_(PadState,buttons)),
load_half_u(R_T4, R_PadRaw, 4 * S_(U1)), /* R_T4 = right_xy; fills R_T5's load-delay slot */
store_half( R_T5, R_PadState, O_(PadState,left_x)), /* R_T5 settled, store left_xy */
store_half( R_T4, R_PadState, O_(PadState,right_x)),
store_byte( R_RawId, R_PadState, O_(PadState,id)),
jump_rel(atom_offset(analog_pad, snap_end)),
mac_yield_load(),
atom_label(try_unsupported) /* === Case 7: Unsupported — fall through from the AnalogPad range-check miss. */
add_ui( R_T4, R_0, PadStatus_Unsupported),
store_word(R_T4, R_PadState, O_(PadState,status)),
store_half(R_0, R_PadState, O_(PadState,buttons)),
/* axes = 0x80808080 (centered) — single sw writes the 4-byte axes block at offset 8 (left_x, left_y, right_x, right_y). */
load_upper_i(R_T4, 0x8080), or_i_self(R_T4, 0x8080),
store_word( R_T4, R_PadState, O_(PadState,left_x)),
add_ui( R_T4, R_0, 0xFF), /* 0xFF sentinel: "unknown id" */
store_byte( R_T4, R_PadState, O_(PadState,id)),
/* Fall through to snap_end. */
atom_label(no_jump_fallthrough)
mac_yield_load(),
atom_label(snap_end)
/* NOT mac_yield() — R_AtomJmp was already loaded in the BD-slot of the case-exit branch. */
mac_yield_tail(),
};
#pragma endregion Baked Atoms
+73
View File
@@ -0,0 +1,73 @@
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "dsl.h"
#endif
/* PSX button bit positions — 1:1 with PSX-SPX docs at docs/psx-spx/docs/controllersandmemorycards.md:405-421.
* Wire is active-low (0 = pressed).
* The decoder atom computes buttons = (~raw_buttons) & 0xFFFF; the active-low-to-active-high inversion is applied bit-by-bit. */
enum {
Bit_(Pad_Select, 0),
Bit_(Pad_L3, 1),
Bit_(Pad_R3, 2),
Bit_(Pad_Start, 3),
Bit_(Pad_Up, 4),
Bit_(Pad_Right, 5),
Bit_(Pad_Down, 6),
Bit_(Pad_Left, 7),
Bit_(Pad_L2, 8),
Bit_(Pad_R2, 9),
Bit_(Pad_L1, 10),
Bit_(Pad_R1, 11),
Bit_(Pad_Triangle, 12),
Bit_(Pad_Circle, 13),
Bit_(Pad_Cross, 14),
Bit_(Pad_Square, 15),
};
enum {
PadId_Offset = 4,
Pad0 = 0 << PadId_Offset,
Pad1 = 1 << PadId_Offset,
};
#define pad0_(btn_id) (btn_id << Pad0)
#define pad1_(btn_id) (btn_id << Pad1)
/* ============================================================
* BIOS pad-buffer subsystem: docs/psx-spx/docs/kernelbios.md (B(12h) + B(13h))
* ============================================================ */
enum {
PAD_BIOS_RAW_SIZE = 0x22,
};
typedef Struct_(PadBiosRaw) {
U1 bytes[PAD_BIOS_RAW_SIZE];
};
typedef Enum_(U4, PadStatus) {
PadStatus_Disconnected,
PadStatus_Digital,
PadStatus_AnalogStick,
PadStatus_AnalogPad,
PadStatus_Unsupported,
PadStatus_Pending,
PadStatus_Invalid,
};
/* PadState — per-port normalized runtime state.
* Field order is chosen so that the 4 axes (left_x, left_y, right_x, right_y)
* form a contiguous 4-byte block at offset 8, allowing a single `store_word` to clear-or-write all 4 axes in one MIPS instruction.
* The struct size stays 12 bytes (unchanged from the prior order,
* which left the C compiler to insert 1 byte of trailing pad to reach the 4-byte struct alignment). */
typedef Struct_(PadState) {
PadStatus status; /* offset 0, size 4 (U4) */
U2 buttons; /* offset 4, size 2 */
U1 id; /* offset 6, size 1 */
U1 pad; /* offset 7, size 1 — explicit pad to align the axes block */
U1 left_x; /* offset 8, size 1 — store_word target (4-byte aligned) */
U1 left_y; /* offset 9, size 1 */
U1 right_x; /* offset 10, size 1 */
U1 right_y; /* offset 11, size 1 */
};
+7
View File
@@ -0,0 +1,7 @@
#ifdef INTELLISENSE_DIRECTIVES
# include "gen/macs.h"
# include "gen/offsets.h"
# include "psyq.h"
#endif
ATOM_FILE_DEBUGGER_LINE_MARKER(pysq_atom_c);
+103
View File
@@ -0,0 +1,103 @@
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "dsl.h"
# include "math.h"
# include "gp.h"
#endif
typedef Struct_(DrawEnv_Packed) { U4 tag; U4 code[15]; };
typedef Struct_(DrawEnv) {
Rect_S2 clip_area;
V2_S2 drawing_offset[2];
Rect_S2 texture_window;
S2 texture_page;
B1 flag_dither;
B1 flag_draw_on_display;
B1 enable_auto_clear;
RGB8 initial_bg_color;
DrawEnv_Packed dr_env; // reserved
};
typedef Struct_(DisplayEnv) {
Rect_S2 display_area;
Rect_S2 screen;
B1 vinterlace;
B1 color24;
B1 pad0;
B1 pad1;
};
typedef Array_(DrawEnv, 2);
typedef Array_(DisplayEnv, 2);
typedef Struct_(DoubleBuffer) {
A2_DrawEnv draw;
A2_DisplayEnv display;
};
DisplayEnv* displayenv_init(DisplayEnv* env, S4 x, S4 y, S4 w, S4 h) asm("SetDefDispEnv");
DrawEnv* drawenv_init (DrawEnv* env, S4 x, S4 y, S4 w, S4 h) asm("SetDefDrawEnv");
DisplayEnv* displayenv_put(DisplayEnv* env) asm("PutDispEnv");
DrawEnv* drawenv_put (DrawEnv* env) asm("PutDrawEnv");
U4 geom_init(void) asm("InitGeom");
void geom_set_offset(U4 x, U4 y) asm("SetGeomOffset");
void geom_set_screen(U4 h) asm("SetGeomScreen");
U4* orderingtbl_clear_reverse(U4* ot, U4 len) asm("ClearOTagR");
U4 reset_graph(U4 mode) asm("ResetGraph");
void set_display_enabled(U4 mask) asm("SetDispMask");
U4 draw_sync(U4 mode) asm("DrawSync");
U4 vsync(U4 mode) asm("VSync");
void draw_orderingtbl(U4* buf) asm("DrawOTag");
typedef Struct_(Tile) {
U4 tag;
RGB8 color;
B1 code;
Rect_S2 rect;
};
/*
Linear Algebra
*/
M3_S2* m3s2_rotation (V3_S2* vec, M3_S2* mat) asm("RotMatrix");
M3_S2* m3s2_translation(M3_S2* mat, V3_S4* vec) asm("TransMatrix");
M3_S2* m3s2_scale (M3_S2* mat, V3_S4* vec) asm("ScaleMatrix");
// Rotation, Translation, Perspective
S4 rtp_v3s2_raw(V3_S2* vec, S4* xy, S4* pp, S4* flag) asm("RotTransPers");
FI_ S4 rtp_v3s2(V3_S2* vec, V2_S2* xy, A2_S2* pp, S4* flag) { return rtp_v3s2_raw(vec, C_(S4*R_, & xy->x), C_(S4*R_, pp), r_(flag)); }
S4 rtp_avg_nclip_a3_v3s2_raw(V3_S2* v0, V3_S2* v1, V3_S2* v2, S4* xy1, S4* xy2, S4* xy3, S4* pp, S4* otz, S4* flag) asm("RotAverageNclip3");
FI_ S4 rtp_avg_nclip_a3_v3s2(
V3_S2* v0, V3_S2* v1, V3_S2* v2,
V2_S2* xy0, V2_S2* xy1, V2_S2* xy2,
A2_S2* pp, S4* otz, S4* flag
){
return rtp_avg_nclip_a3_v3s2_raw(
v0, v1, v2,
C_(S4*R_, xy0), C_(S4*R_, xy1), C_(S4*R_, xy2),
C_(S4*R_, pp), C_(S4*R_, otz), C_(S4*R_, flag)
);
}
S4 rtp_avg_nclip_a4_v3s2_raw(V3_S2* v0, V3_S2* v1, V3_S2* v2, V3_S2* v3, S4* xy1, S4* xy2, S4* xy3, S4* xy4, S4* pp, S4* otz, S4* flag) asm("RotAverageNclip4");
FI_ S4 rtp_avg_nclip_a4_v3s2(
V3_S2* v0, V3_S2* v1, V3_S2* v2, V3_S2* v3,
V2_S2* xy0, V2_S2* xy1, V2_S2* xy2, V2_S2* xy3,
A2_S2* pp, S4* otz, S4* flag
){
return rtp_avg_nclip_a4_v3s2_raw(
v0, v1, v2, v3,
C_(S4*R_, xy0), C_(S4*R_, xy1), C_(S4*R_, xy2), C_(S4*R_, xy3),
C_(S4*R_, pp), C_(S4*R_, otz), C_(S4*R_, flag)
);
}
void gte_matrix_set_rotation (M3_S2* mat) asm("SetRotMatrix");
void gte_matrix_set_translation(M3_S2* mat) asm("SetTransMatrix");
+60
View File
@@ -0,0 +1,60 @@
// word_count.metadata.h
// Single source of truth for instruction-word counts.
// Used by C (to define compile-time constants) AND Python (to count positions).
//
// Format: WORD_COUNT(MACRO_NAME, COUNT)
// One line per macro that appears in your atom sources.
//
// This file is encoding-macros-only.
// The auto-generated component macros (mac_X) live in the source directory's own gen/macs.h (per-directory aggregation; included separately by the unity build).
// The unity build should include THIS file and the .macs.h file in the same TU, with both wrapped
// (or the include guard order handled) to avoid WORD_COUNT redeclaration.
//
// To regenerate: hand-count the instructions in each macro definition.
// (You'll only need to do this once per macro — they don't change often.)
#define WORD_COUNT(name, count) enum { words_##name = (count) };
WORD_COUNT(nop, 1)
WORD_COUNT(load_upper_i, 1)
WORD_COUNT(jump_reg, 1)
WORD_COUNT(jump_link, 1)
WORD_COUNT(call_reg, 1)
WORD_COUNT(call_addr, 1)
WORD_COUNT(branch_le_zero, 1)
WORD_COUNT(branch_equal, 1)
WORD_COUNT(branch_ne, 1)
WORD_COUNT(add_ui, 1)
WORD_COUNT(set_lt_u, 1)
WORD_COUNT(set_lt_s, 1)
WORD_COUNT(set_lt_si, 1)
WORD_COUNT(set_lt_ui, 1)
WORD_COUNT(load_word, 1)
WORD_COUNT(load_half_u, 1)
WORD_COUNT(load_byte_u, 1)
WORD_COUNT(store_word, 1)
WORD_COUNT(store_byte, 1)
WORD_COUNT(add_ui_self, 1)
WORD_COUNT(add_u_self, 1)
WORD_COUNT(add_u, 1)
WORD_COUNT(or_i, 1)
WORD_COUNT(or_i_self, 1)
WORD_COUNT(or_u, 1)
WORD_COUNT(or_u_self, 1)
WORD_COUNT(nor_u, 1)
WORD_COUNT(shift_lleft, 1)
WORD_COUNT(shift_lleft_self, 1)
WORD_COUNT(shift_lright, 1)
WORD_COUNT(shift_aright, 1)
WORD_COUNT(mask_upper, 2)
WORD_COUNT(gte_mv_from_data_r, 1)
WORD_COUNT(gte_mv_from_ctrl_r, 1)
WORD_COUNT(gte_mv_to_data_r, 1)
WORD_COUNT(gte_mv_to_ctrl_r, 1)
WORD_COUNT(gte_sw, 1)
WORD_COUNT(gte_cmdw_rtpt, 1)
WORD_COUNT(gte_cmdw_nclip, 1)
WORD_COUNT(gte_avg_sort_z3, 1)
WORD_COUNT(sub_u, 1)
WORD_COUNT(nop2, 2)
#undef WORD_COUNT
+15 -15
View File
@@ -17,19 +17,19 @@ enum {
};
typedef U4 OrderingTable_Buffer[OrderingTbl_Len];
typedef def_farray(OrderingTable_Buffer, 2);
typedef Array_(OrderingTable_Buffer, 2);
typedef B1 PrimitiveBuffer[PrimitiveBuff_Len];
typedef def_farray(PrimitiveBuffer, 2);
typedef def_struct(PrimitiveArena) {
typedef Array_(PrimitiveBuffer, 2);
typedef Struct_(PrimitiveArena) {
A2_PrimitiveBuffer buf;
U4 used;
};
#define Cube_num_verts 8
typedef def_farray(V3_S2, Cube_num_verts);
typedef Array_(V3_S2, Cube_num_verts);
#define Cube_num_faces 6
typedef def_farray(V4_S2, Cube_num_faces);
typedef Array_(V4_S2, Cube_num_faces);
void ent_cube128_init(A8_V3_S2* verts, A6_V4_S2* faces) {
memory_copy(verts, & (A8_V3_S2) {
{ -128, -128, -128 },
@@ -40,7 +40,7 @@ void ent_cube128_init(A8_V3_S2* verts, A6_V4_S2* faces) {
{ 128, 128, -128 },
{ 128, 128, 128 },
{ -128, 128, 128 }
}, size_of(A8_V3_S2) );
}, S_(A8_V3_S2) );
memory_copy(faces, & (A6_V4_S2) {
{ 3, 2, 0, 1 },
{ 0, 1, 4, 5 },
@@ -48,10 +48,10 @@ void ent_cube128_init(A8_V3_S2* verts, A6_V4_S2* faces) {
{ 1, 2, 5, 6 },
{ 2, 3, 6, 7 },
{ 3, 0, 7, 4 },
}, size_of(A6_V4_S2) );
}, S_(A6_V4_S2) );
return;
}
typedef def_struct(Ent_Cube) {
typedef Struct_(Ent_Cube) {
V3_S4 accel;
V3_S4 vel;
V3_S4 pos;
@@ -62,22 +62,22 @@ typedef def_struct(Ent_Cube) {
};
#define Floor_num_verts 4
typedef def_farray(V3_S2, Floor_num_verts);
typedef Array_(V3_S2, Floor_num_verts);
#define Floor_num_faces 2
typedef def_farray(V3_S2, Floor_num_faces);
typedef Array_(V3_S2, Floor_num_faces);
void ent_floor_init(A4_V3_S2* verts, A2_V3_S2* faces) {
memory_copy(verts, &(A4_V3_S2) {
{ -900, 0, -900 },
{ -900, 0, 900 },
{ 900, 0, -900 },
{ 900, 0, 900 },
}, size_of(A8_V3_S2));
}, S_(A8_V3_S2));
memory_copy(faces, & (A2_V3_S2) {
{ 0, 1, 2 },
{ 1, 3, 2 },
}, size_of(A2_V3_S2));
}, S_(A2_V3_S2));
};
typedef def_struct(Ent_Floor) {
typedef Struct_(Ent_Floor) {
V3_S4 accel;
V3_S4 pos;
V3_S4 scale;
@@ -86,7 +86,7 @@ typedef def_struct(Ent_Floor) {
A2_V3_S2 faces;
};
typedef def_struct(SMemory) {
typedef Struct_(SMemory) {
DoubleBuffer screen_buf;
A2_OrderingTable_Buffer ordering_tbl;
PrimitiveArena primitives;
@@ -108,7 +108,7 @@ B1* prim__alloc(U4 type_width, Str8 type_name) {
pa->used += type_width;
return next;
}
#define prim_alloc(type) (type*)prim__alloc(size_of(type), txt( stringify(type)))
#define prim_alloc(type) (type*)prim__alloc(S_(type), slit( stringify(type)))
void gp_screen_init_c11(DoubleBuffer* screen_buf, S2* active_buf_id)
{
+17 -17
View File
@@ -5,8 +5,8 @@
# include "duffle/gp.h"
#endif
typedef def_struct(DrawEnv_Packed) { U4 tag; U4 code[15]; };
typedef def_struct(DrawEnv) {
typedef Struct_(DrawEnv_Packed) { U4 tag; U4 code[15]; };
typedef Struct_(DrawEnv) {
Rect_S2 clip_area;
A2_S2 drawing_offset;
Rect_S2 texture_window;
@@ -17,7 +17,7 @@ typedef def_struct(DrawEnv) {
RGB8 initial_bg_color;
DrawEnv_Packed dr_env; // reserved
};
typedef def_struct(DisplayEnv) {
typedef Struct_(DisplayEnv) {
Rect_S2 display_area;
Rect_S2 screen;
B1 vinterlace;
@@ -25,9 +25,9 @@ typedef def_struct(DisplayEnv) {
B1 pad0;
B1 pad1;
};
typedef def_farray(DrawEnv, 2);
typedef def_farray(DisplayEnv, 2);
typedef def_struct(DoubleBuffer) {
typedef Array_(DrawEnv, 2);
typedef Array_(DisplayEnv, 2);
typedef Struct_(DoubleBuffer) {
A2_DrawEnv draw;
A2_DisplayEnv display;
};
@@ -58,7 +58,7 @@ U4 vsync(U4 mode) __asm__("VSync");
void draw_orderingtbl(U4* buf) __asm__("DrawOTag");
typedef def_struct(PolyTag) {
typedef Struct_(PolyTag) {
U4 addr: 24;
U4 len: 8;
RGB8 color;
@@ -106,7 +106,7 @@ typedef def_struct(PolyTag) {
// #define setLineF4(p) set_len(p, 6), set_code(p, 0x4c),(p)->pad = 0x55555555
// #define setLineG4(p) set_len(p, 9), set_code(p, 0x5c),(p)->pad = 0x55555555, (p)->p2 = 0, (p)->p3 = 0
typedef def_struct(Poly_F3) {
typedef Struct_(Poly_F3) {
U4 tag;
RGB8 color;
B1 code;
@@ -120,14 +120,14 @@ typedef def_struct(Poly_F3) {
};
};
typedef def_struct(Poly_G3) {
typedef Struct_(Poly_G3) {
U4 tag; RGB8 c0; B1 code;
V2_S2 p0; RGB8 c1; B1 pad1;
V2_S2 p1; RGB8 c2; B1 pad2;
V2_S2 p2;
};
typedef def_struct(Poly_F4) {
typedef Struct_(Poly_F4) {
U4 tag;
RGB8 color;
B1 code;
@@ -142,7 +142,7 @@ typedef def_struct(Poly_F4) {
};
};
typedef def_struct(Poly_G4) {
typedef Struct_(Poly_G4) {
U4 tag; RGB8 c0; B1 code;
V2_S2 p0; RGB8 c1; B1 pad1;
V2_S2 p1; RGB8 c2; B1 pad2;
@@ -150,7 +150,7 @@ typedef def_struct(Poly_G4) {
V2_S2 p3;
};
typedef def_struct(Tile) {
typedef Struct_(Tile) {
U4 tag;
RGB8 color;
B1 code;
@@ -169,7 +169,7 @@ M3_S2* m3s2_scale (M3_S2* mat, V3_S4* vec) __asm__("ScaleMatrix");
// Rotation, Translation, Perspective
S4 rtp_v3s2_raw(V3_S2* vec, S4* xy, S4* pp, S4* flag) __asm__("RotTransPers");
FI_ S4 rtp_v3s2(V3_S2* vec, V2_S2* xy, A2_S2* pp, S4* flag) { return rtp_v3s2_raw(vec, cast(S4*R_, & xy->x), cast(S4*R_, pp), r_(flag)); }
FI_ S4 rtp_v3s2(V3_S2* vec, V2_S2* xy, A2_S2* pp, S4* flag) { return rtp_v3s2_raw(vec, C_(S4*R_, & xy->x), C_(S4*R_, pp), r_(flag)); }
S4 rtp_avg_nclip_a3_v3s2_raw(V3_S2* v0, V3_S2* v1, V3_S2* v2, S4* xy1, S4* xy2, S4* xy3, S4* pp, S4* otz, S4* flag) __asm__("RotAverageNclip3");
FI_ S4 rtp_avg_nclip_a3_v3s2(
@@ -179,8 +179,8 @@ FI_ S4 rtp_avg_nclip_a3_v3s2(
){
return rtp_avg_nclip_a3_v3s2_raw(
v0, v1, v2,
cast(S4*R_, xy0), cast(S4*R_, xy1), cast(S4*R_, xy2),
cast(S4*R_, pp), cast(S4*R_, otz), cast(S4*R_, flag)
C_(S4*R_, xy0), C_(S4*R_, xy1), C_(S4*R_, xy2),
C_(S4*R_, pp), C_(S4*R_, otz), C_(S4*R_, flag)
);
}
@@ -192,8 +192,8 @@ FI_ S4 rtp_avg_nclip_a4_v3s2(
){
return rtp_avg_nclip_a4_v3s2_raw(
v0, v1, v2, v3,
cast(S4*R_, xy0), cast(S4*R_, xy1), cast(S4*R_, xy2), cast(S4*R_, xy3),
cast(S4*R_, pp), cast(S4*R_, otz), cast(S4*R_, flag)
C_(S4*R_, xy0), C_(S4*R_, xy1), C_(S4*R_, xy2), C_(S4*R_, xy3),
C_(S4*R_, pp), C_(S4*R_, otz), C_(S4*R_, flag)
);
}
@@ -1,19 +0,0 @@
// Auto-generated by tape_atom_offset_gen.meta.lua — DO NOT EDIT
// Source: C:\projects\Pikuma\ps1\code\gte_hello\hello_gte_tape.c
#pragma once
#pragma region hello_gte_tape
// --- atom: floor_tri (51 words) ---
#define _atom_offset_culling_floor_tri_exit 17
#define _atom_offset_bounds_chk_floor_tri_exit 3
enum {
atom_offset_culling_floor_tri_exit = _atom_offset_culling_floor_tri_exit,
atom_offset_bounds_chk_floor_tri_exit = _atom_offset_bounds_chk_floor_tri_exit,
};
#pragma endregion hello_gte_tape
-233
View File
@@ -1,233 +0,0 @@
#ifdef INTELLISENSE_DIRECTIVES
# include "duffle/lottes_tape.h"
# include "duffle/atom_dsl.h"
# include "hello_gte.h"
# include "tape_atom.metadata.h"
# include "gen/hello_gte_tape.offsets.h"
#endif
#pragma region MACs (Mips Atom components)
/* Words: 3; High: 0x20/B, Low: G/R */
#define mac_format_f3_color(color_hi, color_lo) \
load_ui(R_AT, color_hi), or_i(R_AT, R_AT, color_lo) \
, store_word(R_AT, R_PrimCursor, O_(Poly_F3,color)) \
/* Words: 3 */
#define mac_gte_store_f3() \
gte_sw(C2_SXY0, R_PrimCursor, O_(Poly_F3,p0)) \
, gte_sw(C2_SXY1, R_PrimCursor, O_(Poly_F3,p1)) \
, gte_sw(C2_SXY2, R_PrimCursor, O_(Poly_F3,p2))
#pragma endregion MACs
#pragma region Baked Atoms
typedef Struct_(Binds_CubeTri) {
U4 PrimCursor;
U4 FaceCursor;
U4 VertBase;
U4 OtBase;
};
internal MipsAtom_(rbind_cube_tri) {
/* Pop 4 arguments from the tape directly into the workspace registers */
load_word(R_PrimCursor, R_TapePtr, O_(Binds_CubeTri,PrimCursor)),
load_word(R_FaceCursor, R_TapePtr, O_(Binds_CubeTri,FaceCursor)),
load_word(R_VertBase, R_TapePtr, O_(Binds_CubeTri,VertBase)),
load_word(R_OtBase, R_TapePtr, O_(Binds_CubeTri,OtBase)),
add_ui_1( R_TapePtr, S_(Binds_CubeTri)),
// Note(Ed): This entire thing is argument shuffle?
// TODO(Ed): Eliminate
mac_yield()
};
/* ============================================================================
* cube_tri — Draw one cube face (Gouraud-shaded quad) via the GTE tape pipeline
* ============================================================================
*
* Reads 4 indices from R_FaceCur (V4_S2 = 8 bytes), loads 4 vertices into
* the GTE, runs the PsyQ RotAverageNclip4 sequence, and renders a Poly_G4.
*/
atom_region (cube_tri, REGION_PRIM_ARENA)
atom_group (cube_tri, GROUP_RENDER_PRIMS)
atom_cadence (cube_tri, CADENCE_FRAME)
atom_annot(cube_tri, phase_work,
tape_regs(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase),
tape_regs(R_PrimCursor, R_FaceCursor))
internal
MipsAtom_(cube_tri) {
/* ── 1. Load 4 face indices from R_FaceCur ──────────────────────────── */
load_half_u(R_T0, R_FaceCursor, 0), /* T0 = face->x (vertex 0 index) */
load_half_u(R_T1, R_FaceCursor, 2), /* T1 = face->y (vertex 1 index) */
load_half_u(R_T2, R_FaceCursor, 4), /* T2 = face->z (vertex 2 index) */
load_half_u(R_T3, R_FaceCursor, 6), /* T3 = face->w (vertex 3 index) */
/* ── 2. Load V0, V1, V2 into GTE ────────────────────────────────────── */
/* V0 = verts[face->x] */
shift_ll(R_AT, R_T0, 3), add_u(R_AT, R_AT, R_VertBase),
load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4),
gte_mt(R_V0, C2_VXY0), gte_mt(R_V1, C2_VZ0),
/* V1 = verts[face->y] */
shift_ll(R_AT, R_T1, 3), add_u(R_AT, R_AT, R_VertBase),
load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4),
gte_mt(R_V0, C2_VXY1), gte_mt(R_V1, C2_VZ1),
/* V2 = verts[face->z] */
shift_ll(R_AT, R_T2, 3), add_u(R_AT, R_AT, R_VertBase),
load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4),
gte_mt(R_V0, C2_VXY2), gte_mt(R_V1, C2_VZ2),
/* ── 3. RTPT — transforms V0/V1/V2 → SXY0/SXY1/SXY2 + SZ1/SZ2/SZ3 ─── */
nop, nop, gte_cmdw_rtpt,
/* ── 4. NCLIP — backface culling on SXY0/SXY1/SXY2 (p0,p1,p2) ──────── */
/* MUST be done BEFORE RTPS overwrites SXY0 with p3! */
nop, nop, gte_cmdw_nclip,
nop, nop,
/* ── 5. Cull check: skip format/insert if MAC0 ≤ 0 (backface) ───────── */
gte_mf(R_T0, C2_MAC0),
nop,
branch_le_zero(R_T0, 49), /* Skip 49 if MAC0 ≤ 0 (backface) → cull */
nop, /* BD slot */
/* ── 6. Store p0,p1,p2 to primitive buffer (BEFORE RTPS overwrites) ─── */
store_word(R_0, R_PrimCursor, 0),
/* Word 1: c0 (BGR) + code = 0x38FF00FF (magenta, opcode 0x38) */
load_ui(R_AT, 0x38FF), or_i(R_AT, R_AT, 0x00FF),
store_word(R_AT, R_PrimCursor, 4),
/* Word 2: p0 = SXY0 (stored BEFORE RTPS overwrites it) */
gte_sw(C2_SXY0, R_PrimCursor, 8),
/* Word 3: c1 (BGR) + pad = 0x0000FFFF (yellow) */
load_ui(R_AT, 0x0000), or_i(R_AT, R_AT, 0xFFFF),
store_word(R_AT, R_PrimCursor, 12),
/* Word 4: p1 = SXY1 */
gte_sw(C2_SXY1, R_PrimCursor, 16),
/* Word 5: c2 (BGR) + pad = 0x00FFFF00 (cyan) */
load_ui(R_AT, 0x00FF), or_i(R_AT, R_AT, 0xFF00),
store_word(R_AT, R_PrimCursor, 20),
/* Word 6: p2 = SXY2 */
gte_sw(C2_SXY2, R_PrimCursor, 24),
/* Word 7: c3 (BGR) + pad = 0x0000FF00 (green) */
load_ui(R_AT, 0x0000), or_i(R_AT, R_AT, 0xFF00),
store_word(R_AT, R_PrimCursor, 28),
/* ── 7. Load V3 = verts[face->w] into V0 ─────────────────────────────── */
shift_ll(R_AT, R_T3, 3), add_u(R_AT, R_AT, R_VertBase),
load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4),
gte_mt(R_V0, C2_VXY0), gte_mt(R_V1, C2_VZ0),
/* ── 8. RTPS — transforms V0 (now V3) → SXY0 (p3) + SZ0 ─────────────── */
nop, nop, gte_cmdw_rtps,
/* Word 8: p3 = SXY0 (written AFTER RTPS with V3's screen coords) */
gte_sw(C2_SXY0, R_PrimCursor, 32),
/* ── 9. AVSZ4 — average Z from SZ0/SZ1/SZ2/SZ3 ────────────── */
nop, nop, gte_cmdw_avsz4,
nop, nop,
gte_mf(R_T1, C2_OTZ),
/* ── 10. Bounds check OTZ < 2048 ─────────────────────────────────────── */
add_ui( R_AT, R_0, 2048),
slt_u( R_AT, R_T1, R_AT),
branch_equal(R_AT, R_0, 13), /* Skip 13 → land at add_ui(R_FaceCur,...) */
nop, /* BD slot */
/* ── 11. Insert into Ordering Table (length = 8 for Poly_G4) ─────────── */
mac_insert_ot_tag(R_T1, 0x0800), /* 0x0800 = 8 << 8 = length 8 in tag */
/* ── 12. Advance cursors & yield ─────────────────────────────────────── */
add_ui(R_PrimCursor, R_PrimCursor, 36), /* 9 words × 4 bytes */
add_ui(R_FaceCursor, R_FaceCursor, 8), /* 4 × S2 = 8 bytes */
mac_yield()
};
typedef Struct_(Binds_FloorTri) {
U4 PrimCursor;
U4 FaceCursor;
U4 VertBase;
U4 OtBase;
};
atom_region(rbind_floor_tri, REGION_PRIM_ARENA)
atom_group(rbind_floor_tri, GROUP_RENDER_FLOOR)
atom_cadence(rbind_floor_tri, CADENCE_FRAME)
atom_annot(rbind_floor_tri, phase_bind
, atom_reads()
, atom_writes(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase))
internal
MipsAtom_(rbind_floor_tri) {
/* Pop 4 arguments from the tape directly into the workspace registers */
load_word(R_PrimCursor, R_TapePtr, O_(Binds_FloorTri,PrimCursor)),
load_word(R_FaceCursor, R_TapePtr, O_(Binds_FloorTri,FaceCursor)),
load_word(R_VertBase, R_TapePtr, O_(Binds_FloorTri,VertBase)),
load_word(R_OtBase, R_TapePtr, O_(Binds_FloorTri,OtBase)),
add_ui_1( R_TapePtr, S_(Binds_FloorTri)),
mac_yield()
};
atom_region( floor_tri, REGION_PRIM_ARENA)
atom_group( floor_tri, GROUP_RENDER_FLOOR)
atom_cadence(floor_tri, CADENCE_FRAME)
atom_annot( floor_tri, phase_work,
atom_reads( R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase),
atom_writes(R_PrimCursor, R_FaceCursor))
internal
MipsAtom_(floor_tri) {
mac_load_tri_indices(R_T0, R_T1, R_T2),
mac_load_tri_verts( R_T0, R_T1, R_T2),
nop, nop, gte_cmdw_rotate_translate_perspective_triple,
nop, nop, gte_cmdw_nclip,
nop, nop,
/* Culling (Branch forward if Backface) */
gte_mf(R_T0, C2_MAC0),
nop, branch_le_zero(R_T0, atom_offset(culling, floor_tri_exit)),
nop,
/* Format Primitive */
mac_format_f3_color(0x20FF, 0xFFFF),
mac_gte_store_f3(),
/* Calculate Depth */
nop, nop, gte_avg_sort_z3,
nop, nop, gte_mf(R_T1, C2_OTZ),
/* Bounds Check OTZ < 2048 (Branch forward to skip insertion) */
add_ui( R_AT, R_0, OrderingTbl_Len),
slt_u( R_AT, R_T1, R_AT),
branch_equal(R_AT, R_0, atom_offset(bounds_chk, floor_tri_exit)),
nop,
/* Insert into Ordering Table Linked List */
mac_insert_ot_tag(R_T1, 0x0400),
add_ui_1(R_PrimCursor, S_(Poly_F3)), /* Advance Prim Cursor (5 words) */
// Note(Ed): No bounds checking, should be checked before atom runs.
/* Advance Input Cursor & Yield (Both branch targets land here) */
atom_label(floor_tri_exit)
add_ui_1(R_FaceCursor, S_(S2) * 4), /* Advance Face Cursor (4 * S2 = 8 bytes) */
mac_yield()
};
typedef Struct_(Binds_SyncPrimitiveArena) { U4 used; U4 cursor; };
atom_region( sync_primitive_arena, REGION_PRIM_ARENA)
atom_group( sync_primitive_arena, GROUP_RENDER_FLOOR)
atom_cadence(sync_primitive_arena, CADENCE_FRAME)
atom_annot( sync_primitive_arena, phase_work,
atom_reads( R_TapePtr, R_PrimCursor),
atom_writes(R_TapePtr))
internal MipsAtom_(sync_primitive_arena) {
load_word(R_AT, R_TapePtr, O_(Binds_SyncPrimitiveArena,used)),
load_word(R_T0, R_TapePtr, O_(Binds_SyncPrimitiveArena,cursor)),
add_ui_1( R_TapePtr, S_(Binds_SyncPrimitiveArena)),
/* Calculate byte offset and store directly back to RAM */
sub_u( R_T0, R_PrimCursor, R_T0), // R_T0 = R_PrimCursor - binds.cursor
store_word(R_T0, R_AT, 0), // R_AT[0] = R_T0
mac_yield()
};
#pragma endregion Baked Atoms
-43
View File
@@ -1,43 +0,0 @@
// tape_atom.metadata.h
// Single source of truth for instruction-word counts.
// Used by C (to define compile-time constants) AND Python (to count positions).
//
// Format: WORD_COUNT(MACRO_NAME, COUNT)
// One line per macro that appears in your atom sources.
//
// To regenerate: hand-count the instructions in each macro definition.
// (You'll only need to do this once per macro — they don't change often.)
#define WORD_COUNT(name, count) enum { words_##name = (count) };
WORD_COUNT(nop, 1)
WORD_COUNT(jump_reg, 1)
WORD_COUNT(jump_link, 1)
WORD_COUNT(branch_le_zero, 1)
WORD_COUNT(branch_equal, 1)
WORD_COUNT(add_ui, 1)
WORD_COUNT(slt_u, 1)
WORD_COUNT(load_ui, 1)
WORD_COUNT(load_word, 1)
WORD_COUNT(load_half_u, 1)
WORD_COUNT(store_word, 1)
WORD_COUNT(add_ui_1, 1)
WORD_COUNT(add_u, 1)
WORD_COUNT(or_i, 1)
WORD_COUNT(or_u, 1)
WORD_COUNT(shift_ll, 1)
WORD_COUNT(shift_lr, 1)
WORD_COUNT(gte_mf, 1)
WORD_COUNT(gte_mt, 1)
WORD_COUNT(gte_ct, 1)
WORD_COUNT(gte_sw, 1)
WORD_COUNT(gte_cmdw_rtpt, 1)
WORD_COUNT(gte_cmdw_nclip, 1)
WORD_COUNT(gte_avg_sort_z3, 1)
WORD_COUNT(mac_load_tri_indices, 3)
WORD_COUNT(mac_load_tri_verts, 18)
WORD_COUNT(mac_format_f3_color, 3)
WORD_COUNT(mac_gte_store_f3, 3)
WORD_COUNT(mac_insert_ot_tag, 11)
WORD_COUNT(mac_yield, 4)
#undef WORD_COUNT
+41
View File
@@ -0,0 +1,41 @@
#ifdef INTELLISENSE_DIRECTIVES
#pragma once
#endif
// Auto-generated by ps1_meta.lua — DO NOT EDIT
// Directory: C:\projects\Pikuma\ps1\code\hello_camera/
// source: C:\projects\Pikuma\ps1\code\hello_camera\hello_camera.c
// source: C:\projects\Pikuma\ps1\code\hello_camera\hello_camera.h
// source: C:\projects\Pikuma\ps1\code\hello_camera\hello_camera.atom.c
// Component atoms (MipsAtomComp_(ac_*)) -> macro variants (mac_*)
#ifndef WORD_COUNT
#define WORD_COUNT(name, count) enum { words_##name = (count) };
#endif
#define mac_put_disp_env(reg_transfer, reg_base, port) \
mac_gcmd_push(gp0_word_draw_area_top_left_origin, reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_draw_area_bottom_right_320x240, reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_set_mask_bit(), reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_draw_area_top_left_origin, reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_draw_area_bottom_right_320x240, reg_transfer, reg_base, port)
WORD_COUNT(mac_put_disp_env, 5)
#define mac_put_draw_env(reg_transfer, reg_base, port) \
mac_gcmd_push(gp0_dr_env_tag, reg_transfer, reg_base, port) /* tag (length=15 << 24, addr=0) — packet header for the DR_ENV sequence. The GPU needs this to recognize the next 15 words as a DR_ENV packet and trigger the isbg auto-clear. */ \
, mac_gcmd_push(gp0_word_draw_mode_drawing_allowed, reg_transfer, reg_base, port) /* code[0] DrawMode (dfe=1, dtd=0, tpage=0) */ \
, mac_gcmd_push(gp0_word_set_texture_window(), reg_transfer, reg_base, port) /* code[1] TextureWindow (tw=(0,0)) */ \
, mac_gcmd_push(enc_gp0_draw_area_tl_word(0, ScreenRes_Y), reg_transfer, reg_base, port) /* code[2] DrawArea top-left (clip.x=0, clip.y=ScreenRes_Y=240) */ \
, mac_gcmd_push(gp0_word_draw_area_bottom_right_320x240, reg_transfer, reg_base, port) /* code[3] DrawArea bottom-right (clip.x+w=320, clip.y+h=480) */ \
, mac_gcmd_push(gp0_word_set_draw_offset(), reg_transfer, reg_base, port) /* code[4] DrawOffset (ofs=(0,0)) — bare-cmd word; the GPU uses the current state machine. */ \
, mac_gcmd_push(gp0_word_dr_env_mask(), reg_transfer, reg_base, port) /* code[5] Mask (dtd=0, dfe=1, isbg=1) — 0xE6 cmd + isbg bit. */ \
, mac_gcmd_push(gp0_word_dr_env_bg_color_cmd(1, 7, 7, 7), reg_transfer, reg_base, port) /* code[6] Initial-bg-color + auto-clear (isbg=1, r=7, g=7, b=7). */ \
, mac_gcmd_push(gp0_word_dr_env_draw_mode(1), reg_transfer, reg_base, port) /* code[7] Re-assert DrawMode with isbg=1 (isbg-flag set; the 0xE1 cmd byte plus isbg only). */ /* code[8..10] Padding (NOP — GPU discards; the DR_ENV requires 16 words total). */ \
, mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port) /* code[11..12] TextureWindow bottom-right (tw.x+tw.w=0, tw.y+tw.h=0) — libpsyx emits twice. */ \
, mac_gcmd_push(gp0_word_set_texture_window(), reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_set_texture_window(), reg_transfer, reg_base, port) /* code[13..14] Padding (NOP) — completes the 16-word packet. */ \
, mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port)
WORD_COUNT(mac_put_draw_env, 16)
+50
View File
@@ -0,0 +1,50 @@
// Auto-generated by ps1_meta.lua (passes/offsets.lua) — DO NOT EDIT
// Directory: C:\projects\Pikuma\ps1\code\hello_camera\
// source: C:\projects\Pikuma\ps1\code\hello_camera\hello_camera.c
// source: C:\projects\Pikuma\ps1\code\hello_camera\hello_camera.h
// source: C:\projects\Pikuma\ps1\code\hello_camera\hello_camera.atom.c
#pragma once
#pragma region hello_camera
// --- atom: pad_apply_input (60 words) ---
#define _atom_offset_dpad_left_exit_dpad_left 6
#define _atom_offset_dpad_right_exit_dpad_right 6
#define _atom_offset_dead_zone_low_check_dead_low_active 8
#define _atom_offset_dead_zone_high_check_dead_high_active 15
#define _atom_offset_dead_zone_skip_exit_stick 24
#define _atom_offset_end_low_exit_stick 12
enum {
atom_offset_dpad_left_exit_dpad_left = _atom_offset_dpad_left_exit_dpad_left,
atom_offset_dpad_right_exit_dpad_right = _atom_offset_dpad_right_exit_dpad_right,
atom_offset_dead_zone_low_check_dead_low_active = _atom_offset_dead_zone_low_check_dead_low_active,
atom_offset_dead_zone_high_check_dead_high_active = _atom_offset_dead_zone_high_check_dead_high_active,
atom_offset_dead_zone_skip_exit_stick = _atom_offset_dead_zone_skip_exit_stick,
atom_offset_end_low_exit_stick = _atom_offset_end_low_exit_stick,
};
// --- atom: cube_g4_face (76 words) ---
#define _atom_offset_cull_cube_g4_face_exit 41
#define _atom_offset_bounds_chk_cube_g4_face_exit 24
enum {
atom_offset_cull_cube_g4_face_exit = _atom_offset_cull_cube_g4_face_exit,
atom_offset_bounds_chk_cube_g4_face_exit = _atom_offset_bounds_chk_cube_g4_face_exit,
};
// --- atom: floor_f3_face (58 words) ---
#define _atom_offset_culling_floor_f3_face_exit 25
#define _atom_offset_bounds_chk_floor_f3_face_exit 16
enum {
atom_offset_culling_floor_f3_face_exit = _atom_offset_culling_floor_f3_face_exit,
atom_offset_bounds_chk_floor_f3_face_exit = _atom_offset_bounds_chk_floor_f3_face_exit,
};
#pragma endregion hello_camera
+468
View File
@@ -0,0 +1,468 @@
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "duffle/gen/macs.h"
# include "duffle/gen/offsets.h"
# include "duffle/dsl.atom.h"
# include "duffle/lottes_tape.h"
# include "duffle/mips.h"
# include "duffle/gte.h"
# include "duffle/gp.h"
# include "duffle/pad.h"
# include "duffle/word_count.metadata.h"
# include "duffle/psyq.h"
# include "duffle/math.atom.c"
# include "duffle/mips.atom.c"
# include "duffle/gte.atom.c"
# include "duffle/gp.atom.c"
# include "duffle/psyq.atom.c"
# include "gen/offsets.h"
# include "gen/macs.h"
# include "hello_camera.h"
#endif
ATOM_FILE_DEBUGGER_LINE_MARKER(hello_joypad_atom_c);
#pragma region MACs (Mips Atom components)
FI_ Slice_MipsCode ac_put_disp_env(U4 reg_transfer, U4 reg_base, U2 port)
MipsAtomComp_Proc_(ac_put_disp_env, {
// Emits 5 GP0 commands for buffer 0 (display_area = (0,0,320,240)).
// Sequence per libpsyx PutDispEnv: DrawArea TL → DrawArea BR → Mask → DrawArea TL → DrawArea BR
mac_gcmd_push(gp0_word_draw_area_top_left_origin, reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_draw_area_bottom_right_320x240, reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_set_mask_bit(), reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_draw_area_top_left_origin, reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_draw_area_bottom_right_320x240, reg_transfer, reg_base, port),
})
FI_ Slice_MipsCode ac_put_draw_env(U4 reg_transfer, U4 reg_base, U2 port)
MipsAtomComp_Proc_(ac_put_draw_env, {
/*
* ORIGIN: each code word corresponds to the EXACT value libpsyx's PutDrawEnv function would compute for the same DrawEnv settings.
* References:
* - libpsyx source: `toolchain/psyq-4_7/lib/libgpu.a` (binary, function `PutDrawEnv`)
* - PSX-SPX doc: https://problemkaputt.de/psx-spx.htm#gputdrawingcommands
* - PSYQ SDK: `setdrawenv` / `makelongdr_env` source
* - NOCASH PSX spec: §"GP0(E1h) Draw Mode setting" through §"DR_ENV"
*
* The 16-word format is documented in the PSYQ SDK manual and on NOCASH's PSX-spec.txt. The libpsyx reference is at:
* ./toolchain/psyq-4_7/lib/libgpu.a
* (binary; the PutDrawEnv implementation builds the 16-word DR_ENV from the user's DRAWENV struct and emits it via GP0 GPU commands.)
*
* Word indices (libpsyx PutDrawEnv / SetDrawEnv order):
* tag = (length << 24) | addr — 16-word packet (1 tag + 15 code)
* code[0] = DrawMode (dfe=1, dtd=0, tpage=0) — must come first per libpsyx
* code[1] = TextureWindow (tw=(0,0)) — bare-cmd word; GPU uses current state
* code[2] = DrawArea top-left (clip.x=0, clip.y=240)
* code[3] = DrawArea bottom-right (clip.x+w=320, clip.y+h=480)
* code[4] = DrawOffset (ofs=(0,0)) — bare-cmd word
* code[5] = Mask (dtd=0, dfe=1, isbg=1) — 0xE6 cmd + isbg bit
* code[6] = Initial-bg-color (isbg=1, r=7, g=7, b=7)
* code[7] = DrawMode (isbg=1, tpage=0) — re-asserts DrawMode with isbg
* code[8..10] = padding (NOP) — 3 words to fill the packet
* code[11..12] = TextureWindow bottom-right — defaults to (0,0,0,0)
* code[13..14] = padding (NOP) — completes the 16-word packet
*/
mac_gcmd_push(gp0_dr_env_tag, reg_transfer, reg_base, port), /* tag (length=15 << 24, addr=0) — packet header for the DR_ENV sequence. The GPU needs this to recognize the next 15 words as a DR_ENV packet and trigger the isbg auto-clear. */
mac_gcmd_push(gp0_word_draw_mode_drawing_allowed, reg_transfer, reg_base, port), /* code[0] DrawMode (dfe=1, dtd=0, tpage=0) */
mac_gcmd_push(gp0_word_set_texture_window(), reg_transfer, reg_base, port), /* code[1] TextureWindow (tw=(0,0)) */
mac_gcmd_push(enc_gp0_draw_area_tl_word(0, ScreenRes_Y), reg_transfer, reg_base, port), /* code[2] DrawArea top-left (clip.x=0, clip.y=ScreenRes_Y=240) */
mac_gcmd_push(gp0_word_draw_area_bottom_right_320x240, reg_transfer, reg_base, port), /* code[3] DrawArea bottom-right (clip.x+w=320, clip.y+h=480) */
mac_gcmd_push(gp0_word_set_draw_offset(), reg_transfer, reg_base, port), /* code[4] DrawOffset (ofs=(0,0)) — bare-cmd word; the GPU uses the current state machine. */
mac_gcmd_push(gp0_word_dr_env_mask(), reg_transfer, reg_base, port), /* code[5] Mask (dtd=0, dfe=1, isbg=1) — 0xE6 cmd + isbg bit. */
mac_gcmd_push(gp0_word_dr_env_bg_color_cmd(1, 7, 7, 7), reg_transfer, reg_base, port), /* code[6] Initial-bg-color + auto-clear (isbg=1, r=7, g=7, b=7). */
mac_gcmd_push(gp0_word_dr_env_draw_mode(1), reg_transfer, reg_base, port), /* code[7] Re-assert DrawMode with isbg=1 (isbg-flag set; the 0xE1 cmd byte plus isbg only). */
/* code[8..10] Padding (NOP — GPU discards; the DR_ENV requires 16 words total). */
mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port),
/* code[11..12] TextureWindow bottom-right (tw.x+tw.w=0, tw.y+tw.h=0) — libpsyx emits twice. */
mac_gcmd_push(gp0_word_set_texture_window(), reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_set_texture_window(), reg_transfer, reg_base, port),
/* code[13..14] Padding (NOP) — completes the 16-word packet. */
mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port),
})
#pragma endregion MACs
#pragma region Baked Atoms
enum {
R_ScreenX = R_T5 atom_reg atom_type(U2),
R_ScreenY = R_T6 atom_reg atom_type(U2),
R_ScreenBuf = R_T7 atom_reg, /* Caller-pinned: & smem.screen_buf */
#define R_ScreenBuf_Code R_T7_Code
};
//screen_env_init. Mirrors the libpsyx's SetDefDispEnv + SetDefDrawEnv + the manual enable_auto_clear / initial_bg_color writes.
internal MipsAtom_(screen_env_init) atom_info(atom_phase(screen_init)
, atom_reads(R_T0, R_ScreenX, R_ScreenY, R_ScreenBuf)
, atom_writes(R_T0, R_ScreenX, R_ScreenY)
) {
/* display[0] = (0, 0, 320, 240); rest of struct zeroed. */
add_ui(R_ScreenX, R_0, ScreenRes_X), add_ui(R_ScreenY, R_0, ScreenRes_Y),
mac_store_v2s2(R_ScreenX, R_ScreenY, R_ScreenBuf, O_(DisplayEnv,display_area.width) + OA_(DoubleBuffer,display,0)),
store_word(R_0, R_ScreenBuf, O_(DisplayEnv,display_area) + OA_(DoubleBuffer,display,0)),
store_word(R_0, R_ScreenBuf, O_(DisplayEnv,screen) + OA_(DoubleBuffer,display,0)),
store_word(R_0, R_ScreenBuf, O_(DisplayEnv,vinterlace) + OA_(DoubleBuffer,display,0)),
/* display[1] = (0, 240, 320, 240); rest of struct zeroed. */
mac_store_rects2(R_0, R_ScreenY, R_ScreenX, R_ScreenY, R_ScreenBuf, O_(DisplayEnv,display_area) + OA_(DoubleBuffer,display,1)),
store_word(R_0, R_ScreenBuf, O_(DisplayEnv,screen) + OA_(DoubleBuffer,display,1)),
store_word(R_0, R_ScreenBuf, O_(DisplayEnv,vinterlace) + OA_(DoubleBuffer,display,1)),
mac_store_rects2(R_0, R_ScreenY, R_ScreenX, R_ScreenY, R_ScreenBuf, O_(DrawEnv,clip_area) + OA_(DoubleBuffer,draw,0)), /* draw[0].clip_area = (0, 240, 320, 240). C11's SetDefDrawEnv writes clip.y = y_arg. */
mac_store_v2s2( R_0, R_ScreenY, R_ScreenBuf, O_(DrawEnv,drawing_offset[0]) + OA_(DoubleBuffer,draw,0)), /* draw[0].drawing_offset[0] = (0, 240); C11 passes y_arg as ofs. */
mac_store_v2s2(R_ScreenX, R_ScreenY, R_ScreenBuf, O_(DrawEnv,clip_area.width) + OA_(DoubleBuffer,draw,1)),
/* draw[0].texture_window = (0, 0, 0, 0); two word-zeroes cover the full 8-byte tw field. */
store_word(R_0, R_ScreenBuf, O_(DrawEnv,texture_window.x) + OA_(DoubleBuffer,draw,0)),
store_word(R_0, R_ScreenBuf, O_(DrawEnv,texture_window.width) + OA_(DoubleBuffer,draw,0)),
store_word(R_0, R_ScreenBuf, O_(DrawEnv,drawing_offset[0].x) + OA_(DoubleBuffer,draw,1)),
store_word(R_0, R_ScreenBuf, O_(DrawEnv,texture_window.x) + OA_(DoubleBuffer,draw,1)),
store_word(R_0, R_ScreenBuf, O_(DrawEnv,texture_window.width) + OA_(DoubleBuffer,draw,1)),
/* draw[0].texture_page = 10 (gp0_tpage_default). C11 SetDefDrawEnv at C11_only.elf:0x8001273C writes the same 0x0A. . */
add_ui(R_T0, R_0, gp0_tpage_default),
store_half(R_T0, R_ScreenBuf, O_(DrawEnv,texture_page) + OA_(DoubleBuffer,draw,0)),
store_half(R_T0, R_ScreenBuf, O_(DrawEnv,texture_page) + OA_(DoubleBuffer,draw,1)),
/* draw[0] control bytes: flag_dither=1, flag_draw_on_display=1 (the dfe bit per psx-spx; libpsyx sets it via `SetDefDrawEnv`'s conditional at C11_only.elf:0x80012728), enable_auto_clear=1. Each byte is named;
* the previous `store_word(R_0, ..., +20)` overwrote all four with zero. */
add_ui(R_T0, R_0, 1),
store_byte(R_T0, R_ScreenBuf, O_(DrawEnv,flag_dither) + OA_(DoubleBuffer,draw,0)),
store_byte(R_T0, R_ScreenBuf, O_(DrawEnv,flag_draw_on_display) + OA_(DoubleBuffer,draw,0)),
store_byte(R_T0, R_ScreenBuf, O_(DrawEnv,enable_auto_clear) + OA_(DoubleBuffer,draw,0)),
store_byte(R_T0, R_ScreenBuf, O_(DrawEnv,flag_dither) + OA_(DoubleBuffer,draw,1)),
store_byte(R_T0, R_ScreenBuf, O_(DrawEnv,flag_draw_on_display) + OA_(DoubleBuffer,draw,1)),
store_byte(R_T0, R_ScreenBuf, O_(DrawEnv,enable_auto_clear) + OA_(DoubleBuffer,draw,1)),
/* draw[0].initial_bg_color = (r=7, g=7, b=7). */
add_ui(R_T0, R_0, 7),
mac_store_rgb8(R_T0,R_T0,R_T0, R_ScreenBuf, O_(DrawEnv,initial_bg_color) + OA_(DoubleBuffer,draw,0)),
mac_store_rgb8(R_T0,R_T0,R_T0, R_ScreenBuf, O_(DrawEnv,initial_bg_color) + OA_(DoubleBuffer,draw,1)),
mac_yield(),
};
enum {
R_IO_BaseAddr = R_T4 atom_reg, /* Caller-pinned: IO_BASE_ADDR = 0x1F800000 */
#define R_IO_BaseAddr_Code R_T4_Code
};
internal MipsAtom_(gp_screen_init) atom_info(atom_phase(screen_init), atom_reads(R_IO_BaseAddr)) {
store_word(R_0, R_IO_BaseAddr, GPIO_PORT1_OFFSET), /* GP1(00h) Reset */
mac_gcmd_push(gp1_word_ResetCmdBuffer(), R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET), /* GP1(01h) ClearFIFO */
mac_gcmd_push(gp1_word_AcknowledgeIRQ(), R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET), /* GP1(02h) AckIRQ */
mac_gcmd_push(gp1_word_DisplayOn(), R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET), /* GP1(03h) Display ON */
mac_gcmd_push(gp1_word_dma_to_gpu(), R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET), /* GP1(04h) DMADirection=2 (CPU→GPU). libpsyx's per-frame PutDrawEnv/DrawOTag use DMA2; without this the DMA queue never drains. */
mac_gcmd_push(gp1_word_StartDisplayArea(), R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET), /* GP1(05h) StartDisplayArea (X=0, Y=0) */
/* GP1: DisplayMode + Display Ranges */
mac_gcmd_push(gp1_word_display_mode_320x240_15bit_ntsc, R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET),
mac_gcmd_push(gp1_word_horizontal_range_ntsc, R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET),
mac_gcmd_push(gp1_word_vertical_range_ntsc, R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET),
/* GTE: SetGeomOffset (OFX, OFY) — ScreenRes_CenterX, ScreenRes_CenterY. */
load_upper_i(R_T5, ScreenRes_CenterX), gte_mv_to_ctrl_r(R_T5, gte_cr_OFX_Code),
load_upper_i(R_T5, ScreenRes_CenterY), gte_mv_to_ctrl_r(R_T5, gte_cr_OFY_Code),
/* GTE: SetGeomScreen (H) — CR26 (per PSX-SPX / libpsyx), value is the raw projection-plane distance, NOT shifted. */
add_ui(R_T5, R_0, ScreenZ), gte_mv_to_ctrl_r(R_T5, gte_cr_H_Code),
/* GP1: DisplayEnable — bit 0 = 0 (Display ON). */
mac_gcmd_push(gp1_word_DisplayOn(), R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET),
mac_yield(),
};
/* ----- pad_apply_input -----
* Reads pad[0].buttons + pad[0].left_x;
* Applies the input-semantics deltas to cube_rot.y + floor_rot.y:
* - D-pad Left: cube_rot.y += 30, floor_rot.y += 5
* - D-pad Right: cube_rot.y -= 30, floor_rot.y -= 5
* - Analog stick X (dead zone 0x70..0x90):
* cube delta = (0x80 - left_x) >> 2 (range approx -32..+32)
* floor delta = (0x80 - left_x) >> 5 (range approx -4..+4)
* - D-pad + analog deltas add when used together.
*
* Convention:
* pad_state = 0 means no buttons active.
* The fail-safe zero-button value flows through unchanged, so a disconnected/fresh pad produces no rotation.
* The branch_le_zero pattern below matches the existing pad_input_demo convention (atom body lines 248/257).
*
* Signed-delta trick:
* load_byte_u zero-extends left_x to 32 bits; sub_u from 0x80 wraps to a SIGNED two's-complement value in the negative range;
* shift_aright (sra) then correctly sign-extends the shift for both positive (left_x < 0x80) and negative (left_x > 0x80) cases.
* Digital pads publish left_x = 0x80 → delta = 0 → no rotation, so the analog step is naturally a no-op for digital controllers.
*/
typedef Struct_(Binds_PadApplyInput) {
PadState* state;
V3_S2* cube_rot;
V3_S2* floor_rot;
};
enum {
R_PadStateT5 = R_T5 atom_reg,
R_CubeRot = R_T1 atom_reg,
R_FloorRot = R_T2 atom_reg,
};
internal MipsAtom_(pad_apply_input) atom_info(atom_bind(Binds_PadApplyInput)
, atom_reads(R_T0, R_CubeRot, R_FloorRot, R_T3, R_T4, R_PadStateT5, R_TapePtr)
, atom_writes( R_CubeRot, R_FloorRot)
) {
/* Pop Binds from tape (state, cube_rot, floor_rot) */
load_word(R_PadStateT5, R_TapePtr, O_(Binds_PadApplyInput,state)),
load_word(R_CubeRot, R_TapePtr, O_(Binds_PadApplyInput,cube_rot)),
load_word(R_FloorRot, R_TapePtr, O_(Binds_PadApplyInput,floor_rot)),
add_ui_self( R_TapePtr, S_(Binds_PadApplyInput)),
/* Load pad[0].buttons into R_T0. */
load_word(R_T0, R_PadStateT5, O_(PadState,buttons)), nop,
// Note(Ed): Potential op with delay slot?
/* D-pad Left: cube_rot.y += 30, floor_rot.y += 5. */
and_i(R_T3, R_T0, pad0_(Pad_Left)), branch_le_zero(R_T3, atom_offset(dpad_left, exit_dpad_left)),
load_half( R_T4, R_CubeRot, O_(V3_S2,y)), /* BD-slot */
load_half( R_T3, R_FloorRot, O_(V3_S2,y)),
add_si( R_T4, R_T4, 30),
add_si( R_T3, R_T3, 5),
store_half(R_T4, R_CubeRot, O_(V3_S2,y)),
store_half(R_T3, R_FloorRot, O_(V3_S2,y)),
atom_label(exit_dpad_left)
/* D-pad Right: cube_rot.y -= 30, floor_rot.y -= 5. */
and_i(R_T3, R_T0, pad0_(Pad_Right)), branch_le_zero(R_T3, atom_offset(dpad_right, exit_dpad_right)),
load_half( R_T4, R_CubeRot, O_(V3_S2,y)), /* BD-slot */
load_half( R_T3, R_FloorRot, O_(V3_S2,y)),
add_si( R_T4, R_T4, -30),
add_si( R_T3, R_T3, -5),
store_half(R_T4, R_CubeRot, O_(V3_S2,y)),
store_half(R_T3, R_FloorRot, O_(V3_S2,y)),
atom_label(exit_dpad_right)
/* Analog left-stick X: dead zone 0x70..0x90.
* Cube delta = (0x80 - left_x) >> 2; floor delta = (0x80 - left_x) >> 5. */
load_byte_u(R_T3, R_PadStateT5, O_(PadState,left_x)),
/* Dead-zone check: skip analog if left_x in [0x70, 0x90] inclusive. Outside dead zone on LOW side: left_x < 0x70 (strictly).
* set_lt_u(R_T4, R_T3, R_T4=0x70) → R_T4 = (left_x < 0x70) ? 1 : 0. */
add_ui(R_T4, R_0, 0x70), set_lt_u(R_T4, R_T3, R_T4), branch_ne(R_T4, R_0, atom_offset(dead_zone_low_check, dead_low_active)),
add_ui(R_T4, R_0, 0x80), /* BD-slot: pre-load 0x80 for dead_low_active */
atom_label(dead_check_upper)
/* left_x >= 0x70 → check upper bound. */
load_byte_u(R_T3, R_PadStateT5, O_(PadState,left_x)), /* reload */
add_ui( R_T4, R_0, 0x90),
/* R_T4 = (0x90 < left_x) ? 1 : 0 → (left_x > 0x90) ? 1 : 0 */
set_lt_u(R_T4, R_T4, R_T3), branch_ne(R_T4, R_0, atom_offset(dead_zone_high_check, dead_high_active)),
add_ui( R_T4, R_0, 0x80), /* BD-slot: pre-load 0x80 for dead_high_active */
jump_rel(atom_offset(dead_zone_skip, exit_stick)),
mac_yield_load(),
atom_label(dead_low_active)
/* R_T3 = left_x (from line 632 lbu; not clobbered between dead_zone_low_check branch + its BD-slot `add_ui R_T4, 0x80`).
* The earlier `load_byte_u(R_T3, ...)` reload was redundant and introduced a load-use hazard on the next `sub_u`.
* R_T4 = 0x80 from the BD-slot of `dead_zone_low_check`'s branch_ne. */
sub_u( R_T3, R_T4, R_T3), /* R_T3 = 0x80 - left_x */
/* delta = 0x80 - left_x (positive). */
/* R_T4 = cube_delta */
shift_aright(R_T4, R_T3, 2),
load_half( R_T0, R_CubeRot, O_(V3_S2,y)),
nop,
add_u( R_T0, R_T0, R_T4),
store_half( R_T0, R_CubeRot, O_(V3_S2,y)),
/* R_T4 = floor_delta — moved into the load-delay slot of the floor load below (fills the 1-instruction gap;
* doesn't read R_T0; R_T4 settles by the subsequent add_u). */
load_half( R_T0, R_FloorRot, O_(V3_S2,y)),
shift_aright(R_T4, R_T3, 5),
add_u( R_T0, R_T0, R_T4),
store_half( R_T0, R_FloorRot, O_(V3_S2,y)),
jump_rel(atom_offset(end_low, exit_stick)),
mac_yield_load(),
atom_label(dead_high_active)
/* R_T3 = left_x (from line 641 lbu in dead_check_upper; not clobbered between dead_zone_high_check branch + its BD-slot `add_ui R_T4, 0x80`).
* The earlier `load_byte_u(R_T3, ...)` reload was redundant and introduced a load-use hazard on the next `sub_u`.
* R_T4 = 0x80 from the BD-slot of `dead_zone_high_check`'s branch_ne. */
sub_u( R_T3, R_T4, R_T3),
/* delta = 0x80 - left_x (signed negative). */
shift_aright(R_T4, R_T3, 2), /* R_T4 = cube_delta (signed) */
load_half( R_T0, R_CubeRot, O_(V3_S2,y)),
nop,
add_u( R_T0, R_T0, R_T4),
store_half( R_T0, R_CubeRot, O_(V3_S2,y)),
/* R_T4 = floor_delta (signed) — moved into the load-delay slot of the floor load below. */
load_half( R_T0, R_FloorRot, O_(V3_S2,y)),
shift_aright(R_T4, R_T3, 5),
add_u( R_T0, R_T0, R_T4),
store_half( R_T0, R_FloorRot, O_(V3_S2,y)),
atom_label(no_jump_fallthrough)
mac_yield_load(),
atom_label(exit_stick)
/* NOT mac_yield() — R_AtomJmp was already loaded in the BD-slot of the dead-zone/exit branch. */
mac_yield_tail(),
};
enum {
R_PrimCursor = R_T7 atom_reg atom_type(U4*), /* VRAM output cursor (primitive buffer) */
R_FaceCursor = R_T4 atom_reg atom_type(V4_S2*), /* Cube face-index cursor (V4_S2*); floor context switches to V3_S2* via atom_phase */
R_VertBase = R_T5 atom_reg atom_type(V3_S2*), /* Base address of the vertex array */
R_OtBase = R_T6 atom_reg atom_type(U4*), /* Base address of the Ordering Table */
#define R_PrimCursor_Code R_T7_Code
#define R_FaceCursor_Code R_T4_Code
#define R_VertBase_Code R_T5_Code
#define R_OtBase_Code R_T6_Code
};
typedef Struct_(Binds_CubeTri) {
U4 PrimCursor;
V4_S2* FaceCursor;
V3_S2* VertBase;
U4* OtBase;
};
internal MipsAtom_(rbind_cube_g4_face) atom_info(atom_bind(Binds_CubeTri), atom_phase(cube_g4)
, atom_reads(R_TapePtr)
, atom_writes(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase, R_TapePtr)
){
/* Pop 4 arguments from the tape directly into the workspace registers */
load_word(R_PrimCursor, R_TapePtr, O_(Binds_CubeTri,PrimCursor)),
load_word(R_FaceCursor, R_TapePtr, O_(Binds_CubeTri,FaceCursor)),
load_word(R_VertBase, R_TapePtr, O_(Binds_CubeTri,VertBase)),
load_word(R_OtBase, R_TapePtr, O_(Binds_CubeTri,OtBase)),
add_ui_self( R_TapePtr, S_(Binds_CubeTri)),
mac_yield()
};
// cube_g4_face — Draw one cube face (Gouraud-shaded quad) via the GTE tape pipeline
internal
MipsAtom_(cube_g4_face) atom_info(atom_phase(cube_g4),
atom_reads( R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase),
atom_writes(R_PrimCursor, R_FaceCursor)
){
load_half_u(R_T0, R_FaceCursor, 0 * S_(S2)),
load_half_u(R_T1, R_FaceCursor, 1 * S_(S2)),
load_half_u(R_T2, R_FaceCursor, 2 * S_(S2)),
load_half_u(R_T3, R_FaceCursor, 3 * S_(S2)),
mac_gte_load_tri_verts(R_VertBase, R_T0, R_T1, R_T2),
nop2, gte_cmdw_rotate_translate_perspective_triple, // required cpu -> gte delay slot
gte_cmdw_nclip,
gte_mv_from_data_r(R_T0, C2_MAC0), nop,
branch_le_zero(R_T0, atom_offset(cull, cube_g4_face_exit)),
/* BD-slot: write the prim tag (R_0=0; overwrites the legacy tag word in the prim_buffer).
* If branch IS taken (face culled), the body is skipped and this 0-tag is stranded —
* harmless because the OT entry that points to this prim is created later, only on the body path. */
store_word(R_0, R_PrimCursor, O_(Poly_G4, tag)),
shift_lleft(R_AT, R_T3, v3s2_byteoff), add_u(R_AT, R_AT, R_VertBase),
load_word(R_V0, R_AT, O_(V3_S2, x)), load_word(R_V1, R_AT, O_(V3_S2, z)),
gte_mv_to_data_r(R_V0, C2_VXY0), gte_mv_to_data_r(R_V1, C2_VZ0),
mac_gte_store_g4_p012(R_PrimCursor),
gte_cmdw_rotate_translate_perspective_single,
mac_gte_store_g4_p3(R_PrimCursor),
gte_cmdw_avg_sort_z4,
gte_mv_from_data_r(R_T1, C2_OTZ),
add_ui( R_AT, R_0, OrderingTbl_Len),
set_lt_u( R_AT, R_T1, R_AT),
branch_equal(R_AT, R_0, atom_offset(bounds_chk, cube_g4_face_exit)), nop,
mac_insert_ot_tag_g4(R_OtBase, R_PrimCursor),
mac_format_g4_color(R_PrimCursor,
/* c0 magenta */ 0xFF, 0x00, 0xFF,
/* c1 yellow */ 0xFF, 0xFF, 0x00,
/* c2 cyan */ 0x00, 0xFF, 0xFF,
/* c3 green */ 0x00, 0xFF, 0x00),
// end: branch(bounds_chk)
// end: branch(cull)
atom_label(cube_g4_face_exit)
add_ui_self(R_PrimCursor, S_(Poly_G4)), /* 9 words = Poly_G4 */
add_ui_self(R_FaceCursor, S_(S2) * 4), /* 4 × S2 = 8 bytes */
mac_yield()
};
typedef Struct_(Binds_FloorTri) {
U4 PrimCursor;
V3_S2* FaceCursor;
V3_S2* VertBase;
U4* OtBase;
};
internal
MipsAtom_(rbind_floor_f3_face) atom_info(atom_bind(Binds_FloorTri), atom_phase(floor_f3)
, atom_reads(R_TapePtr)
, atom_writes(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase, R_TapePtr)
){
/* Pop 4 arguments from the tape directly into the workspace registers */
load_word(R_PrimCursor, R_TapePtr, O_(Binds_FloorTri,PrimCursor)),
load_word(R_FaceCursor, R_TapePtr, O_(Binds_FloorTri,FaceCursor)),
load_word(R_VertBase, R_TapePtr, O_(Binds_FloorTri,VertBase)),
load_word(R_OtBase, R_TapePtr, O_(Binds_FloorTri,OtBase)),
add_ui_self( R_TapePtr, S_(Binds_FloorTri)),
mac_yield()
};
// atom_dbg_skip
internal
MipsAtom_(floor_f3_face) atom_info(atom_phase(floor_f3)
, atom_reads( R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase)
, atom_writes(R_PrimCursor, R_FaceCursor)
) {
mac_load_tri_indices( R_FaceCursor, R_T0, R_T1, R_T2),
mac_gte_load_tri_verts(R_VertBase, R_T0, R_T1, R_T2),
nop2, gte_cmdw_rotate_translate_perspective_triple, // 2 nops retire the final cpu -> gte writes before RTPT
gte_cmdw_nclip,
/* Culling (Branch forward if Backface) */
gte_mv_from_data_r(R_T0, C2_MAC0),
nop, branch_le_zero(R_T0, atom_offset(culling, floor_f3_face_exit)), nop, // required gte -> cpu load-delay slot.
/* Format Primitive */
mac_gte_store_f3(R_PrimCursor),
/* Calculate Depth */
gte_avg_sort_z3,
gte_mv_from_data_r(R_T1, C2_OTZ),
/* Bounds Check OTZ < 2048 (Branch forward to skip insertion) */
add_ui( R_AT, R_0, OrderingTbl_Len),
set_lt_u( R_AT, R_T1, R_AT),
branch_equal(R_AT, R_0, atom_offset(bounds_chk, floor_f3_face_exit)), nop,
mac_format_f3_color(R_PrimCursor, 0xFF, 0xFF, 0xFF), // RGB-form (R=FF, G=FF, B=FF = white)
mac_insert_ot_tag_f3(R_OtBase, R_PrimCursor), /* Insert into Ordering Table Linked List */
add_ui_self(R_PrimCursor, S_(Poly_F3)), /* Advance Prim Cursor (5 words) */
// Note(Ed): No bounds checking, should be checked before atom runs.
// end: branch(bounds_chk)
// end: branch(culling)
/* Advance Input Cursor & Yield (Both branch targets land here) */
atom_label(floor_f3_face_exit)
add_ui_self(R_FaceCursor, S_(S2) * 4), /* Advance Face Cursor (4 * S2 = 8 bytes) */
mac_yield()
};
typedef Struct_(Binds_SyncPrimitiveArena) { U4 used; U4 cursor; };
internal MipsAtom_(sync_primitive_arena) atom_info(atom_bind(Binds_SyncPrimitiveArena)
, atom_reads( R_TapePtr, R_PrimCursor)
, atom_writes(R_TapePtr)
){
load_word(R_AT, R_TapePtr, O_(Binds_SyncPrimitiveArena,used)),
load_word(R_T0, R_TapePtr, O_(Binds_SyncPrimitiveArena,cursor)),
add_ui_self( R_TapePtr, S_(Binds_SyncPrimitiveArena)),
/* Calculate byte offset and store directly back to RAM */
sub_u( R_T0, R_PrimCursor, R_T0), // R_T0 = R_PrimCursor - binds.cursor
store_word(R_T0, R_AT, 0), // R_AT[0] = R_T0
mac_yield()
};
#pragma endregion Baked Atoms
+342
View File
@@ -0,0 +1,342 @@
#pragma region Vendors
#include <stdio.h>
#include <stdlib.h>
#include <assert.h>
// #include "libgpu.h"
// #include "libetc.h"
// #include "libgte.h"
#pragma endregion Vendors
#pragma region Duffle Headers
# include "duffle/gen/macs.h"
# include "duffle/gen/offsets.h"
#include "duffle/word_count.metadata.h"
#include "duffle/dsl.h"
#include "duffle/memory.h"
#include "duffle/math.h"
#include "duffle/gcc_asm.h"
#include "duffle/mips.h"
#include "duffle/gp.h"
#include "duffle/gte.h"
#include "duffle/pad.h"
#include "duffle/dsl.atom.h"
#include "duffle/lottes_tape.h"
#include "duffle/psyq.h"
#pragma endregion Duffle Headers
#pragma region Duffle TUs
#include "duffle/math.atom.c"
#include "duffle/mips.atom.c"
#include "duffle/gte.atom.c"
#include "duffle/gp.atom.c"
#include "duffle/pad.atom.c"
#include "duffle/psyq.atom.c"
#pragma endregion Duffle TUs
#pragma region Hello Camera Headers
# include "gen/macs.h"
# include "gen/offsets.h"
#include "hello_camera.h"
#pragma endregion Hello Camera Headers
#pragma region Hello Joypad TUs
#include "hello_camera.atom.c"
#pragma endregion Hello Joypad TUs
enum {
Scratchpad_Len = 1024,
MemTape_Len = 512,
};
typedef Struct_(SMemory) {
PrimitiveArena primitives;
A2_OrderingTable_Buffer ordering_tbl;
DoubleBuffer screen_buf;
S4 active_buf_id;
U4 MemTape[MemTape_Len];
M3_S2 tform_world;
Ent_Cube cube;
Ent_Floor floor;
PadBiosRaw pad_raw[2];
PadState pad[2];
U4_V scratchpad; // d-cache
};
global SMemory smem;
extern SMemory smem;
I_ B1* prim__alloc(U4 type_width, Str8 type_name) {
gknown PrimitiveArena* pa = & smem.primitives;
gknown B1* buf = (B1*) r_(smem.primitives.buf)[smem.active_buf_id];
assert(pa->used + type_width < PrimitiveBuff_Len);
B1* next = buf + pa->used;
pa->used += type_width;
return next;
}
#define prim_alloc(type) (type*)prim__alloc(S_(type), slit( stringify(type)))
/* Uses ONE 8-byte frame allocated via the compiler's standard prologue.
* The 4 wasted-arg words for B(12h) InitPAD2 live at [SP+0..15] but are not explicitly allocated.
* The compiler handles the MIPS O32 "wasted stack" convention for us by treating the B-call as a 4-arg call.
*
* The buffer pointers are passed as arguments so the compiler keeps them in callee-saved registers;
* The B(12h) asm volatile block does NOT clobber those registers (it clobbers only the volatile GPRs + the B-table arg registers explicitly).
* The C-level writes after the call re-load the pointers from their callee-saved homes.
*
* The clobber list for both B-calls names the full BIOS destroy set documented in kernelbios.md:167-174 (R1..R15, R24..R25, R31, HI/LO).
* The kernel-ABI "volatile GPRs" subset is clb_system; the rest of the destroy set is enumerated explicitly here. */
NI_ void pad_bios_init_start(PadBiosRaw* raw0, PadBiosRaw* raw1)
{
/* Pin raw0 + raw1 to $a0 + $a1 via rgcc; the B(12h) call uses these directly.
* The `(void)` casts mark them as unread after the call so the compiler doesn't need to move them back. */
register PadBiosRaw* p0 rgcc(R_A0) = raw0;
register PadBiosRaw* p1 rgcc(R_A1) = raw1;
(void)p0; (void)p1;
// TODO(Ed): Properly annotate the raw values in the inline asm instructions.
// Use enums.
/* B(12h) InitPAD2(raw0, 0x22, raw1, 0x22)
* $a0 = raw0 (rgcc-bound; survives the sequence below)
* $a1 = raw1 (preserved into $a2 before $a1 is overwritten)
* $a2 = raw1 (moved from $a1; survives $a1's overwrite)
* $a3 = 0x22 (immediate)
* $t1 = 0x12 (function number)
* $t2 = 0xB0 (BIOS B-table address) */
asm volatile(
asm_words(
or_u( rarg_2, rarg_1, rdiscard), /* $a2 = $a1 = raw1 */
add_ui( rarg_1, rdiscard, 0x22), /* $a1 = 0x22 */
add_ui( rarg_3, rdiscard, 0x22), /* $a3 = 0x22 */
add_ui( rtmp_1, rdiscard, 0x12), /* $t1 = 0x12 */
add_ui( rtmp_2, rdiscard, 0xB0), /* $t2 = 0xB0 */
call_reg(rtmp_2), /* jalr $t2, $ra */
nop /* BD slot */
)
asm_rpins, r_use(p0), r_use(p1)
asm_clobber:
rlit(R_AT),
rlit(R_V0), rlit(R_V1),
rlit(R_T0), rlit(R_T1), rlit(R_T2), rlit(R_T3), rlit(R_T4),
rlit(R_T5), rlit(R_T6), rlit(R_T7), rlit(R_T8), rlit(R_T9),
rlit(R_RA),
clb_mem_drain
);
/* The C-level writes re-load the pointers via the parameter names and write 0xFF to each
* buffer's status byte to mark the initial-state hazard documented in kernelbios.md:1621-1624. */
u1_v(raw0)[0] = 0xFF;
u1_v(raw1)[0] = 0xFF;
/* B(13h) StartPAD2() — no args. The BIOS preserves $sp. */
asm volatile(
asm_words(
add_ui( rtmp_1, rdiscard, 0x13), /* $t1 = 0x13 */
add_ui( rtmp_2, rdiscard, 0xB0), /* $t2 = 0xB0 (re-load) */
call_reg(rtmp_2), /* jalr $t2, $ra */
nop /* BD slot */
)
asm_clobber:
rlit(R_AT),
rlit(R_V0), rlit(R_V1),
rlit(R_T0), rlit(R_T1), rlit(R_T2), rlit(R_T3), rlit(R_T4),
rlit(R_T5), rlit(R_T6), rlit(R_T7), rlit(R_T8), rlit(R_T9),
rlit(R_RA),
clb_mem_drain
);
}
GCC_OPTIMIZATION_DISABLE
void update(PrimitiveArena* pa, U4* ordering_buf)
{
TapeBuilder tb = tb_make(slice_ut_arr(smem.MemTape));
if (1) // Pad Input
{
tb.used = 0; tb_scope_run(& tb) {
/* BIOS-owned polling: per-frame snapshot of both ports. */
tb_emit_(pad_bios_snapshot);
tb_data_(raw, & smem.pad_raw[0]);
tb_data_(state, & smem.pad[0]);
tb_emit_(pad_bios_snapshot);
tb_data_(raw, & smem.pad_raw[1]);
tb_data_(state, & smem.pad[1]);
/* Per-frame rotation apply: consume pad[0].buttons + pad[0].left_x */
tb_emit_(pad_apply_input);
tb_data_(state, & smem.pad[0]);
tb_data_(cube_rot, & smem.cube.rot);
tb_data_(floor_rot, & smem.floor.rot);
}
}
orderingtbl_clear_reverse(ordering_buf, OrderingTbl_Len);
// Update the position based on acceleration and velocity
gknown V3_S4_R pos = & smem.cube.pos;
gknown V3_S4_R vel = & smem.cube.vel;
gknown V3_S4_R acc = & smem.cube.accel;
add_v3s4(vel, acc[0]);
add_v3s4_fp(pos, vel[0]);
// vel->x += acc->x;
// vel->y += acc->y;
// vel->z += acc->z;
// pos->x += vel->x;
// pos->y += vel->y;
// pos->z += vel->z;
if (pos->y + 150 > smem.floor.pos.y) vel->y *= -1;
// Prep
S4 nclip = 0;
S4 orderingtbl_z = 0;
A2_S2 p; //???
S4 flag; //????
// Draw cube
if (1)
{
m3s2_rotation (& smem.cube.rot, & smem.tform_world);
m3s2_translation(& smem.tform_world, & smem.cube.pos);
m3s2_scale (& smem.tform_world, & smem.cube.scale);
gte_matrix_set_rotation (& smem.tform_world);
gte_matrix_set_translation(& smem.tform_world);
U4 prim_base = u4_(pa->buf[smem.active_buf_id]);
U4 prim_cursor = prim_base + pa->used;
tb.used = 0; tb_scope(& tb) {
tb_emit(& tb, rbind_cube_g4_face);
tb_data(& tb, prim_cursor);
tb_data(& tb, u4_(smem.cube.faces));
tb_data(& tb, u4_(smem.cube.verts));
tb_data(& tb, u4_(ordering_buf));
for (U4 i = 0; i < Cube_num_faces; i++) {
// Two triangles per quad face: (x,y,z) and (x,z,w)
tb_emit(& tb, cube_g4_face);
}
tb_emit(& tb, sync_primitive_arena);
tb_data(& tb, u4_(& pa->used));
tb_data(& tb, prim_base);
}
tape_run(tb_slice(tb));
// smem.cube.rot.y += 30;
}
// Draw floor
if (1)
{
m3s2_rotation (& smem.floor.rot, & smem.tform_world);
m3s2_translation(& smem.tform_world, & smem.floor.pos);
m3s2_scale (& smem.tform_world, & smem.floor.scale);
U4 prim_base = u4_(pa->buf[smem.active_buf_id]);
U4 prim_cursor = prim_base + pa->used;
// TODO(Ed): We should do a bounds check beforehand to confirm pa can hold all tris?
// The tape atoms in-flight should not need to care.
// Prepare the tape. (Push protocol to tape)
tb.used = 0; tb_scope(& tb) {
tb_emit(& tb, set_gte_world);
tb_data(& tb, u4_(& smem.tform_world));
tb_emit(& tb, rbind_floor_f3_face);
// TODO(Ed): Just use a single context struct ref
tb_data(& tb, prim_cursor);
tb_data(& tb, u4_(smem.floor.faces));
tb_data(& tb, u4_(smem.floor.verts));
tb_data(& tb, u4_(ordering_buf));
for (U4 i = 0; i < Floor_num_faces; i++) {
tb_emit(& tb, floor_f3_face);
}
// After floor_f3_face iterations complete, the primitive arena's used counter needs updating.
tb_emit(& tb, sync_primitive_arena);
tb_data(& tb, u4_(& pa->used));
tb_data(& tb, prim_base);
}
tape_run(tb_slice(tb));// Fire off the tape.
// C-side state (pa->used) has already been updated by the tape!
// smem.floor.rot.y += 5;
}
}
GCC_OPTIMIZATION_ENABLE
void render(void) {
}
void gp_display_frame(DoubleBuffer* screen_buf, S4* active_buf_id, U4* ordering_buf, PrimitiveArena* pa) {
draw_sync(0);
vsync(0);
displayenv_put(& r_(screen_buf->display)[active_buf_id[0] ]);
drawenv_put (& r_(screen_buf->draw) [active_buf_id[0] ]);
{
draw_orderingtbl(ordering_buf + OrderingTbl_Len - 1);
pa->used = 0;
}
active_buf_id[0] = ! active_buf_id[0]; // Swap current buffer
}
GCC_OPTIMIZATION_DISABLE
void hot_reload_entry(void)
{
smem.primitives.used = 0;
while (1) {
gknown S4* active_buf_id = & smem.active_buf_id;
gknown U4* ordering_buf = r_(smem.ordering_tbl)[active_buf_id[0]];
gknown PrimitiveArena* pa = & smem.primitives;
update(pa, ordering_buf);
render();
gp_display_frame(& smem.screen_buf, active_buf_id, ordering_buf, pa);
}
}
int main(void)
{
smem = (SMemory){0};
smem.scratchpad = C_(U4_V, 0x1F800000);
// smem.primitives.used = 0;
// smem.active_buf_id = 0;
/*Persistent Entity Setup*/{
ent_cube128_init(& smem.cube.verts, & smem.cube.faces); {
Ent_Cube* cube = & smem.cube;
cube->rot = v3s2(0, 0, 0);
cube->scale = v3s4_fp_one();
cube->accel = v3s4(0, 1, 0);
cube->pos = v3s4(0, -400, 1800);
}
ent_floor_init(& smem.floor.verts, & smem.floor.faces); {
Ent_Floor* floor = & smem.floor;
floor->rot = v3s2(0, 0, 0);
floor->pos = v3s4(0, 450, 1800);
floor->scale = v3s4_fp_one();
}
}
TapeBuilder tb = tb_make(slice_ut_arr(smem.MemTape)); {
reset_graph(0);
/* Direct BIOS: poll both ports during VBlank. */
pad_bios_init_start(& smem.pad_raw[0], & smem.pad_raw[1]);
/* Pinned registers for the GPU init atom. */
register U4* io_base_addr rgcc(R_IO_BaseAddr) = u4_r(IO_BASE_ADDR);
register DoubleBuffer* screen_buf rgcc(R_ScreenBuf) = & smem.screen_buf;
tb.used = 0; tb_scope_run(& tb) {
tb_emit(& tb, screen_env_init);
tb_emit(& tb, gp_screen_init);
}
}
hot_reload_entry();
return 0;
}
GCC_OPTIMIZATION_ENABLE
+102
View File
@@ -0,0 +1,102 @@
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "duffle/dsl.h"
# include "duffle/math.h"
# include "duffle/gp.h"
# include "duffle/pad.h"
#endif
enum {
// PrimitiveBuff_Len = 4096,
// OrderingTbl_Len = 2048,
PrimitiveBuff_Len = 131072,
OrderingTbl_Len = 8192,
};
enum {
ScreenRes_X = 320,
ScreenRes_Y = 240,
ScreenZ = 320,
ScreenRes_CenterX = (ScreenRes_X >> 1),
ScreenRes_CenterY = (ScreenRes_Y >> 1),
};
enum {
fp_one = (1 << 12),
};
#define v3s4_fp_one() v3s4(fp_one, fp_one, fp_one)
typedef U4 OrderingTable_Buffer[OrderingTbl_Len];
typedef Array_(OrderingTable_Buffer, 2);
typedef B1 PrimitiveBuffer[PrimitiveBuff_Len];
typedef Array_(PrimitiveBuffer, 2);
typedef Struct_(PrimitiveArena) {
A2_PrimitiveBuffer buf;
U4 used;
};
#define Cube_num_verts 8
typedef Array_(V3_S2, Cube_num_verts);
#define Cube_num_faces 6
typedef Array_(V4_S2, Cube_num_faces);
I_ void ent_cube128_init(A8_V3_S2* verts, A6_V4_S2* faces) {
LP_ A8_V3_S2 baked_verts = (A8_V3_S2) {
{ -128, -128, -128 },
{ 128, -128, -128 },
{ 128, -128, 128 },
{ -128, -128, 128 },
{ -128, 128, -128 },
{ 128, 128, -128 },
{ 128, 128, 128 },
{ -128, 128, 128 }
};
LP_ A6_V4_S2 baked_faces = (A6_V4_S2) {
{ 3, 2, 0, 1 },
{ 0, 1, 4, 5 },
{ 4, 5, 7, 6 },
{ 1, 2, 5, 6 },
{ 2, 3, 6, 7 },
{ 3, 0, 7, 4 },
};
mem_copy(u4_(verts), u4_(& baked_verts), S_(A8_V3_S2) );
mem_copy(u4_(faces), u4_(& baked_faces), S_(A6_V4_S2) );
return;
}
typedef Struct_(Ent_Cube) {
V3_S4 accel;
V3_S4 vel;
V3_S4 pos;
V3_S4 scale;
V3_S2 rot;
A8_V3_S2 verts;
A6_V4_S2 faces;
};
#define Floor_num_verts 4
typedef Array_(V3_S2, Floor_num_verts);
#define Floor_num_faces 2
typedef Array_(V3_S2, Floor_num_faces);
I_ void ent_floor_init(A4_V3_S2* verts, A2_V3_S2* faces) {
LP_ A4_V3_S2 baked_verts = (A4_V3_S2) {
{ -900, 0, -900 },
{ -900, 0, 900 },
{ 900, 0, -900 },
{ 900, 0, 900 },
};
LP_ A2_V3_S2 baked_faces = (A2_V3_S2) {
{ 0, 1, 2 },
{ 1, 3, 2 },
};
mem_copy(u4_(verts), u4_(& baked_verts), S_(A4_V3_S2));
mem_copy(u4_(faces), u4_(& baked_faces), S_(A2_V3_S2));
};
typedef Struct_(Ent_Floor) {
V3_S4 accel;
V3_S4 pos;
V3_S4 scale;
V3_S2 rot;
A4_V3_S2 verts;
A2_V3_S2 faces;
};
+29
View File
@@ -0,0 +1,29 @@
// Auto-generated by ps1_meta.lua (passes/offsets.lua) — DO NOT EDIT
// Source: C:\projects\Pikuma\ps1\code\gte_hello\hello_gte_tape.c
#pragma once
#pragma region hello_gte_tape
// --- atom: cube_g4_face (77 words) ---
#define _atom_offset_cull_cube_g4_face_exit 42
#define _atom_offset_bounds_chk_cube_g4_face_exit 24
enum {
atom_offset_cull_cube_g4_face_exit = _atom_offset_cull_cube_g4_face_exit,
atom_offset_bounds_chk_cube_g4_face_exit = _atom_offset_bounds_chk_cube_g4_face_exit,
};
// --- atom: floor_f3_face (58 words) ---
#define _atom_offset_culling_floor_f3_face_exit 25
#define _atom_offset_bounds_chk_floor_f3_face_exit 16
enum {
atom_offset_culling_floor_f3_face_exit = _atom_offset_culling_floor_f3_face_exit,
atom_offset_bounds_chk_floor_f3_face_exit = _atom_offset_bounds_chk_floor_f3_face_exit,
};
#pragma endregion hello_gte_tape
+29
View File
@@ -0,0 +1,29 @@
// Auto-generated by ps1_meta.lua (passes/offsets.lua) — DO NOT EDIT
// Source: C:\projects\Pikuma\ps1\code\hello_gte\hello_gte.tape.c
#pragma once
#pragma region hello_gte.tape
// --- atom: cube_g4_face (77 words) ---
#define _atom_offset_cull_cube_g4_face_exit 42
#define _atom_offset_bounds_chk_cube_g4_face_exit 24
enum {
atom_offset_cull_cube_g4_face_exit = _atom_offset_cull_cube_g4_face_exit,
atom_offset_bounds_chk_cube_g4_face_exit = _atom_offset_bounds_chk_cube_g4_face_exit,
};
// --- atom: floor_f3_face (58 words) ---
#define _atom_offset_culling_floor_f3_face_exit 25
#define _atom_offset_bounds_chk_floor_f3_face_exit 16
enum {
atom_offset_culling_floor_f3_face_exit = _atom_offset_culling_floor_f3_face_exit,
atom_offset_bounds_chk_floor_f3_face_exit = _atom_offset_bounds_chk_floor_f3_face_exit,
};
#pragma endregion hello_gte.tape
@@ -1,6 +1,6 @@
#include "stdio.h"
#include <stdio.h>
#include <stdlib.h>
#include "assert.h"
#include <assert.h>
// #include "libgpu.h"
// #include "libetc.h"
// #include "libgte.h"
@@ -8,19 +8,22 @@
#include "duffle/dsl.h"
#include "duffle/memory.h"
#include "duffle/math.h"
#include "duffle/gcc_asm.h"
#include "duffle/mips.h"
#include "duffle/gp.h"
#include "duffle/gte.h"
# include "duffle/gen/lottes_tape.offsets.h"
# include "duffle/gen/macs.h"
# include "duffle/gen/offsets.h"
#include "duffle/atom_dsl.h"
#include "duffle/lottes_tape.h"
#include "duffle/word_count.metadata.h"
# include "tape_atom.metadata.h"
# include "gen/hello_gte_tape.offsets.h"
# include "gen/offsets.h"
#include "hello_gte.h"
#include "hello_gte_tape.c"
#include "hello_gte.tape.c"
typedef U4 OrderingTable_Buffer[OrderingTbl_Len];
typedef Array_(OrderingTable_Buffer, 2);
@@ -96,8 +99,13 @@ typedef Struct_(Ent_Floor) {
A2_V3_S2 faces;
};
enum { scratchpad_size = 1024, };
enum {
Scratchpad_Len = 1024,
MemTape_Len = 512,
};
typedef Struct_(SMemory) {
U4 MemTape[MemTape_Len];
DoubleBuffer screen_buf;
A2_OrderingTable_Buffer ordering_tbl;
PrimitiveArena primitives;
@@ -114,7 +122,7 @@ global SMemory smem;
extern SMemory smem;
// TODO(Ed):
FI_ U4* spad_warm(MipsAtom atom) {
FI_ U4* spad_warm(Slice_MipsCode atom) {
return nullptr;
}
@@ -154,6 +162,10 @@ void gp_screen_init_c11(DoubleBuffer* screen_buf, S4* active_buf_id)
// Initialize and setup the GTE geometry offsets
geom_init();
// NOTE: geom_set_offset/geom_set_screen are kept as-is (the libgte versions
// are known to be broken in this PSYQ 4.7 build — see report 2026-07-09).
// The user's research wants the C-side non-tape reference to work as a
// known-good baseline for comparison against the tape.
geom_set_offset(ScreenRes_CenterX, ScreenRes_CenterY);
geom_set_screen(ScreenZ);
@@ -175,6 +187,7 @@ void gp_display_frame(DoubleBuffer* screen_buf, S4* active_buf_id, U4* ordering_
void render(void) {
}
GCC_OPTIMIZATION_DISABLE
void update(PrimitiveArena* pa, U4* ordering_buf)
{
orderingtbl_clear_reverse(ordering_buf, OrderingTbl_Len);
@@ -200,13 +213,15 @@ void update(PrimitiveArena* pa, U4* ordering_buf)
A2_S2 p; //???
S4 flag; //????
TapeBuilder tb = tb_make(slice_ut_arr(smem.MemTape));
// Draw Cube
if (1)
if (0)
{
m3s2_rotation (& smem.cube.rot, & smem.tform_world);
m3s2_translation(& smem.tform_world, & smem.cube.pos);
m3s2_scale (& smem.tform_world, & smem.cube.scale);
gte_matrix_set_rotation (& smem.tform_world);
// gte_matrix_set_rotation (& smem.tform_world);
gte_matrix_set_translation(& smem.tform_world);
for (U4 face_id = 0; face_id < Cube_num_faces; face_id += 1)
{
@@ -241,7 +256,7 @@ void update(PrimitiveArena* pa, U4* ordering_buf)
smem.cube.rot.y += 30;
}
// Draw cube (tape method) - two triangles per face
if (0)
if (1)
{
m3s2_rotation (& smem.cube.rot, & smem.tform_world);
m3s2_translation(& smem.tform_world, & smem.cube.pos);
@@ -252,9 +267,8 @@ void update(PrimitiveArena* pa, U4* ordering_buf)
U4 prim_base = u4_(pa->buf[smem.active_buf_id]);
U4 prim_cursor = prim_base + pa->used;
LP_ U4 mem_temp_tape[512]; FArena tape_arena; farena_init(& tape_arena, slice_ut_arr(mem_temp_tape));
TapeBuilder tb = tb_make_old(&tape_arena); tb_scope(& tb) {
tb_emit(& tb, code_rbind_cube_tri);
tb.used = 0; tb_scope(& tb) {
tb_emit(& tb, rbind_cube_g4_face);
tb_data(& tb, prim_cursor);
tb_data(& tb, u4_(smem.cube.faces));
tb_data(& tb, u4_(smem.cube.verts));
@@ -262,10 +276,10 @@ void update(PrimitiveArena* pa, U4* ordering_buf)
for (U4 i = 0; i < Cube_num_faces; i++) {
// Two triangles per quad face: (x,y,z) and (x,z,w)
tb_emit(& tb, code_cube_tri);
tb_emit(& tb, cube_g4_face);
}
tb_emit(& tb, code_sync_primitive_arena);
tb_emit(& tb, sync_primitive_arena);
tb_data(& tb, u4_(& pa->used));
tb_data(& tb, prim_base);
}
@@ -291,7 +305,6 @@ void update(PrimitiveArena* pa, U4* ordering_buf)
register V3_S2* p1 rgcc(R_T5) = & smem.floor.verts[face->y];
register V3_S2* p2 rgcc(R_T6) = & smem.floor.verts[face->z];
// Three independent bases — full register discretion at the call site
gte_load_v0(p0, R_T4);
/*
asm volatile( ".word " "%0" ", %1" : :
@@ -334,38 +347,32 @@ void update(PrimitiveArena* pa, U4* ordering_buf)
m3s2_rotation (& smem.floor.rot, & smem.tform_world);
m3s2_translation(& smem.tform_world, & smem.floor.pos);
m3s2_scale (& smem.tform_world, & smem.floor.scale);
// TODO(Ed): This can either be in the tape or here...
gte_matrix_set_rotation (& smem.tform_world);
gte_matrix_set_translation(& smem.tform_world);
U4 prim_base = u4_(pa->buf[smem.active_buf_id]);
U4 prim_cursor = prim_base + pa->used;
// TODO(Ed): We should do a bounds check beforehand to confirm pa can hold all tris.
// TODO(Ed): We should do a bounds check beforehand to confirm pa can hold all tris?
// The tape atoms in-flight should not need to care.
// Prepare the tape. (Push protocol to tape)
LP_ U4 mem_temp_tape[512];
TapeBuilder tb = tb_make(slice_ut_arr(mem_temp_tape)); tb_scope(& tb) {
// TODO(Ed): This is bugged.
// tb_emit(& tb, code_set_gte_world);
// tb_data(& tb, u4_(& smem.tform_world));
tb.used = 0; tb_scope(& tb) {
tb_emit(& tb, set_gte_world);
tb_data(& tb, u4_(& smem.tform_world));
tb_emit(& tb, code_rbind_floor_tri);
tb_emit(& tb, rbind_floor_f3_face);
// TODO(Ed): Just use a single context struct ref
tb_data(& tb, prim_cursor);
tb_data(& tb, u4_(smem.floor.faces));
tb_data(& tb, u4_(smem.floor.verts));
tb_data(& tb, u4_(ordering_buf));
for (U4 i = 0; i < Floor_num_faces; i++) {
tb_emit(& tb, code_floor_tri);
tb_emit(& tb, floor_f3_face);
}
// After code_floor_tri iterations complete, the primitive arena's used counter needs updating.
tb_emit(& tb, code_sync_primitive_arena);
// After floor_f3_face iterations complete, the primitive arena's used counter needs updating.
tb_emit(& tb, sync_primitive_arena);
tb_data(& tb, u4_(& pa->used));
tb_data(& tb, prim_base);
}
tape_run(tb_slice(tb));// Fire off the tape.
// C-side state (pa->used) has already been updated by the tape!
@@ -378,23 +385,17 @@ void update(PrimitiveArena* pa, U4* ordering_buf)
TapeBuilder tb = tb_make_old(& tape_arena); tb_scope(& tb) {
// Skip set_gte_world atom for diagnostics to isolate the triangle loop
for (U4 i = 0; i < Floor_num_faces; i++) {
// =======================================================
// SWAP EMIT TO TEST DIFFERENT PARTS OF THE PIPELINE:
// =======================================================
// 1. code_diag_yield -> Tests Tape Engine jump logic
// 2. code_diag_color -> Tests OT and Prim Arena memory
// 3. code_diag_gte -> Tests Vertex arrays and GTE Math
// tb_emit(& tb, code_diag_yield);
// tb_emit(& tb, code_diag_color); //TODO(Ed): Stopped working
// tb_emit(& tb, code_diag_color);
// tb_emit(& tb, code_diag_gte);
}
}
B1* prim_cursor = (B1*)r_(pa->buf)[smem.active_buf_id] + pa->used;
tape_run(tb_slice(tb));
pa->used = (U4)prim_cursor - (U4)r_(pa->buf)[smem.active_buf_id];
smem.floor.rot.y += 5;
}
}
GCC_OPTIMIZATION_ENABLE
int main(void)
{
@@ -63,98 +63,6 @@ U4 vsync(U4 mode) __asm__("VSync");
void draw_orderingtbl(U4* buf) __asm__("DrawOTag");
typedef Struct_(PolyTag) {
U4 addr: 24;
U4 len: 8;
RGB8 color;
B1 code;
};
/*
* Primitive Handling Macros
*/
#define set_len( p, _len) (((PolyTag*R_)(p))->len = (B1)(_len))
#define set_addr(p, _addr) (((PolyTag*R_)(p))->addr = (U4)(_addr))
#define set_code(p, _code) (((PolyTag*R_)(p))->code = (B1)(_code))
#define get_len(p) (B1)(((PolyTag*R_)(p))->len)
#define get_code(p) (B1)(((PolyTag*R_)(p))->code)
#define get_addr(p) (U4)(((PolyTag*R_)(p))->addr)
#define orderingtbl_add_primitive(ot, p) set_addr(p, get_addr(ot)), set_addr(ot, p)
#define orderingtbl_add_primitives(ot, p0, p1) set_addr(p1, get_addr(ot)), set_addr(ot, p0)
/* Primitive Length Code */
#define set_poly_f3(p) set_len(p, 4), set_code(p, 0x20)
#define set_poly_ft3(p) set_len(p, 7), set_code(p, 0x24)
#define set_poly_g3(p) set_len(p, 6), set_code(p, 0x30)
#define set_poly_gt3(p) set_len(p, 9), set_code(p, 0x34)
#define set_poly_f4(p) set_len(p, 5), set_code(p, 0x28)
#define set_poly_ft4(p) set_len(p, 9), set_code(p, 0x2c)
#define set_poly_g4(p) set_len(p, 8), set_code(p, 0x38)
#define set_poly_gt4(p) set_len(p, 12), set_code(p, 0x3c)
// #define setSprt8(p) setlen(p, 3), setcode(p, 0x74)
// #define setSprt16(p) setlen(p, 3), setcode(p, 0x7c)
// #define setSprt(p) setlen(p, 4), setcode(p, 0x64)
// #define setTile1(p) set_len(p, 2), set_code(p, 0x68)
// #define setTile8(p) set_len(p, 2), set_code(p, 0x70)
// #define setTile16(p) set_len(p, 2), set_code(p, 0x78)
#define set_tile(p) set_len(p, 3), set_code(p, 0x60)
// #define setLineF2(p) set_len(p, 3), set_code(p, 0x40)
// #define setLineG2(p) set_len(p, 4), set_code(p, 0x50)
// #define setLineF3(p) set_len(p, 5), set_code(p, 0x48),(p)->pad = 0x55555555
// #define setLineG3(p) set_len(p, 7), set_code(p, 0x58),(p)->pad = 0x55555555, (p)->p2 = 0
// #define setLineF4(p) set_len(p, 6), set_code(p, 0x4c),(p)->pad = 0x55555555
// #define setLineG4(p) set_len(p, 9), set_code(p, 0x5c),(p)->pad = 0x55555555, (p)->p2 = 0, (p)->p3 = 0
typedef Struct_(Poly_F3) {
U4 tag;
RGB8 color;
B1 code;
union {
struct {
V2_S2 p0;
V2_S2 p1;
V2_S2 p2;
};
A3_V2_S2 points;
};
};
typedef Struct_(Poly_G3) {
U4 tag; RGB8 c0; B1 code;
V2_S2 p0; RGB8 c1; B1 pad1;
V2_S2 p1; RGB8 c2; B1 pad2;
V2_S2 p2;
};
typedef Struct_(Poly_F4) {
U4 tag;
RGB8 color;
B1 code;
union {
struct {
V2_S2 p0;
V2_S2 p1;
V2_S2 p2;
V2_S2 p3;
};
A4_V2_S2 points;
};
};
typedef Struct_(Poly_G4) {
U4 tag; RGB8 c0; B1 code;
V2_S2 p0; RGB8 c1; B1 pad1;
V2_S2 p1; RGB8 c2; B1 pad2;
V2_S2 p2; RGB8 c3; B1 pad3;
V2_S2 p3;
};
typedef Struct_(Tile) {
U4 tag;
RGB8 color;
@@ -162,7 +70,6 @@ typedef Struct_(Tile) {
Rect_S2 rect;
};
/*
Linear Algebra
*/
+218
View File
@@ -0,0 +1,218 @@
#ifdef INTELLISENSE_DIRECTIVES
# include "duffle/gen/macs.h"
# include "duffle/gen/offsets.h"
# include "duffle/atom_dsl.h"
# include "duffle/lottes_tape.h"
# include "duffle/word_count.metadata.h"
# include "gen/offsets.h"
# include "hello_gte.h"
#endif
#pragma region MACs (Mips Atom components)
#pragma endregion MACs
#pragma region Baked Atoms
/* DIAGNOSTIC 1: Pure tape loop test */
internal MipsAtom_(diag_yield) { mac_yield() };
/* DIAGNOSTIC 2: Pure memory test (No GTE). Draws a fixed cyan triangle. */
internal MipsAtom_(diag_color) {
store_word( R_0, R_T7, 0),
load_upper_i(R_AT, gp0_cmd_poly_f3 << 8 | 0xFF), /* High: MipsCode Poly_F3(0x20) + Color B:FF */
or_i_self( R_AT, 0xFF00), /* Low: Color G:FF, R:00 (Cyan) */
store_word( R_AT, R_T7, 4),
/* Fake coordinates - Swapped winding order to prevent GPU culling! */
load_upper_i(R_AT, 0x0010), or_i_self(R_AT, 0x0010), store_word(R_AT, R_T7, 8), /* (16, 16) */
load_upper_i(R_AT, 0x0050), or_i_self(R_AT, 0x0010), store_word(R_AT, R_T7, 12), /* (80, 16) */
load_upper_i(R_AT, 0x0010), or_i_self(R_AT, 0x0050), store_word(R_AT, R_T7, 16), /* (16, 80) */
add_ui( R_T1, R_0, 10),
shift_lleft_self(R_T1, S_(U4)/2),
add_u_self( R_T1, R_T6),
load_word( R_AT, R_T1, 0),
load_upper_i(R_V0, (S_(Poly_F3)/S_(U4) - S_(PolyTag)/S_(U4)) << PolyTag_len_bits),
store_word( R_AT, R_T7, 0),
shift_lleft(R_AT, R_T7, S_(PolyTag_len_bits)), shift_lright(R_AT, R_AT, S_(PolyTag_len_bits)),
or_u_self( R_AT, R_V0),
store_word( R_AT, R_T1, 0),
add_ui(R_T7, R_T7, 20),
mac_yield()
};
/* DIAGNOSTIC 3: Pure GTE test (No Memory Writes) */
internal MipsAtom_(diag_gte) {
/* Load 3 indices */
load_half_u(R_T0, R_T4, 0),
load_half_u(R_T1, R_T4, 2),
load_half_u(R_T2, R_T4, 4),
/* Load Vertices into GTE */
shift_lleft( R_AT, R_T0, 3), add_u( R_AT, R_AT, R_T5),
load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4),
gte_mv_to_data_r(R_V0, C2_VXY0), gte_mv_to_data_r(R_V1, C2_VZ0),
shift_lleft( R_AT, R_T1, 3), add_u(R_AT, R_AT, R_T5),
load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4),
gte_mv_to_data_r(R_V0, C2_VXY1), gte_mv_to_data_r(R_V1, C2_VZ1),
shift_lleft(R_AT, R_T2, 3), add_u(R_AT, R_AT, R_T5),
load_word(R_V0, R_AT, 0), load_word(R_V1, R_AT, 4),
gte_mv_to_data_r(R_V0, C2_VXY2), gte_mv_to_data_r(R_V1, C2_VZ2),
/* Run Math */
nop2, gte_cmdw_rtpt,
nop2, gte_cmdw_nclip,
nop2,
/* Advance Face Cursor and Yield */
add_ui(R_T4, R_T4, 8),
mac_yield()
};
typedef Struct_(Binds_CubeTri) {
U4 PrimCursor;
V4_S2* FaceCursor;
V3_S2* VertBase;
U4* OtBase;
};
internal MipsAtom_(rbind_cube_g4_face) atom_info(atom_bind(Binds_CubeTri), atom_phase(cube_g4)
, atom_reads(R_TapePtr)
, atom_writes(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase)
){
/* Pop 4 arguments from the tape directly into the workspace registers */
load_word(R_PrimCursor, R_TapePtr, O_(Binds_CubeTri,PrimCursor)),
load_word(R_FaceCursor, R_TapePtr, O_(Binds_CubeTri,FaceCursor)),
load_word(R_VertBase, R_TapePtr, O_(Binds_CubeTri,VertBase)),
load_word(R_OtBase, R_TapePtr, O_(Binds_CubeTri,OtBase)),
add_ui_self( R_TapePtr, S_(Binds_CubeTri)),
mac_yield()
};
// cube_g4_face — Draw one cube face (Gouraud-shaded quad) via the GTE tape pipeline
internal
MipsAtom_(cube_g4_face) atom_info(atom_phase(cube_g4),
atom_reads( R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase),
atom_writes(R_PrimCursor, R_FaceCursor)
){
load_half_u(R_T0, R_FaceCursor, 0 * S_(S2)),
load_half_u(R_T1, R_FaceCursor, 1 * S_(S2)),
load_half_u(R_T2, R_FaceCursor, 2 * S_(S2)),
load_half_u(R_T3, R_FaceCursor, 3 * S_(S2)),
mac_gte_load_tri_verts(R_T0, R_T1, R_T2),
nop2, gte_cmdw_rotate_translate_perspective_triple, // required cpu -> gte delay slot
gte_cmdw_nclip,
gte_mv_from_data_r(R_T0, C2_MAC0), nop,
branch_le_zero(R_T0, atom_offset(cull, cube_g4_face_exit)), nop,
store_word(R_0, R_PrimCursor, O_(Poly_G4, tag)),
shift_lleft(R_AT, R_T3, v3s2_byteoff), add_u(R_AT, R_AT, R_VertBase),
load_word(R_V0, R_AT, O_(V3_S2, x)), load_word(R_V1, R_AT, O_(V3_S2, z)),
gte_mv_to_data_r(R_V0, C2_VXY0), gte_mv_to_data_r(R_V1, C2_VZ0),
mac_gte_store_g4_p012(),
gte_cmdw_rotate_translate_perspective_single,
mac_gte_store_g4_p3(),
gte_cmdw_avg_sort_z4,
gte_mv_from_data_r(R_T1, C2_OTZ),
add_ui( R_AT, R_0, OrderingTbl_Len),
set_lt_u( R_AT, R_T1, R_AT),
branch_equal(R_AT, R_0, atom_offset(bounds_chk, cube_g4_face_exit)), nop,
mac_insert_ot_tag_g4(),
mac_format_g4_color(
/* c0 magenta */ 0xFF, 0x00, 0xFF,
/* c1 yellow */ 0xFF, 0xFF, 0x00,
/* c2 cyan */ 0x00, 0xFF, 0xFF,
/* c3 green */ 0x00, 0xFF, 0x00),
// end: branch(bounds_chk)
// end: branch(cull)
atom_label(cube_g4_face_exit)
add_ui_self(R_PrimCursor, S_(Poly_G4)), /* 9 words = Poly_G4 */
add_ui_self(R_FaceCursor, S_(S2) * 4), /* 4 × S2 = 8 bytes */
mac_yield()
};
typedef Struct_(Binds_FloorTri) {
U4 PrimCursor;
V3_S2* FaceCursor;
V3_S2* VertBase;
U4* OtBase;
};
internal
MipsAtom_(rbind_floor_f3_face) atom_info(atom_bind(Binds_FloorTri), atom_phase(floor_f3)
, atom_reads(R_TapePtr)
, atom_writes(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase)
){
/* Pop 4 arguments from the tape directly into the workspace registers */
load_word(R_PrimCursor, R_TapePtr, O_(Binds_FloorTri,PrimCursor)),
load_word(R_FaceCursor, R_TapePtr, O_(Binds_FloorTri,FaceCursor)),
load_word(R_VertBase, R_TapePtr, O_(Binds_FloorTri,VertBase)),
load_word(R_OtBase, R_TapePtr, O_(Binds_FloorTri,OtBase)),
add_ui_self( R_TapePtr, S_(Binds_FloorTri)),
mac_yield()
};
// atom_dbg_skip
internal
MipsAtom_(floor_f3_face) atom_info(atom_phase(floor_f3)
, atom_reads( R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase)
, atom_writes(R_PrimCursor, R_FaceCursor)
) {
mac_load_tri_indices( R_T0, R_T1, R_T2),
mac_gte_load_tri_verts(R_T0, R_T1, R_T2),
nop2, gte_cmdw_rotate_translate_perspective_triple, // 2 nops retire the final cpu -> gte writes before RTPT
gte_cmdw_nclip,
/* Culling (Branch forward if Backface) */
gte_mv_from_data_r(R_T0, C2_MAC0),
nop, branch_le_zero(R_T0, atom_offset(culling, floor_f3_face_exit)), nop, // required gte -> cpu load-delay slot.
/* Format Primitive */
mac_gte_store_f3(),
/* Calculate Depth */
gte_avg_sort_z3,
gte_mv_from_data_r(R_T1, C2_OTZ),
/* Bounds Check OTZ < OrderingTbl_Len (Branch forward to skip insertion) */
add_ui( R_AT, R_0, OrderingTbl_Len),
set_lt_u( R_AT, R_T1, R_AT),
branch_equal(R_AT, R_0, atom_offset(bounds_chk, floor_f3_face_exit)), nop,
mac_format_f3_color(0xFF, 0xFF, 0xFF), // RGB-form (R=FF, G=FF, B=FF = white)
mac_insert_ot_tag_f3(), /* Insert into Ordering Table Linked List */
add_ui_self(R_PrimCursor, S_(Poly_F3)), /* Advance Prim Cursor (5 words) */
// Note(Ed): No bounds checking, should be checked before atom runs.
// end: branch(bounds_chk)
// end: branch(culling)
/* Advance Input Cursor & Yield (Both branch targets land here) */
atom_label(floor_f3_face_exit)
add_ui_self(R_FaceCursor, S_(S2) * 4), /* Advance Face Cursor (4 * S2 = 8 bytes) */
mac_yield()
};
typedef Struct_(Binds_SyncPrimitiveArena) { U4 used; U4 cursor; };
internal MipsAtom_(sync_primitive_arena) atom_info(atom_bind(Binds_SyncPrimitiveArena)
, atom_reads( R_TapePtr, R_PrimCursor)
, atom_writes(R_TapePtr)
){
load_word(R_AT, R_TapePtr, O_(Binds_SyncPrimitiveArena,used)),
load_word(R_T0, R_TapePtr, O_(Binds_SyncPrimitiveArena,cursor)),
add_ui_self( R_TapePtr, S_(Binds_SyncPrimitiveArena)),
/* Calculate byte offset and store directly back to RAM */
sub_u( R_T0, R_PrimCursor, R_T0), // R_T0 = R_PrimCursor - binds.cursor
store_word(R_T0, R_AT, 0), // R_AT[0] = R_T0
mac_yield()
};
#pragma endregion Baked Atoms
+41
View File
@@ -0,0 +1,41 @@
#ifdef INTELLISENSE_DIRECTIVES
#pragma once
#endif
// Auto-generated by ps1_meta.lua — DO NOT EDIT
// Directory: C:\projects\Pikuma\ps1\code\hello_joypad/
// source: C:\projects\Pikuma\ps1\code\hello_joypad\hello_joypad.c
// source: C:\projects\Pikuma\ps1\code\hello_joypad\hello_joypad.h
// source: C:\projects\Pikuma\ps1\code\hello_joypad\hello_joypad.atom.c
// Component atoms (MipsAtomComp_(ac_*)) -> macro variants (mac_*)
#ifndef WORD_COUNT
#define WORD_COUNT(name, count) enum { words_##name = (count) };
#endif
#define mac_put_disp_env(reg_transfer, reg_base, port) \
mac_gcmd_push(gp0_word_draw_area_top_left_origin, reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_draw_area_bottom_right_320x240, reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_set_mask_bit(), reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_draw_area_top_left_origin, reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_draw_area_bottom_right_320x240, reg_transfer, reg_base, port)
WORD_COUNT(mac_put_disp_env, 5)
#define mac_put_draw_env(reg_transfer, reg_base, port) \
mac_gcmd_push(gp0_dr_env_tag, reg_transfer, reg_base, port) /* tag (length=15 << 24, addr=0) — packet header for the DR_ENV sequence. The GPU needs this to recognize the next 15 words as a DR_ENV packet and trigger the isbg auto-clear. */ \
, mac_gcmd_push(gp0_word_draw_mode_drawing_allowed, reg_transfer, reg_base, port) /* code[0] DrawMode (dfe=1, dtd=0, tpage=0) */ \
, mac_gcmd_push(gp0_word_set_texture_window(), reg_transfer, reg_base, port) /* code[1] TextureWindow (tw=(0,0)) */ \
, mac_gcmd_push(enc_gp0_draw_area_tl_word(0, ScreenRes_Y), reg_transfer, reg_base, port) /* code[2] DrawArea top-left (clip.x=0, clip.y=ScreenRes_Y=240) */ \
, mac_gcmd_push(gp0_word_draw_area_bottom_right_320x240, reg_transfer, reg_base, port) /* code[3] DrawArea bottom-right (clip.x+w=320, clip.y+h=480) */ \
, mac_gcmd_push(gp0_word_set_draw_offset(), reg_transfer, reg_base, port) /* code[4] DrawOffset (ofs=(0,0)) — bare-cmd word; the GPU uses the current state machine. */ \
, mac_gcmd_push(gp0_word_dr_env_mask(), reg_transfer, reg_base, port) /* code[5] Mask (dtd=0, dfe=1, isbg=1) — 0xE6 cmd + isbg bit. */ \
, mac_gcmd_push(gp0_word_dr_env_bg_color_cmd(1, 7, 7, 7), reg_transfer, reg_base, port) /* code[6] Initial-bg-color + auto-clear (isbg=1, r=7, g=7, b=7). */ \
, mac_gcmd_push(gp0_word_dr_env_draw_mode(1), reg_transfer, reg_base, port) /* code[7] Re-assert DrawMode with isbg=1 (isbg-flag set; the 0xE1 cmd byte plus isbg only). */ /* code[8..10] Padding (NOP — GPU discards; the DR_ENV requires 16 words total). */ \
, mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port) /* code[11..12] TextureWindow bottom-right (tw.x+tw.w=0, tw.y+tw.h=0) — libpsyx emits twice. */ \
, mac_gcmd_push(gp0_word_set_texture_window(), reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_set_texture_window(), reg_transfer, reg_base, port) /* code[13..14] Padding (NOP) — completes the 16-word packet. */ \
, mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port) \
, mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port)
WORD_COUNT(mac_put_draw_env, 16)
+76
View File
@@ -0,0 +1,76 @@
// Auto-generated by ps1_meta.lua (passes/offsets.lua) — DO NOT EDIT
// Directory: C:\projects\Pikuma\ps1\code\hello_joypad\
// source: C:\projects\Pikuma\ps1\code\hello_joypad\hello_joypad.c
// source: C:\projects\Pikuma\ps1\code\hello_joypad\hello_joypad.h
// source: C:\projects\Pikuma\ps1\code\hello_joypad\hello_joypad.atom.c
#pragma once
#pragma region hello_joypad
// --- atom: cube_g4_face (76 words) ---
#define _atom_offset_cull_cube_g4_face_exit 41
#define _atom_offset_bounds_chk_cube_g4_face_exit 24
enum {
atom_offset_cull_cube_g4_face_exit = _atom_offset_cull_cube_g4_face_exit,
atom_offset_bounds_chk_cube_g4_face_exit = _atom_offset_bounds_chk_cube_g4_face_exit,
};
// --- atom: floor_f3_face (58 words) ---
#define _atom_offset_culling_floor_f3_face_exit 25
#define _atom_offset_bounds_chk_floor_f3_face_exit 16
enum {
atom_offset_culling_floor_f3_face_exit = _atom_offset_culling_floor_f3_face_exit,
atom_offset_bounds_chk_floor_f3_face_exit = _atom_offset_bounds_chk_floor_f3_face_exit,
};
// --- atom: pad_bios_snapshot (78 words) ---
#define _atom_offset_snap_root_skip_disconnected 8
#define _atom_offset_disconnected_snap_end 61
#define _atom_offset_case_2_id_dispatch 8
#define _atom_offset_pending_snap_end 51
#define _atom_offset_id_dispatch_try_analog_stick 11
#define _atom_offset_id_dispatch_snap_end 38
#define _atom_offset_try_analog_stick_try_analog_pad 12
#define _atom_offset_analog_stick_snap_end 24
#define _atom_offset_try_analog_pad_try_unsupported 11
#define _atom_offset_analog_pad_snap_end 10
enum {
atom_offset_snap_root_skip_disconnected = _atom_offset_snap_root_skip_disconnected,
atom_offset_disconnected_snap_end = _atom_offset_disconnected_snap_end,
atom_offset_case_2_id_dispatch = _atom_offset_case_2_id_dispatch,
atom_offset_pending_snap_end = _atom_offset_pending_snap_end,
atom_offset_id_dispatch_try_analog_stick = _atom_offset_id_dispatch_try_analog_stick,
atom_offset_id_dispatch_snap_end = _atom_offset_id_dispatch_snap_end,
atom_offset_try_analog_stick_try_analog_pad = _atom_offset_try_analog_stick_try_analog_pad,
atom_offset_analog_stick_snap_end = _atom_offset_analog_stick_snap_end,
atom_offset_try_analog_pad_try_unsupported = _atom_offset_try_analog_pad_try_unsupported,
atom_offset_analog_pad_snap_end = _atom_offset_analog_pad_snap_end,
};
// --- atom: pad_apply_input (60 words) ---
#define _atom_offset_dpad_left_exit_dpad_left 6
#define _atom_offset_dpad_right_exit_dpad_right 6
#define _atom_offset_dead_zone_low_check_dead_low_active 8
#define _atom_offset_dead_zone_high_check_dead_high_active 15
#define _atom_offset_dead_zone_skip_exit_stick 24
#define _atom_offset_end_low_exit_stick 12
enum {
atom_offset_dpad_left_exit_dpad_left = _atom_offset_dpad_left_exit_dpad_left,
atom_offset_dpad_right_exit_dpad_right = _atom_offset_dpad_right_exit_dpad_right,
atom_offset_dead_zone_low_check_dead_low_active = _atom_offset_dead_zone_low_check_dead_low_active,
atom_offset_dead_zone_high_check_dead_high_active = _atom_offset_dead_zone_high_check_dead_high_active,
atom_offset_dead_zone_skip_exit_stick = _atom_offset_dead_zone_skip_exit_stick,
atom_offset_end_low_exit_stick = _atom_offset_end_low_exit_stick,
};
#pragma endregion hello_joypad
+638
View File
@@ -0,0 +1,638 @@
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "duffle/gen/macs.h"
# include "duffle/gen/offsets.h"
# include "duffle/dsl.atom.h"
# include "duffle/lottes_tape.h"
# include "duffle/mips.h"
# include "duffle/gte.h"
# include "duffle/gp.h"
# include "duffle/pad.h"
# include "duffle/word_count.metadata.h"
# include "duffle/psyq.h"
# include "duffle/math.atom.c"
# include "duffle/mips.atom.c"
# include "duffle/gte.atom.c"
# include "duffle/gp.atom.c"
# include "duffle/psyq.atom.c"
# include "gen/offsets.h"
# include "gen/macs.h"
# include "hello_joypad.h"
#endif
ATOM_FILE_DEBUGGER_LINE_MARKER(hello_joypad_atom_c);
#pragma region MACs (Mips Atom components)
FI_ Slice_MipsCode ac_put_disp_env(U4 reg_transfer, U4 reg_base, U2 port)
MipsAtomComp_Proc_(ac_put_disp_env, {
// Emits 5 GP0 commands for buffer 0 (display_area = (0,0,320,240)).
// Sequence per libpsyx PutDispEnv: DrawArea TL → DrawArea BR → Mask → DrawArea TL → DrawArea BR
mac_gcmd_push(gp0_word_draw_area_top_left_origin, reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_draw_area_bottom_right_320x240, reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_set_mask_bit(), reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_draw_area_top_left_origin, reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_draw_area_bottom_right_320x240, reg_transfer, reg_base, port),
})
FI_ Slice_MipsCode ac_put_draw_env(U4 reg_transfer, U4 reg_base, U2 port)
MipsAtomComp_Proc_(ac_put_draw_env, {
/*
* ORIGIN: each code word corresponds to the EXACT value libpsyx's PutDrawEnv function would compute for the same DrawEnv settings.
* References:
* - libpsyx source: `toolchain/psyq-4_7/lib/libgpu.a` (binary, function `PutDrawEnv`)
* - PSX-SPX doc: https://problemkaputt.de/psx-spx.htm#gputdrawingcommands
* - PSYQ SDK: `setdrawenv` / `makelongdr_env` source
* - NOCASH PSX spec: §"GP0(E1h) Draw Mode setting" through §"DR_ENV"
*
* The 16-word format is documented in the PSYQ SDK manual and on NOCASH's PSX-spec.txt. The libpsyx reference is at:
* ./toolchain/psyq-4_7/lib/libgpu.a
* (binary; the PutDrawEnv implementation builds the 16-word DR_ENV from the user's DRAWENV struct and emits it via GP0 GPU commands.)
*
* Word indices (libpsyx PutDrawEnv / SetDrawEnv order):
* tag = (length << 24) | addr — 16-word packet (1 tag + 15 code)
* code[0] = DrawMode (dfe=1, dtd=0, tpage=0) — must come first per libpsyx
* code[1] = TextureWindow (tw=(0,0)) — bare-cmd word; GPU uses current state
* code[2] = DrawArea top-left (clip.x=0, clip.y=240)
* code[3] = DrawArea bottom-right (clip.x+w=320, clip.y+h=480)
* code[4] = DrawOffset (ofs=(0,0)) — bare-cmd word
* code[5] = Mask (dtd=0, dfe=1, isbg=1) — 0xE6 cmd + isbg bit
* code[6] = Initial-bg-color (isbg=1, r=7, g=7, b=7)
* code[7] = DrawMode (isbg=1, tpage=0) — re-asserts DrawMode with isbg
* code[8..10] = padding (NOP) — 3 words to fill the packet
* code[11..12] = TextureWindow bottom-right — defaults to (0,0,0,0)
* code[13..14] = padding (NOP) — completes the 16-word packet
*/
mac_gcmd_push(gp0_dr_env_tag, reg_transfer, reg_base, port), /* tag (length=15 << 24, addr=0) — packet header for the DR_ENV sequence. The GPU needs this to recognize the next 15 words as a DR_ENV packet and trigger the isbg auto-clear. */
mac_gcmd_push(gp0_word_draw_mode_drawing_allowed, reg_transfer, reg_base, port), /* code[0] DrawMode (dfe=1, dtd=0, tpage=0) */
mac_gcmd_push(gp0_word_set_texture_window(), reg_transfer, reg_base, port), /* code[1] TextureWindow (tw=(0,0)) */
mac_gcmd_push(enc_gp0_draw_area_tl_word(0, ScreenRes_Y), reg_transfer, reg_base, port), /* code[2] DrawArea top-left (clip.x=0, clip.y=ScreenRes_Y=240) */
mac_gcmd_push(gp0_word_draw_area_bottom_right_320x240, reg_transfer, reg_base, port), /* code[3] DrawArea bottom-right (clip.x+w=320, clip.y+h=480) */
mac_gcmd_push(gp0_word_set_draw_offset(), reg_transfer, reg_base, port), /* code[4] DrawOffset (ofs=(0,0)) — bare-cmd word; the GPU uses the current state machine. */
mac_gcmd_push(gp0_word_dr_env_mask(), reg_transfer, reg_base, port), /* code[5] Mask (dtd=0, dfe=1, isbg=1) — 0xE6 cmd + isbg bit. */
mac_gcmd_push(gp0_word_dr_env_bg_color_cmd(1, 7, 7, 7), reg_transfer, reg_base, port), /* code[6] Initial-bg-color + auto-clear (isbg=1, r=7, g=7, b=7). */
mac_gcmd_push(gp0_word_dr_env_draw_mode(1), reg_transfer, reg_base, port), /* code[7] Re-assert DrawMode with isbg=1 (isbg-flag set; the 0xE1 cmd byte plus isbg only). */
/* code[8..10] Padding (NOP — GPU discards; the DR_ENV requires 16 words total). */
mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port),
/* code[11..12] TextureWindow bottom-right (tw.x+tw.w=0, tw.y+tw.h=0) — libpsyx emits twice. */
mac_gcmd_push(gp0_word_set_texture_window(), reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_set_texture_window(), reg_transfer, reg_base, port),
/* code[13..14] Padding (NOP) — completes the 16-word packet. */
mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port),
mac_gcmd_push(gp0_word_nop(), reg_transfer, reg_base, port),
})
#pragma endregion MACs
#pragma region Baked Atoms
enum {
R_ScreenX = R_T5 atom_reg atom_type(U2),
R_ScreenY = R_T6 atom_reg atom_type(U2),
R_ScreenBuf = R_T7 atom_reg, /* Caller-pinned: & smem.screen_buf */
#define R_ScreenBuf_Code R_T7_Code
};
//screen_env_init. Mirrors the libpsyx's SetDefDispEnv + SetDefDrawEnv + the manual enable_auto_clear / initial_bg_color writes.
internal MipsAtom_(screen_env_init) atom_info(atom_phase(screen_init)
, atom_reads(R_T0, R_ScreenX, R_ScreenY, R_ScreenBuf)
, atom_writes(R_T0, R_ScreenX, R_ScreenY)
) {
/* display[0] = (0, 0, 320, 240); rest of struct zeroed. */
add_ui(R_ScreenX, R_0, ScreenRes_X), add_ui(R_ScreenY, R_0, ScreenRes_Y),
mac_store_v2s2(R_ScreenX, R_ScreenY, R_ScreenBuf, O_(DisplayEnv,display_area.width) + OA_(DoubleBuffer,display,0)),
store_word(R_0, R_ScreenBuf, O_(DisplayEnv,display_area) + OA_(DoubleBuffer,display,0)),
store_word(R_0, R_ScreenBuf, O_(DisplayEnv,screen) + OA_(DoubleBuffer,display,0)),
store_word(R_0, R_ScreenBuf, O_(DisplayEnv,vinterlace) + OA_(DoubleBuffer,display,0)),
/* display[1] = (0, 240, 320, 240); rest of struct zeroed. */
mac_store_rects2(R_0, R_ScreenY, R_ScreenX, R_ScreenY, R_ScreenBuf, O_(DisplayEnv,display_area) + OA_(DoubleBuffer,display,1)),
store_word(R_0, R_ScreenBuf, O_(DisplayEnv,screen) + OA_(DoubleBuffer,display,1)),
store_word(R_0, R_ScreenBuf, O_(DisplayEnv,vinterlace) + OA_(DoubleBuffer,display,1)),
mac_store_rects2(R_0, R_ScreenY, R_ScreenX, R_ScreenY, R_ScreenBuf, O_(DrawEnv,clip_area) + OA_(DoubleBuffer,draw,0)), /* draw[0].clip_area = (0, 240, 320, 240). C11's SetDefDrawEnv writes clip.y = y_arg. */
mac_store_v2s2( R_0, R_ScreenY, R_ScreenBuf, O_(DrawEnv,drawing_offset[0]) + OA_(DoubleBuffer,draw,0)), /* draw[0].drawing_offset[0] = (0, 240); C11 passes y_arg as ofs. */
mac_store_v2s2(R_ScreenX, R_ScreenY, R_ScreenBuf, O_(DrawEnv,clip_area.width) + OA_(DoubleBuffer,draw,1)),
/* draw[0].texture_window = (0, 0, 0, 0); two word-zeroes cover the full 8-byte tw field. */
store_word(R_0, R_ScreenBuf, O_(DrawEnv,texture_window.x) + OA_(DoubleBuffer,draw,0)),
store_word(R_0, R_ScreenBuf, O_(DrawEnv,texture_window.width) + OA_(DoubleBuffer,draw,0)),
store_word(R_0, R_ScreenBuf, O_(DrawEnv,drawing_offset[0].x) + OA_(DoubleBuffer,draw,1)),
store_word(R_0, R_ScreenBuf, O_(DrawEnv,texture_window.x) + OA_(DoubleBuffer,draw,1)),
store_word(R_0, R_ScreenBuf, O_(DrawEnv,texture_window.width) + OA_(DoubleBuffer,draw,1)),
/* draw[0].texture_page = 10 (gp0_tpage_default). C11 SetDefDrawEnv at C11_only.elf:0x8001273C writes the same 0x0A. . */
add_ui(R_T0, R_0, gp0_tpage_default),
store_half(R_T0, R_ScreenBuf, O_(DrawEnv,texture_page) + OA_(DoubleBuffer,draw,0)),
store_half(R_T0, R_ScreenBuf, O_(DrawEnv,texture_page) + OA_(DoubleBuffer,draw,1)),
/* draw[0] control bytes: flag_dither=1, flag_draw_on_display=1 (the dfe bit per psx-spx; libpsyx sets it via `SetDefDrawEnv`'s conditional at C11_only.elf:0x80012728), enable_auto_clear=1. Each byte is named;
* the previous `store_word(R_0, ..., +20)` overwrote all four with zero. */
add_ui(R_T0, R_0, 1),
store_byte(R_T0, R_ScreenBuf, O_(DrawEnv,flag_dither) + OA_(DoubleBuffer,draw,0)),
store_byte(R_T0, R_ScreenBuf, O_(DrawEnv,flag_draw_on_display) + OA_(DoubleBuffer,draw,0)),
store_byte(R_T0, R_ScreenBuf, O_(DrawEnv,enable_auto_clear) + OA_(DoubleBuffer,draw,0)),
store_byte(R_T0, R_ScreenBuf, O_(DrawEnv,flag_dither) + OA_(DoubleBuffer,draw,1)),
store_byte(R_T0, R_ScreenBuf, O_(DrawEnv,flag_draw_on_display) + OA_(DoubleBuffer,draw,1)),
store_byte(R_T0, R_ScreenBuf, O_(DrawEnv,enable_auto_clear) + OA_(DoubleBuffer,draw,1)),
/* draw[0].initial_bg_color = (r=7, g=7, b=7). */
add_ui(R_T0, R_0, 7),
mac_store_rgb8(R_T0,R_T0,R_T0, R_ScreenBuf, O_(DrawEnv,initial_bg_color) + OA_(DoubleBuffer,draw,0)),
mac_store_rgb8(R_T0,R_T0,R_T0, R_ScreenBuf, O_(DrawEnv,initial_bg_color) + OA_(DoubleBuffer,draw,1)),
mac_yield(),
};
enum {
R_IO_BaseAddr = R_T4 atom_reg, /* Caller-pinned: IO_BASE_ADDR = 0x1F800000 */
#define R_IO_BaseAddr_Code R_T4_Code
};
internal MipsAtom_(gp_screen_init) atom_info(atom_phase(screen_init), atom_reads(R_IO_BaseAddr)) {
store_word(R_0, R_IO_BaseAddr, GPIO_PORT1_OFFSET), /* GP1(00h) Reset */
mac_gcmd_push(gp1_word_ResetCmdBuffer(), R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET), /* GP1(01h) ClearFIFO */
mac_gcmd_push(gp1_word_AcknowledgeIRQ(), R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET), /* GP1(02h) AckIRQ */
mac_gcmd_push(gp1_word_DisplayOn(), R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET), /* GP1(03h) Display ON */
mac_gcmd_push(gp1_word_dma_to_gpu(), R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET), /* GP1(04h) DMADirection=2 (CPU→GPU). libpsyx's per-frame PutDrawEnv/DrawOTag use DMA2; without this the DMA queue never drains. */
mac_gcmd_push(gp1_word_StartDisplayArea(), R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET), /* GP1(05h) StartDisplayArea (X=0, Y=0) */
/* GP1: DisplayMode + Display Ranges */
mac_gcmd_push(gp1_word_display_mode_320x240_15bit_ntsc, R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET),
mac_gcmd_push(gp1_word_horizontal_range_ntsc, R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET),
mac_gcmd_push(gp1_word_vertical_range_ntsc, R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET),
/* GTE: SetGeomOffset (OFX, OFY) — ScreenRes_CenterX, ScreenRes_CenterY. */
load_upper_i(R_T5, ScreenRes_CenterX), gte_mv_to_ctrl_r(R_T5, gte_cr_OFX_Code),
load_upper_i(R_T5, ScreenRes_CenterY), gte_mv_to_ctrl_r(R_T5, gte_cr_OFY_Code),
/* GTE: SetGeomScreen (H) — CR26 (per PSX-SPX / libpsyx), value is the raw projection-plane distance, NOT shifted. */
add_ui(R_T5, R_0, ScreenZ), gte_mv_to_ctrl_r(R_T5, gte_cr_H_Code),
/* GP1: DisplayEnable — bit 0 = 0 (Display ON). */
mac_gcmd_push(gp1_word_DisplayOn(), R_T5, R_IO_BaseAddr, GPIO_PORT1_OFFSET),
mac_yield(),
};
enum {
R_PrimCursor = R_T7 atom_reg atom_type(U4*), /* VRAM output cursor (primitive buffer) */
R_FaceCursor = R_T4 atom_reg atom_type(V4_S2*), /* Cube face-index cursor (V4_S2*); floor context switches to V3_S2* via atom_phase */
R_VertBase = R_T5 atom_reg atom_type(V3_S2*), /* Base address of the vertex array */
R_OtBase = R_T6 atom_reg atom_type(U4*), /* Base address of the Ordering Table */
#define R_PrimCursor_Code R_T7_Code
#define R_FaceCursor_Code R_T4_Code
#define R_VertBase_Code R_T5_Code
#define R_OtBase_Code R_T6_Code
};
typedef Struct_(Binds_CubeTri) {
U4 PrimCursor;
V4_S2* FaceCursor;
V3_S2* VertBase;
U4* OtBase;
};
internal MipsAtom_(rbind_cube_g4_face) atom_info(atom_bind(Binds_CubeTri), atom_phase(cube_g4)
, atom_reads(R_TapePtr)
, atom_writes(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase, R_TapePtr)
){
/* Pop 4 arguments from the tape directly into the workspace registers */
load_word(R_PrimCursor, R_TapePtr, O_(Binds_CubeTri,PrimCursor)),
load_word(R_FaceCursor, R_TapePtr, O_(Binds_CubeTri,FaceCursor)),
load_word(R_VertBase, R_TapePtr, O_(Binds_CubeTri,VertBase)),
load_word(R_OtBase, R_TapePtr, O_(Binds_CubeTri,OtBase)),
add_ui_self( R_TapePtr, S_(Binds_CubeTri)),
mac_yield()
};
// cube_g4_face — Draw one cube face (Gouraud-shaded quad) via the GTE tape pipeline
internal
MipsAtom_(cube_g4_face) atom_info(atom_phase(cube_g4),
atom_reads( R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase),
atom_writes(R_PrimCursor, R_FaceCursor)
){
load_half_u(R_T0, R_FaceCursor, 0 * S_(S2)),
load_half_u(R_T1, R_FaceCursor, 1 * S_(S2)),
load_half_u(R_T2, R_FaceCursor, 2 * S_(S2)),
load_half_u(R_T3, R_FaceCursor, 3 * S_(S2)),
mac_gte_load_tri_verts(R_VertBase, R_T0, R_T1, R_T2),
nop2, gte_cmdw_rotate_translate_perspective_triple, // required cpu -> gte delay slot
gte_cmdw_nclip,
gte_mv_from_data_r(R_T0, C2_MAC0), nop,
branch_le_zero(R_T0, atom_offset(cull, cube_g4_face_exit)),
/* BD-slot: write the prim tag (R_0=0; overwrites the legacy tag word in the prim_buffer).
* If branch IS taken (face culled), the body is skipped and this 0-tag is stranded —
* harmless because the OT entry that points to this prim is created later, only on the body path. */
store_word(R_0, R_PrimCursor, O_(Poly_G4, tag)),
shift_lleft(R_AT, R_T3, v3s2_byteoff), add_u(R_AT, R_AT, R_VertBase),
load_word(R_V0, R_AT, O_(V3_S2, x)), load_word(R_V1, R_AT, O_(V3_S2, z)),
gte_mv_to_data_r(R_V0, C2_VXY0), gte_mv_to_data_r(R_V1, C2_VZ0),
mac_gte_store_g4_p012(R_PrimCursor),
gte_cmdw_rotate_translate_perspective_single,
mac_gte_store_g4_p3(R_PrimCursor),
gte_cmdw_avg_sort_z4,
gte_mv_from_data_r(R_T1, C2_OTZ),
add_ui( R_AT, R_0, OrderingTbl_Len),
set_lt_u( R_AT, R_T1, R_AT),
branch_equal(R_AT, R_0, atom_offset(bounds_chk, cube_g4_face_exit)), nop,
mac_insert_ot_tag_g4(R_OtBase, R_PrimCursor),
mac_format_g4_color(R_PrimCursor,
/* c0 magenta */ 0xFF, 0x00, 0xFF,
/* c1 yellow */ 0xFF, 0xFF, 0x00,
/* c2 cyan */ 0x00, 0xFF, 0xFF,
/* c3 green */ 0x00, 0xFF, 0x00),
// end: branch(bounds_chk)
// end: branch(cull)
atom_label(cube_g4_face_exit)
add_ui_self(R_PrimCursor, S_(Poly_G4)), /* 9 words = Poly_G4 */
add_ui_self(R_FaceCursor, S_(S2) * 4), /* 4 × S2 = 8 bytes */
mac_yield()
};
typedef Struct_(Binds_FloorTri) {
U4 PrimCursor;
V3_S2* FaceCursor;
V3_S2* VertBase;
U4* OtBase;
};
internal
MipsAtom_(rbind_floor_f3_face) atom_info(atom_bind(Binds_FloorTri), atom_phase(floor_f3)
, atom_reads(R_TapePtr)
, atom_writes(R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase, R_TapePtr)
){
/* Pop 4 arguments from the tape directly into the workspace registers */
load_word(R_PrimCursor, R_TapePtr, O_(Binds_FloorTri,PrimCursor)),
load_word(R_FaceCursor, R_TapePtr, O_(Binds_FloorTri,FaceCursor)),
load_word(R_VertBase, R_TapePtr, O_(Binds_FloorTri,VertBase)),
load_word(R_OtBase, R_TapePtr, O_(Binds_FloorTri,OtBase)),
add_ui_self( R_TapePtr, S_(Binds_FloorTri)),
mac_yield()
};
// atom_dbg_skip
internal
MipsAtom_(floor_f3_face) atom_info(atom_phase(floor_f3)
, atom_reads( R_PrimCursor, R_FaceCursor, R_VertBase, R_OtBase)
, atom_writes(R_PrimCursor, R_FaceCursor)
) {
mac_load_tri_indices( R_FaceCursor, R_T0, R_T1, R_T2),
mac_gte_load_tri_verts(R_VertBase, R_T0, R_T1, R_T2),
nop2, gte_cmdw_rotate_translate_perspective_triple, // 2 nops retire the final cpu -> gte writes before RTPT
gte_cmdw_nclip,
/* Culling (Branch forward if Backface) */
gte_mv_from_data_r(R_T0, C2_MAC0),
nop, branch_le_zero(R_T0, atom_offset(culling, floor_f3_face_exit)), nop, // required gte -> cpu load-delay slot.
/* Format Primitive */
mac_gte_store_f3(R_PrimCursor),
/* Calculate Depth */
gte_avg_sort_z3,
gte_mv_from_data_r(R_T1, C2_OTZ),
/* Bounds Check OTZ < 2048 (Branch forward to skip insertion) */
add_ui( R_AT, R_0, OrderingTbl_Len),
set_lt_u( R_AT, R_T1, R_AT),
branch_equal(R_AT, R_0, atom_offset(bounds_chk, floor_f3_face_exit)), nop,
mac_format_f3_color(R_PrimCursor, 0xFF, 0xFF, 0xFF), // RGB-form (R=FF, G=FF, B=FF = white)
mac_insert_ot_tag_f3(R_OtBase, R_PrimCursor), /* Insert into Ordering Table Linked List */
add_ui_self(R_PrimCursor, S_(Poly_F3)), /* Advance Prim Cursor (5 words) */
// Note(Ed): No bounds checking, should be checked before atom runs.
// end: branch(bounds_chk)
// end: branch(culling)
/* Advance Input Cursor & Yield (Both branch targets land here) */
atom_label(floor_f3_face_exit)
add_ui_self(R_FaceCursor, S_(S2) * 4), /* Advance Face Cursor (4 * S2 = 8 bytes) */
mac_yield()
};
typedef Struct_(Binds_SyncPrimitiveArena) { U4 used; U4 cursor; };
internal MipsAtom_(sync_primitive_arena) atom_info(atom_bind(Binds_SyncPrimitiveArena)
, atom_reads( R_TapePtr, R_PrimCursor)
, atom_writes(R_TapePtr)
){
load_word(R_AT, R_TapePtr, O_(Binds_SyncPrimitiveArena,used)),
load_word(R_T0, R_TapePtr, O_(Binds_SyncPrimitiveArena,cursor)),
add_ui_self( R_TapePtr, S_(Binds_SyncPrimitiveArena)),
/* Calculate byte offset and store directly back to RAM */
sub_u( R_T0, R_PrimCursor, R_T0), // R_T0 = R_PrimCursor - binds.cursor
store_word(R_T0, R_AT, 0), // R_AT[0] = R_T0
mac_yield()
};
/* ----- pad_bios_snapshot -----
* Per-frame snapshot of one BIOS pad buffer into PadState.
* Decoder (branch ladder on raw[0] status + raw[1] id):
* 1. raw[0] == 0xFF -> Disconnected (buttons=0, axes=0x80)
* 2. raw[0]==0 && raw[1]==0 -> Pending (buttons=0, axes=0x80)
* 3. raw[1] == 0x41 -> Digital (buttons normalized; axes=0x80)
* 4. raw[1] == 0x53 -> AnalogStick (buttons normalized; axes from raw[4..7])
* 5. raw[1] in 0x7x -> AnalogPad (buttons normalized; axes from raw[4..7])
* 6. else -> Unsupported (buttons=0, axes=0x80)
*
* Buttons normalization: byte_swap16((~raw_buttons) & 0xFFFF).
* raw_buttons = load_half_u(raw, 2) = raw[2] | (raw[3] << 8).
* byte_swap16(x) = (x >> 8) | (x << 8); nor(x, R_0) = ~x. store_half truncates to 16 bits so the upper-16 mask is implicit in the store.
*
* Register use (atom-local; no wave-context touched):
* R_T0 = raw base (kept throughout; axes loads read raw[4..7] from R_T0)
* R_T1 = state base (kept throughout; all stores go through R_T1)
* R_T2 = raw[0] status (alive across the disc/pending/id dispatch, then dead)
* R_T3 = raw[1] id (alive across the id dispatch, then dead)
* R_T4 = scratch (shifts, compares, immediate loads, store values)
* R_T5 = scratch (parallel lui+ori for the 0x80808080 axes constant + byte-swap target)
*/
enum {
R_PadRaw = R_T0 atom_reg atom_type(U1),
R_PadState = R_T1 atom_reg,
R_RawStatus = R_T2 atom_reg,
R_RawId = R_T3 atom_reg,
};
typedef Struct_(Binds_PadBiosSnapshot) {
PadBiosRaw* raw;
PadState* state;
};
internal MipsAtom_(pad_bios_snapshot) atom_info(atom_bind(Binds_PadBiosSnapshot)
, atom_reads( R_PadRaw, R_PadState, R_RawStatus, R_RawId, R_T4, R_T5, R_TapePtr)
, atom_writes(R_PadRaw, R_PadState, R_RawStatus, R_RawId, R_T4, R_T5, R_TapePtr)
) {
/* === Bind consumption: T0 = raw, T1 = state, advance R_TapePtr by 8. */
load_word(R_PadRaw, R_TapePtr, O_(Binds_PadBiosSnapshot,raw)),
load_word(R_PadState, R_TapePtr, O_(Binds_PadBiosSnapshot,state)),
add_ui_self( R_TapePtr, S_(Binds_PadBiosSnapshot)),
/* === Read raw[0] (status) + raw[1] (id) */
load_byte_u(R_RawStatus, R_PadRaw, 0),
load_byte_u(R_RawId, R_PadRaw, 1),
atom_label(snap_root) /* === Case 1: Disconnected (status == 0xFF). */
add_ui(R_T4, R_0, 0xFF), branch_ne(R_RawStatus, R_T4, atom_offset(snap_root, skip_disconnected)),
/* BD-slot: pre-compute PadStatus_Disconnected. Branch reads R_T4=0xFF in EX before this WB completes.
* If branch NOT taken (fall through to pending/id_dispatch), R_T4 is overwritten by the next case body's add_ui — harmless. */
atom_label(disconnected) /* === Disconnected body. */
/* R_T4 = PadStatus_Disconnected from snap_root BD-slot. */
store_word(R_T4, R_PadState, O_(PadState,status)),
store_half(R_0, R_PadState, O_(PadState,buttons)),
/* axes = 0x80808080 (centered) — single sw writes the 4-byte axes block at offset 8 (left_x, left_y, right_x, right_y). */
load_upper_i(R_T4, 0x8080), or_i_self(R_T4, 0x8080),
store_word( R_T4, R_PadState, O_(PadState,left_x)),
store_byte( R_RawId, R_PadState, O_(PadState,id)),
jump_rel(atom_offset(disconnected, snap_end)),
/* BD-slot: load next atom's entry point (replaces the nop).
* The unconditional branch always jumps to snap_end, where mac_yield_tail()
* transfers control to R_AtomJmp without re-loading it. */
mac_yield_load(),
atom_label(skip_disconnected)
/* === Case 2: Pending (status == 0 && id == 0)
* Combined check: if (status | id) != 0 then skip to id_dispatch.
* Falls through to the Pending case only when both are zero. */
or_u_self(R_RawStatus, R_RawId), branch_ne(R_RawStatus, R_0, atom_offset(case_2, id_dispatch)),
/* BD-slot: pre-compute PadStatus_Pending. Branch reads R_RawStatus in EX before this WB completes.
* If branch NOT taken (fall through to id_dispatch), R_T4 is overwritten by the digital/analog body add_ui — harmless. */
atom_label(pending) /* === Pending body */
/* R_T4 = PadStatus_Pending from case_2 BD-slot. */
store_word(R_T4, R_PadState, O_(PadState,status)),
store_half(R_0, R_PadState, O_(PadState,buttons)),
/* axes = 0x80808080 (centered) — single sw writes the 4-byte axes block at offset 8 (left_x, left_y, right_x, right_y). */
load_upper_i(R_T4, 0x8080), or_i_self(R_T4, 0x8080),
store_word( R_T4, R_PadState, O_(PadState,left_x)),
store_byte( R_RawId, R_PadState, O_(PadState,id)),
jump_rel(atom_offset(pending, snap_end)),
mac_yield_load(),
atom_label(id_dispatch) /* === Case 3-6: ID dispatch */
add_ui(R_T4, R_0, 0x41), branch_ne(R_RawId, R_T4, atom_offset(id_dispatch, try_analog_stick)),
/* BD-slot: pre-compute PadStatus_Digital. Branch reads R_RawId in EX before this WB completes.
* If branch NOT taken (fall through to try_analog_stick), R_T4 is overwritten by the analog body add_ui. */
/* === Digital body (status, buttons normalize, axes=0x80, id, branch. */
/* R_T4 = PadStatus_Digital from id_dispatch BD-slot. */
store_word( R_T4, R_PadState, O_(PadState,status)),
load_half_u(R_T4, R_PadRaw, 2 * S_(U1)),
/* Fill R_T4's load-delay slot with the 0x80808080 axes constant into R_T5
* (R_T5 is dead on this path; it's only consumed at the analog_pad range check). */
load_upper_i(R_T5, 0x8080), or_i_self(R_T5, 0x8080),
nor_u( R_T4, R_T4, R_0), /* raw_buttons is already in host bit order; no swap needed */
store_half( R_T4, R_PadState, O_(PadState,buttons)),
/* axes = 0x80808080 (centered) — single sw writes the 4-byte axes block at offset 8 (left_x, left_y, right_x, right_y). */
store_word( R_T5, R_PadState, O_(PadState,left_x)),
add_ui( R_T4, R_0, 0x41),
store_byte( R_T4, R_PadState, O_(PadState,id)),
jump_rel(atom_offset(id_dispatch, snap_end)),
mac_yield_load(),
atom_label(try_analog_stick) /* === Case 4: AnalogStick (id == 0x53)*/
add_ui(R_T4, R_0, 0x53), branch_ne(R_RawId, R_T4, atom_offset(try_analog_stick, try_analog_pad)),
/* BD-slot: pre-compute PadStatus_AnalogStick. Branch reads R_RawId in EX before this WB completes.
* If branch NOT taken (fall through to try_analog_pad), R_T4 is overwritten by the analog_pad body add_ui. */
atom_label(analog_stick) /* === AnalogStick body
* Axes are loaded as two halfwords: raw[6..7] → left_xy (sh at offset 8), raw[4..5] → right_xy (sh at offset 10).
* R_T5 holds left_xy / id-value in turn (it's dead on this path — only consumed at the analog_pad range check). */
/* R_T4 = PadStatus_AnalogStick from try_analog_stick BD-slot. */
store_word( R_T4, R_PadState, O_(PadState,status)),
load_half_u( R_T4, R_PadRaw, 2 * S_(U1)), /* R_T4 = raw_buttons */
load_half_u( R_T5, R_PadRaw, 6 * S_(U1)), /* R_T5 = left_xy; fills R_T4's load-delay slot (doesn't read R_T4) */
nor_u( R_T4, R_T4, R_0), /* R_T4 = ~raw_buttons */
store_half( R_T4, R_PadState, O_(PadState,buttons)),
load_half_u( R_T4, R_PadRaw, 4 * S_(U1)), /* R_T4 = right_xy; fills R_T5's load-delay slot */
store_half( R_T5, R_PadState, O_(PadState,left_x)), /* R_T5 settled, store left_xy */
store_half( R_T4, R_PadState, O_(PadState,right_x)),
add_ui( R_T5, R_0, 0x53), /* R_T5 = id value (clobbers left_xy, already stored) */
store_byte( R_T5, R_PadState, O_(PadState,id)),
jump_rel(atom_offset(analog_stick, snap_end)),
mac_yield_load(),
atom_label(try_analog_pad) /* === Case 5-6: AnalogPad (id & 0xF0 == 0x70) */
and_i( R_T4, R_RawId, 0xF0),
add_ui( R_T5, R_0, 0x70),
branch_ne(R_T4, R_T5, atom_offset(try_analog_pad, try_unsupported)),
/* BD-slot: pre-compute PadStatus_AnalogPad. Branch reads R_T4 in EX before this WB completes.
* If branch NOT taken (fall through to try_unsupported), R_T4 is overwritten by the unsupported body add_ui. */
atom_label(analog_pad) /* === AnalogPad body
* Same shape as AnalogStick with AnalogPad status. R_T5 holds left_xy (it's dead on this path). */
/* R_T4 = PadStatus_AnalogPad from try_analog_pad BD-slot. */
store_word( R_T4, R_PadState, O_(PadState,status)),
load_half_u(R_T4, R_PadRaw, 2 * S_(U1)), /* R_T4 = raw_buttons */
load_half_u(R_T5, R_PadRaw, 6 * S_(U1)), /* R_T5 = left_xy; fills R_T4's load-delay slot */
nor_u( R_T4, R_T4, R_0), /* R_T4 = ~raw_buttons */
store_half( R_T4, R_PadState, O_(PadState,buttons)),
load_half_u(R_T4, R_PadRaw, 4 * S_(U1)), /* R_T4 = right_xy; fills R_T5's load-delay slot */
store_half( R_T5, R_PadState, O_(PadState,left_x)), /* R_T5 settled, store left_xy */
store_half( R_T4, R_PadState, O_(PadState,right_x)),
store_byte( R_RawId, R_PadState, O_(PadState,id)),
jump_rel(atom_offset(analog_pad, snap_end)),
mac_yield_load(),
atom_label(try_unsupported) /* === Case 7: Unsupported — fall through from the AnalogPad range-check miss. */
add_ui( R_T4, R_0, PadStatus_Unsupported),
store_word(R_T4, R_PadState, O_(PadState,status)),
store_half(R_0, R_PadState, O_(PadState,buttons)),
/* axes = 0x80808080 (centered) — single sw writes the 4-byte axes block at offset 8 (left_x, left_y, right_x, right_y). */
load_upper_i(R_T4, 0x8080), or_i_self(R_T4, 0x8080),
store_word( R_T4, R_PadState, O_(PadState,left_x)),
add_ui( R_T4, R_0, 0xFF), /* 0xFF sentinel: "unknown id" */
store_byte( R_T4, R_PadState, O_(PadState,id)),
/* Fall through to snap_end. */
atom_label(no_jump_fallthrough)
mac_yield_load(),
atom_label(snap_end)
/* NOT mac_yield() — R_AtomJmp was already loaded in the BD-slot of the case-exit branch. */
mac_yield_tail(),
};
/* ----- pad_apply_input -----
* Reads pad[0].buttons + pad[0].left_x;
* Applies the input-semantics deltas to cube_rot.y + floor_rot.y:
* - D-pad Left: cube_rot.y += 30, floor_rot.y += 5
* - D-pad Right: cube_rot.y -= 30, floor_rot.y -= 5
* - Analog stick X (dead zone 0x70..0x90):
* cube delta = (0x80 - left_x) >> 2 (range approx -32..+32)
* floor delta = (0x80 - left_x) >> 5 (range approx -4..+4)
* - D-pad + analog deltas add when used together.
*
* Convention:
* pad_state = 0 means no buttons active.
* The fail-safe zero-button value flows through unchanged, so a disconnected/fresh pad produces no rotation.
* The branch_le_zero pattern below matches the existing pad_input_demo convention (atom body lines 248/257).
*
* Signed-delta trick:
* load_byte_u zero-extends left_x to 32 bits; sub_u from 0x80 wraps to a SIGNED two's-complement value in the negative range;
* shift_aright (sra) then correctly sign-extends the shift for both positive (left_x < 0x80) and negative (left_x > 0x80) cases.
* Digital pads publish left_x = 0x80 → delta = 0 → no rotation, so the analog step is naturally a no-op for digital controllers.
*/
typedef Struct_(Binds_PadApplyInput) {
PadState* state;
V3_S2* cube_rot;
V3_S2* floor_rot;
};
enum {
R_PadStateT5 = R_T5 atom_reg,
R_CubeRot = R_T1 atom_reg,
R_FloorRot = R_T2 atom_reg,
};
internal MipsAtom_(pad_apply_input) atom_info(atom_bind(Binds_PadApplyInput)
, atom_reads(R_T0, R_CubeRot, R_FloorRot, R_T3, R_T4, R_PadStateT5, R_TapePtr)
, atom_writes( R_CubeRot, R_FloorRot)
) {
/* Pop Binds from tape (state, cube_rot, floor_rot) */
load_word(R_PadStateT5, R_TapePtr, O_(Binds_PadApplyInput,state)),
load_word(R_CubeRot, R_TapePtr, O_(Binds_PadApplyInput,cube_rot)),
load_word(R_FloorRot, R_TapePtr, O_(Binds_PadApplyInput,floor_rot)),
add_ui_self( R_TapePtr, S_(Binds_PadApplyInput)),
/* Load pad[0].buttons into R_T0. */
load_word(R_T0, R_PadStateT5, O_(PadState,buttons)), nop,
// Note(Ed): Potential op with delay slot?
/* D-pad Left: cube_rot.y += 30, floor_rot.y += 5. */
and_i(R_T3, R_T0, pad0_(Pad_Left)), branch_le_zero(R_T3, atom_offset(dpad_left, exit_dpad_left)),
load_half( R_T4, R_CubeRot, O_(V3_S2,y)), /* BD-slot */
load_half( R_T3, R_FloorRot, O_(V3_S2,y)),
add_si( R_T4, R_T4, 30),
add_si( R_T3, R_T3, 5),
store_half(R_T4, R_CubeRot, O_(V3_S2,y)),
store_half(R_T3, R_FloorRot, O_(V3_S2,y)),
atom_label(exit_dpad_left)
/* D-pad Right: cube_rot.y -= 30, floor_rot.y -= 5. */
and_i(R_T3, R_T0, pad0_(Pad_Right)), branch_le_zero(R_T3, atom_offset(dpad_right, exit_dpad_right)),
load_half( R_T4, R_CubeRot, O_(V3_S2,y)), /* BD-slot */
load_half( R_T3, R_FloorRot, O_(V3_S2,y)),
add_si( R_T4, R_T4, -30),
add_si( R_T3, R_T3, -5),
store_half(R_T4, R_CubeRot, O_(V3_S2,y)),
store_half(R_T3, R_FloorRot, O_(V3_S2,y)),
atom_label(exit_dpad_right)
/* Analog left-stick X: dead zone 0x70..0x90.
* Cube delta = (0x80 - left_x) >> 2; floor delta = (0x80 - left_x) >> 5. */
load_byte_u(R_T3, R_PadStateT5, O_(PadState,left_x)),
/* Dead-zone check: skip analog if left_x in [0x70, 0x90] inclusive. Outside dead zone on LOW side: left_x < 0x70 (strictly).
* set_lt_u(R_T4, R_T3, R_T4=0x70) → R_T4 = (left_x < 0x70) ? 1 : 0. */
add_ui(R_T4, R_0, 0x70), set_lt_u(R_T4, R_T3, R_T4), branch_ne(R_T4, R_0, atom_offset(dead_zone_low_check, dead_low_active)),
add_ui(R_T4, R_0, 0x80), /* BD-slot: pre-load 0x80 for dead_low_active */
atom_label(dead_check_upper)
/* left_x >= 0x70 → check upper bound. */
load_byte_u(R_T3, R_PadStateT5, O_(PadState,left_x)), /* reload */
add_ui( R_T4, R_0, 0x90),
/* R_T4 = (0x90 < left_x) ? 1 : 0 → (left_x > 0x90) ? 1 : 0 */
set_lt_u(R_T4, R_T4, R_T3), branch_ne(R_T4, R_0, atom_offset(dead_zone_high_check, dead_high_active)),
add_ui( R_T4, R_0, 0x80), /* BD-slot: pre-load 0x80 for dead_high_active */
jump_rel(atom_offset(dead_zone_skip, exit_stick)),
mac_yield_load(),
atom_label(dead_low_active)
/* R_T3 = left_x (from line 632 lbu; not clobbered between dead_zone_low_check branch + its BD-slot `add_ui R_T4, 0x80`).
* The earlier `load_byte_u(R_T3, ...)` reload was redundant and introduced a load-use hazard on the next `sub_u`.
* R_T4 = 0x80 from the BD-slot of `dead_zone_low_check`'s branch_ne. */
sub_u( R_T3, R_T4, R_T3), /* R_T3 = 0x80 - left_x */
/* delta = 0x80 - left_x (positive). */
/* R_T4 = cube_delta */
shift_aright(R_T4, R_T3, 2),
load_half( R_T0, R_CubeRot, O_(V3_S2,y)),
nop,
add_u( R_T0, R_T0, R_T4),
store_half( R_T0, R_CubeRot, O_(V3_S2,y)),
/* R_T4 = floor_delta — moved into the load-delay slot of the floor load below (fills the 1-instruction gap;
* doesn't read R_T0; R_T4 settles by the subsequent add_u). */
load_half( R_T0, R_FloorRot, O_(V3_S2,y)),
shift_aright(R_T4, R_T3, 5),
add_u( R_T0, R_T0, R_T4),
store_half( R_T0, R_FloorRot, O_(V3_S2,y)),
jump_rel(atom_offset(end_low, exit_stick)),
mac_yield_load(),
atom_label(dead_high_active)
/* R_T3 = left_x (from line 641 lbu in dead_check_upper; not clobbered between dead_zone_high_check branch + its BD-slot `add_ui R_T4, 0x80`).
* The earlier `load_byte_u(R_T3, ...)` reload was redundant and introduced a load-use hazard on the next `sub_u`.
* R_T4 = 0x80 from the BD-slot of `dead_zone_high_check`'s branch_ne. */
sub_u( R_T3, R_T4, R_T3),
/* delta = 0x80 - left_x (signed negative). */
shift_aright(R_T4, R_T3, 2), /* R_T4 = cube_delta (signed) */
load_half( R_T0, R_CubeRot, O_(V3_S2,y)),
nop,
add_u( R_T0, R_T0, R_T4),
store_half( R_T0, R_CubeRot, O_(V3_S2,y)),
/* R_T4 = floor_delta (signed) — moved into the load-delay slot of the floor load below. */
load_half( R_T0, R_FloorRot, O_(V3_S2,y)),
shift_aright(R_T4, R_T3, 5),
add_u( R_T0, R_T0, R_T4),
store_half( R_T0, R_FloorRot, O_(V3_S2,y)),
atom_label(no_jump_fallthrough)
mac_yield_load(),
atom_label(exit_stick)
/* NOT mac_yield() — R_AtomJmp was already loaded in the BD-slot of the dead-zone/exit branch. */
mac_yield_tail(),
};
#pragma endregion Baked Atoms
+441
View File
@@ -0,0 +1,441 @@
#pragma region Vendors
#include <stdio.h>
#include <stdlib.h>
#include <assert.h>
// #include "libgpu.h"
// #include "libetc.h"
// #include "libgte.h"
#pragma endregion Vendors
#pragma region Duffle Headers
# include "duffle/gen/macs.h"
# include "duffle/gen/offsets.h"
#include "duffle/word_count.metadata.h"
#include "duffle/dsl.h"
#include "duffle/memory.h"
#include "duffle/math.h"
#include "duffle/gcc_asm.h"
#include "duffle/mips.h"
#include "duffle/gp.h"
#include "duffle/gte.h"
#include "duffle/pad.h"
#include "duffle/dsl.atom.h"
#include "duffle/lottes_tape.h"
#include "duffle/psyq.h"
#pragma endregion Duffle Headers
#pragma region Duffle TUs
#include "duffle/math.atom.c"
#include "duffle/mips.atom.c"
#include "duffle/gte.atom.c"
#include "duffle/gp.atom.c"
#include "duffle/psyq.atom.c"
#pragma endregion Duffle TUs
#pragma region Joypade Headers
# include "gen/macs.h"
# include "gen/offsets.h"
#include "hello_joypad.h"
#pragma region Joypad Headers
#pragma region Hello Joypad TUs
#include "hello_joypad.atom.c"
#pragma endregion Hello Joypad TUs
enum {
Scratchpad_Len = 1024,
MemTape_Len = 512,
};
typedef Struct_(SMemory) {
PrimitiveArena primitives;
A2_OrderingTable_Buffer ordering_tbl;
DoubleBuffer screen_buf;
S4 active_buf_id;
U4 MemTape[MemTape_Len];
M3_S2 tform_world;
Ent_Cube cube;
Ent_Floor floor;
PadBiosRaw pad_raw[2];
PadState pad[2];
U4_V scratchpad; // d-cache
};
global SMemory smem;
extern SMemory smem;
I_ B1* prim__alloc(U4 type_width, Str8 type_name) {
gknown PrimitiveArena* pa = & smem.primitives;
gknown B1* buf = (B1*) r_(smem.primitives.buf)[smem.active_buf_id];
assert(pa->used + type_width < PrimitiveBuff_Len);
B1* next = buf + pa->used;
pa->used += type_width;
return next;
}
#define prim_alloc(type) (type*)prim__alloc(S_(type), slit( stringify(type)))
/* Uses ONE 8-byte frame allocated via the compiler's standard prologue.
* The 4 wasted-arg words for B(12h) InitPAD2 live at [SP+0..15] but are not explicitly allocated.
* The compiler handles the MIPS O32 "wasted stack" convention for us by treating the B-call as a 4-arg call.
*
* The buffer pointers are passed as arguments so the compiler keeps them in callee-saved registers;
* The B(12h) asm volatile block does NOT clobber those registers (it clobbers only the volatile GPRs + the B-table arg registers explicitly).
* The C-level writes after the call re-load the pointers from their callee-saved homes.
*
* The clobber list for both B-calls names the full BIOS destroy set documented in kernelbios.md:167-174 (R1..R15, R24..R25, R31, HI/LO).
* The kernel-ABI "volatile GPRs" subset is clb_system; the rest of the destroy set is enumerated explicitly here. */
NI_ void pad_bios_init_start(PadBiosRaw* raw0, PadBiosRaw* raw1)
{
/* Pin raw0 + raw1 to $a0 + $a1 via rgcc; the B(12h) call uses these directly.
* The `(void)` casts mark them as unread after the call so the compiler doesn't need to move them back. */
register PadBiosRaw* p0 rgcc(R_A0) = raw0;
register PadBiosRaw* p1 rgcc(R_A1) = raw1;
(void)p0; (void)p1;
// TODO(Ed): Properly annotate the raw values in the inline asm instructions.
// Use enums.
/* B(12h) InitPAD2(raw0, 0x22, raw1, 0x22)
* $a0 = raw0 (rgcc-bound; survives the sequence below)
* $a1 = raw1 (preserved into $a2 before $a1 is overwritten)
* $a2 = raw1 (moved from $a1; survives $a1's overwrite)
* $a3 = 0x22 (immediate)
* $t1 = 0x12 (function number)
* $t2 = 0xB0 (BIOS B-table address) */
asm volatile(
asm_words(
or_u( rarg_2, rarg_1, rdiscard), /* $a2 = $a1 = raw1 */
add_ui( rarg_1, rdiscard, 0x22), /* $a1 = 0x22 */
add_ui( rarg_3, rdiscard, 0x22), /* $a3 = 0x22 */
add_ui( rtmp_1, rdiscard, 0x12), /* $t1 = 0x12 */
add_ui( rtmp_2, rdiscard, 0xB0), /* $t2 = 0xB0 */
call_reg(rtmp_2), /* jalr $t2, $ra */
nop /* BD slot */
)
asm_rpins, r_use(p0), r_use(p1)
asm_clobber:
rlit(R_AT),
rlit(R_V0), rlit(R_V1),
rlit(R_T0), rlit(R_T1), rlit(R_T2), rlit(R_T3), rlit(R_T4),
rlit(R_T5), rlit(R_T6), rlit(R_T7), rlit(R_T8), rlit(R_T9),
rlit(R_RA),
clb_mem_drain
);
/* The C-level writes re-load the pointers via the parameter names and write 0xFF to each
* buffer's status byte to mark the initial-state hazard documented in kernelbios.md:1621-1624. */
u1_v(raw0)[0] = 0xFF;
u1_v(raw1)[0] = 0xFF;
/* B(13h) StartPAD2() — no args. The BIOS preserves $sp. */
asm volatile(
asm_words(
add_ui( rtmp_1, rdiscard, 0x13), /* $t1 = 0x13 */
add_ui( rtmp_2, rdiscard, 0xB0), /* $t2 = 0xB0 (re-load) */
call_reg(rtmp_2), /* jalr $t2, $ra */
nop /* BD slot */
)
asm_clobber:
rlit(R_AT),
rlit(R_V0), rlit(R_V1),
rlit(R_T0), rlit(R_T1), rlit(R_T2), rlit(R_T3), rlit(R_T4),
rlit(R_T5), rlit(R_T6), rlit(R_T7), rlit(R_T8), rlit(R_T9),
rlit(R_RA),
clb_mem_drain
);
}
GCC_OPTIMIZATION_DISABLE
void update(PrimitiveArena* pa, U4* ordering_buf)
{
TapeBuilder tb = tb_make(slice_ut_arr(smem.MemTape));
if (0) // Pad Input (dead — kept for the source-as-written record; references the deleted `pad_state` field)
{
(void)Pad_Left; (void)Pad_Right; /* suppress unused-token warnings */
if (false) {
smem.cube.rot.y += 30;
smem.floor.rot.y += 5;
}
if (false) {
smem.cube.rot.y -= 30;
smem.floor.rot.y -= 5;
}
}
if (1) // Pad Input (Tape version)
{
tb.used = 0; tb_scope_run(& tb) {
/* BIOS-owned polling: per-frame snapshot of both ports. */
tb_emit_(pad_bios_snapshot);
tb_data_(raw, & smem.pad_raw[0]);
tb_data_(state, & smem.pad[0]);
tb_emit_(pad_bios_snapshot);
tb_data_(raw, & smem.pad_raw[1]);
tb_data_(state, & smem.pad[1]);
/* Per-frame rotation apply: consume pad[0].buttons + pad[0].left_x */
tb_emit_(pad_apply_input);
tb_data_(state, & smem.pad[0]);
tb_data_(cube_rot, & smem.cube.rot);
tb_data_(floor_rot, & smem.floor.rot);
}
}
orderingtbl_clear_reverse(ordering_buf, OrderingTbl_Len);
// Update the position based on acceleration and velocity
gknown V3_S4_R pos = & smem.cube.pos;
gknown V3_S4_R vel = & smem.cube.vel;
gknown V3_S4_R acc = & smem.cube.accel;
add_v3s4(vel, acc[0]);
add_v3s4_fp(pos, vel[0]);
// vel->x += acc->x;
// vel->y += acc->y;
// vel->z += acc->z;
// pos->x += vel->x;
// pos->y += vel->y;
// pos->z += vel->z;
if (pos->y + 150 > smem.floor.pos.y) vel->y *= -1;
// Prep
S4 nclip = 0;
S4 orderingtbl_z = 0;
A2_S2 p; //???
S4 flag; //????
// Draw Cube
if (0)
{
m3s2_rotation (& smem.cube.rot, & smem.tform_world);
m3s2_translation(& smem.tform_world, & smem.cube.pos);
m3s2_scale (& smem.tform_world, & smem.cube.scale);
// gte_matrix_set_rotation (& smem.tform_world);
gte_matrix_set_translation(& smem.tform_world);
for (U4 face_id = 0; face_id < Cube_num_faces; face_id += 1)
{
Poly_G4* quad = prim_alloc(Poly_G4); set_poly_g4(quad);
quad->c0 = rgb8(255, 0, 255);
quad->c1 = rgb8(255, 255, 0);
quad->c2 = rgb8( 0, 255, 255);
quad->c3 = rgb8( 0, 255, 0);
V4_S2* face = & smem.cube.faces[face_id];
V3_S2* p0 = & smem.cube.verts[face->x];
V3_S2* p1 = & smem.cube.verts[face->y];
V3_S2* p2 = & smem.cube.verts[face->z];
V3_S2* p3 = & smem.cube.verts[face->w];
nclip = rtp_avg_nclip_a4_v3s2(
p0, p1, p2, p3,
& quad->p0, & quad->p1, & quad->p2, & quad->p3,
& p, & orderingtbl_z, & flag
);
if (nclip <= 0) {
continue;
}
if ((orderingtbl_z > 0) && (orderingtbl_z < OrderingTbl_Len)) {
orderingtbl_add_primitive(ordering_buf[orderingtbl_z], quad);
}
}
// smem.cube.rot.x += 6;
// smem.cube.rot.y += 8;
// smem.cube.rot.z += 12;
smem.cube.rot.y += 30;
}
// Draw cube (tape method) - two triangles per face
if (1)
{
m3s2_rotation (& smem.cube.rot, & smem.tform_world);
m3s2_translation(& smem.tform_world, & smem.cube.pos);
m3s2_scale (& smem.tform_world, & smem.cube.scale);
gte_matrix_set_rotation (& smem.tform_world);
gte_matrix_set_translation(& smem.tform_world);
U4 prim_base = u4_(pa->buf[smem.active_buf_id]);
U4 prim_cursor = prim_base + pa->used;
tb.used = 0; tb_scope(& tb) {
tb_emit(& tb, rbind_cube_g4_face);
tb_data(& tb, prim_cursor);
tb_data(& tb, u4_(smem.cube.faces));
tb_data(& tb, u4_(smem.cube.verts));
tb_data(& tb, u4_(ordering_buf));
for (U4 i = 0; i < Cube_num_faces; i++) {
// Two triangles per quad face: (x,y,z) and (x,z,w)
tb_emit(& tb, cube_g4_face);
}
tb_emit(& tb, sync_primitive_arena);
tb_data(& tb, u4_(& pa->used));
tb_data(& tb, prim_base);
}
tape_run(tb_slice(tb));
// smem.cube.rot.y += 30;
}
// Draw Floor
if (0)
{
m3s2_rotation (& smem.floor.rot, & smem.tform_world);
m3s2_translation(& smem.tform_world, & smem.floor.pos);
m3s2_scale (& smem.tform_world, & smem.floor.scale);
gte_matrix_set_rotation (& smem.tform_world);
gte_matrix_set_translation(& smem.tform_world);
for (U4 face_id = 0; face_id < Floor_num_faces; face_id += 1)
{
Poly_F3* tri = prim_alloc(Poly_F3); set_poly_f3(tri);
tri->color = rgb8(255, 255, 255);
V3_S2* face = & smem.floor.faces[face_id];
register V3_S2* p0 rgcc(R_T4) = & smem.floor.verts[face->x];
register V3_S2* p1 rgcc(R_T5) = & smem.floor.verts[face->y];
register V3_S2* p2 rgcc(R_T6) = & smem.floor.verts[face->z];
gte_load_v0(p0, R_T4);
/*
asm volatile( ".word " "%0" ", %1" : :
"i"(((op_lwc2 & OPCODE_MASK) << OPCODE_SHIFT) | ((R_T4 & REG_MASK) << RS_SHIFT) | ((gte_in_v0_xy & REG_MASK) << RT_SHIFT) | (0 & IMM_MASK)),
"i"(((op_lwc2 & OPCODE_MASK) << OPCODE_SHIFT) | ((R_T4 & REG_MASK) << RS_SHIFT) | ((gte_in_v0_z & REG_MASK) << RT_SHIFT) | (GTE_Z_Offset & IMM_MASK)),
"r"(p0) :
"$2", "$8", "$9", "$31", "memory"
);
*/
gte_load_v1(p1, R_T5);
gte_load_v2(p2, R_T6);
gte_rtpt();
gte_nclip();
gte_stotz(& nclip);
// nclip = rtp_avg_nclip_a3_v3s2(p0, p1, p2
// , & tri->p0, & tri->p1, & tri->p2
// , & p, & orderingtbl_z, & flag
// );
// if (nclip <= 0) {
// continue;
// }
if (nclip > 0 ) {
gte_stsxy3(& tri->p0, & tri->p1, & tri->p2);
gte_avsz3();
gte_stotz(& orderingtbl_z);
if ((orderingtbl_z > 0) && (orderingtbl_z < OrderingTbl_Len)) {
orderingtbl_add_primitive(ordering_buf[orderingtbl_z], tri);
}
}
}
smem.floor.rot.y += 5;
}
// Draw floor tape method
if (1)
{
m3s2_rotation (& smem.floor.rot, & smem.tform_world);
m3s2_translation(& smem.tform_world, & smem.floor.pos);
m3s2_scale (& smem.tform_world, & smem.floor.scale);
U4 prim_base = u4_(pa->buf[smem.active_buf_id]);
U4 prim_cursor = prim_base + pa->used;
// TODO(Ed): We should do a bounds check beforehand to confirm pa can hold all tris?
// The tape atoms in-flight should not need to care.
// Prepare the tape. (Push protocol to tape)
tb.used = 0; tb_scope(& tb) {
tb_emit(& tb, set_gte_world);
tb_data(& tb, u4_(& smem.tform_world));
tb_emit(& tb, rbind_floor_f3_face);
// TODO(Ed): Just use a single context struct ref
tb_data(& tb, prim_cursor);
tb_data(& tb, u4_(smem.floor.faces));
tb_data(& tb, u4_(smem.floor.verts));
tb_data(& tb, u4_(ordering_buf));
for (U4 i = 0; i < Floor_num_faces; i++) {
tb_emit(& tb, floor_f3_face);
}
// After floor_f3_face iterations complete, the primitive arena's used counter needs updating.
tb_emit(& tb, sync_primitive_arena);
tb_data(& tb, u4_(& pa->used));
tb_data(& tb, prim_base);
}
tape_run(tb_slice(tb));// Fire off the tape.
// C-side state (pa->used) has already been updated by the tape!
// smem.floor.rot.y += 5;
}
}
GCC_OPTIMIZATION_ENABLE
void render(void) {
}
void gp_display_frame(DoubleBuffer* screen_buf, S4* active_buf_id, U4* ordering_buf, PrimitiveArena* pa) {
draw_sync(0);
vsync(0);
displayenv_put(& r_(screen_buf->display)[active_buf_id[0] ]);
drawenv_put (& r_(screen_buf->draw) [active_buf_id[0] ]);
{
draw_orderingtbl(ordering_buf + OrderingTbl_Len - 1);
pa->used = 0;
}
active_buf_id[0] = ! active_buf_id[0]; // Swap current buffer
}
GCC_OPTIMIZATION_DISABLE
int main(void)
{
smem = (SMemory){0};
smem.scratchpad = C_(U4_V, 0x1F800000);
// smem.primitives.used = 0;
// smem.active_buf_id = 0;
/*Persistent Entity Setup*/{
ent_cube128_init(& smem.cube.verts, & smem.cube.faces); {
Ent_Cube* cube = & smem.cube;
cube->rot = v3s2(0, 0, 0);
cube->scale = v3s4_fp_one();
cube->accel = v3s4(0, 1, 0);
cube->pos = v3s4(0, -400, 1800);
}
ent_floor_init(& smem.floor.verts, & smem.floor.faces); {
Ent_Floor* floor = & smem.floor;
floor->rot = v3s2(0, 0, 0);
floor->pos = v3s4(0, 450, 1800);
floor->scale = v3s4_fp_one();
}
}
TapeBuilder tb = tb_make(slice_ut_arr(smem.MemTape)); {
reset_graph(0);
/* Direct BIOS: poll both ports during VBlank. */
pad_bios_init_start(& smem.pad_raw[0], & smem.pad_raw[1]);
/* Pinned registers for the GPU init atom. */
register U4* io_base_addr rgcc(R_IO_BaseAddr) = u4_r(IO_BASE_ADDR);
register DoubleBuffer* screen_buf rgcc(R_ScreenBuf) = & smem.screen_buf;
tb.used = 0; tb_scope_run(& tb) {
tb_emit(& tb, screen_env_init);
tb_emit(& tb, gp_screen_init);
}
}
while (1) {
gknown S4* active_buf_id = & smem.active_buf_id;
gknown U4* ordering_buf = r_(smem.ordering_tbl)[active_buf_id[0]];
gknown PrimitiveArena* pa = & smem.primitives;
update(pa, ordering_buf);
render();
gp_display_frame(& smem.screen_buf, active_buf_id, ordering_buf, pa);
};
return 0;
}
GCC_OPTIMIZATION_ENABLE
+102
View File
@@ -0,0 +1,102 @@
#ifdef INTELLISENSE_DIRECTIVES
# pragma once
# include "duffle/dsl.h"
# include "duffle/math.h"
# include "duffle/gp.h"
# include "duffle/pad.h"
#endif
enum {
// PrimitiveBuff_Len = 4096,
// OrderingTbl_Len = 2048,
PrimitiveBuff_Len = 131072,
OrderingTbl_Len = 8192,
};
enum {
ScreenRes_X = 320,
ScreenRes_Y = 240,
ScreenZ = 320,
ScreenRes_CenterX = (ScreenRes_X >> 1),
ScreenRes_CenterY = (ScreenRes_Y >> 1),
};
enum {
fp_one = (1 << 12),
};
#define v3s4_fp_one() v3s4(fp_one, fp_one, fp_one)
typedef U4 OrderingTable_Buffer[OrderingTbl_Len];
typedef Array_(OrderingTable_Buffer, 2);
typedef B1 PrimitiveBuffer[PrimitiveBuff_Len];
typedef Array_(PrimitiveBuffer, 2);
typedef Struct_(PrimitiveArena) {
A2_PrimitiveBuffer buf;
U4 used;
};
#define Cube_num_verts 8
typedef Array_(V3_S2, Cube_num_verts);
#define Cube_num_faces 6
typedef Array_(V4_S2, Cube_num_faces);
I_ void ent_cube128_init(A8_V3_S2* verts, A6_V4_S2* faces) {
LP_ A8_V3_S2 baked_verts = (A8_V3_S2) {
{ -128, -128, -128 },
{ 128, -128, -128 },
{ 128, -128, 128 },
{ -128, -128, 128 },
{ -128, 128, -128 },
{ 128, 128, -128 },
{ 128, 128, 128 },
{ -128, 128, 128 }
};
LP_ A6_V4_S2 baked_faces = (A6_V4_S2) {
{ 3, 2, 0, 1 },
{ 0, 1, 4, 5 },
{ 4, 5, 7, 6 },
{ 1, 2, 5, 6 },
{ 2, 3, 6, 7 },
{ 3, 0, 7, 4 },
};
mem_copy(u4_(verts), u4_(& baked_verts), S_(A8_V3_S2) );
mem_copy(u4_(faces), u4_(& baked_faces), S_(A6_V4_S2) );
return;
}
typedef Struct_(Ent_Cube) {
V3_S4 accel;
V3_S4 vel;
V3_S4 pos;
V3_S4 scale;
V3_S2 rot;
A8_V3_S2 verts;
A6_V4_S2 faces;
};
#define Floor_num_verts 4
typedef Array_(V3_S2, Floor_num_verts);
#define Floor_num_faces 2
typedef Array_(V3_S2, Floor_num_faces);
I_ void ent_floor_init(A4_V3_S2* verts, A2_V3_S2* faces) {
LP_ A4_V3_S2 baked_verts = (A4_V3_S2) {
{ -900, 0, -900 },
{ -900, 0, 900 },
{ 900, 0, -900 },
{ 900, 0, 900 },
};
LP_ A2_V3_S2 baked_faces = (A2_V3_S2) {
{ 0, 1, 2 },
{ 1, 3, 2 },
};
mem_copy(u4_(verts), u4_(& baked_verts), S_(A4_V3_S2));
mem_copy(u4_(faces), u4_(& baked_faces), S_(A2_V3_S2));
};
typedef Struct_(Ent_Floor) {
V3_S4 accel;
V3_S4 pos;
V3_S4 scale;
V3_S2 rot;
A4_V3_S2 verts;
A2_V3_S2 faces;
};
+659
View File
@@ -0,0 +1,659 @@
#if 0 /* ac_pad_sio_write_pad_state — superseded by pad_bios_snapshot */
/* ============================================================
* raw_sio_pad_poll_20260802 — superseded by bios_pad_buffer_snapshot_20260803.
* The doomed raw-SIO production atoms (ac_pad_sio_write_pad_state,
* pad_sio_init, pad_sio_step, pad_sio_diag_pin, pad_sio_diag_byte_exchange)
* reference symbols that were removed from code/duffle/pad.h during
* Phase 1. Each is wrapped in a narrow `#if 0` so the C compile skips
* the body while the source-as-written text stays in place for the
* Phase 5.1 deletion pass. The wrap is removed (and the bodies are
* deleted) by Phase 5.1 of this track.
* ============================================================ */
* Writes the per-port PadState in 5 instructions plus 4 store_word calls (status,
* buttons, left_x/y/right_x/right_y packed, attempt). The provisional decode publishes
* 0x0000FFFF buttons + centered axes on every path until response-byte decode lands.
*
* Args:
* status_val - the PadSioStatus enum value to publish
* state_ptr_reg - the PadState* base (R_PadState at the call site)
* scratch_reg - scratch register for the value being stored (e.g., R_T0)
*
* Emits 9 instructions (status/buttons/axes/attempt stores plus the
* two-instruction zero-extended buttons load).
*/
FI_ Slice_MipsCode ac_pad_sio_write_pad_state(U4 status_val, U4 state_ptr_reg, U4 scratch_reg)
MipsAtomComp_Proc_(ac_pad_sio_write_pad_state, {
add_ui(scratch_reg, R_0, status_val),
store_word(scratch_reg, state_ptr_reg, O_(PadState,status)),
/* FIX 2026-08-02: buttons = 0x0000FFFF = "no buttons pressed" in
* libetc convention. Build it with LUI + ORI so addiu does not
* sign-extend 0xFFFF to 0xFFFFFFFF. */
load_upper_i(scratch_reg, 0x0000),
or_i(scratch_reg, scratch_reg, 0xFFFF),
store_word(scratch_reg, state_ptr_reg, O_(PadState,buttons)),
add_ui(scratch_reg, R_0, 0x80808080),
store_word(scratch_reg, state_ptr_reg, O_(PadState,left_x)),
add_ui(scratch_reg, R_0, 0),
store_word(scratch_reg, state_ptr_reg, O_(PadState,attempt))
})
#endif /* end ac_pad_sio_write_pad_state wrap */
/* ----- pad_sio_init -----
* Boot-time SIO0 init. Caller pins R_T6 = sio_base_addr0.
* Issues SIO CTRL=0x0040 (reset), MODE=0x000D, BAUD=0x0088.
* (Phase 2 fills the body.)
*/
#if 0 /* pad_sio_init — superseded by pad_bios_init_start (Phase 1.3) */
internal MipsAtom_(pad_sio_init) atom_info(atom_phase(pad_init)
, atom_reads(R_T5, R_T6)
, atom_writes(R_T5, R_T6)
) {
/* FIX 2026-08-02: explicitly load the KSEG1 base into R_T6 at the top of
* the atom body. The rgcc(R_PadSioBase) binding in main() pins R_T6 = base
* when main() runs, but $12 is caller-saved per the O32 ABI — when tape_run
* is invoked, R_T6 is fair game. The atom body cannot rely on the value. */
load_upper_i(R_T6, pad_IO_KSEG1_BASE >> 16), /* R_T6 high 16 = 0xBF80 */
or_i(R_T6, R_T6, pad_IO_KSEG1_BASE & 0xFFFF), /* R_T6 = 0xBF800000 */
/* SIO CTRL = 0x0040 (reset) */
add_ui(R_T5, R_0, pad_SIO_CTRL_RESET),
store_half(R_T5, R_T6, pad_SIO_CTRL_OFFSET),
/* SIO MODE = 0x000D (MUL1, 8-bit, no parity, idle-high) */
add_ui(R_T5, R_0, pad_SIO_MODE_INIT),
store_half(R_T5, R_T6, pad_SIO_MODE_OFFSET),
/* SIO BAUD = 0x0088 (~250 kHz) */
add_ui(R_T5, R_0, pad_SIO_BAUD_INIT),
store_half(R_T5, R_T6, pad_SIO_BAUD_OFFSET),
mac_yield(),
};
#endif /* end pad_sio_init wrap */
/* ----- pad_sio_step -----
* Per-frame bounded raw-SIO transaction. Reads PadState pointers + SIO
* base addresses from Binds_PadSioStep; writes per-port status +
* buttons + axes into smem.pad[0..1].
* Body shape (per spec §"Transaction model (per port, per pad_sio_step)"):
* port 0: CTRL=CLEANUP → settle → CTRL=port-select → settle → exchange 5
* bytes (addr + 0x42 0x00 0x00 0x00) → decode → write PadState[0]
* → CTRL=CLEANUP.
* port 1: swap scratch regs (sio_base_addr1 → R_PadSioBase, state1 →
* R_PadState) → mirror port 0 sequence.
*
* Bounded-loop semantics: every countdown is wrapped in
* add_ui_self(R_T1, -1) + branch_ne(R_T1, R_0, ...)
* with a known maximum (pad_SIO_SETTLE_BEFORE_TX=1000, pad_SIO_SETTLE_AFTER_TX=2000,
* pad_SIO_WAIT_BUDGET=4096). The static-analysis pass currently reports
* has_loops = true; the follow-up metaprogram track that learns modeled-bounded
* loops is out of scope here (per spec §"Risks").
*
* Scratch register strategy:
* R_PadStatus = R_T4 — RESERVED for port-1 swap (holds state1)
* R_PadCountdown = R_T5 — RESERVED for port-1 swap (holds sio_base_addr1)
* R_T0 — byte-exchange value + STAT read (clobbered freely)
* R_T1 — countdown budget (clobbered freely)
* R_PadState = R_T7 — PadState* (preserved for PadState writes)
* R_PadSioBase = R_T6 — SIO base (preserved through the port)
*
* Response decode (Task 3.1 teaching scope):
* - status = PadSioStatus_Digital (hardcoded)
* - buttons = 0xFFFF (no buttons pressed in the provisional libetc
* convention; full response-byte decode is follow-up)
* - axes = 0x80808080 (centered: left_x=0x80, left_y=0x80,
* right_x=0x80, right_y=0x80)
* - attempt = 0
* - DualShock handshake (0x43 0x01 → 0x44 0x01 0x03 → 0x43 0x00) is
* follow-up scope; the hardcoded digital decode is a placeholder.
*
* Both ports raise /CS (CTRL = pad_SIO_CTRL_CLEANUP) before exit. Both ports
* treat response timeout as PadSioStatus_Disconnected per the spec §"Failure
* handling" + the canonical per-port timeout semantics.
*/
#if 0 /* pad_sio_step — superseded by pad_bios_snapshot (Phase 2.1) */
internal MipsAtom_(pad_sio_step) atom_info(atom_bind(Binds_PadSioStep)
, atom_reads(R_TapePtr, R_PadSioBase, R_PadState, R_PadStatus, R_PadCountdown)
, atom_writes(R_PadStatus, R_PadCountdown)
) {
/* FIX 2026-08-02: explicitly load KSEG1 base into R_PadSioBase (R_T6) at the
* top. The rgcc() binding in main() does NOT survive the tape_run call
* because R_T6 is caller-saved per the O32 ABI. The pad_sio_init atom
* (also in the per-frame tape) reloads R_T6 separately. */
load_upper_i(R_PadSioBase, pad_IO_KSEG1_BASE >> 16),
or_i(R_PadSioBase, R_PadSioBase, pad_IO_KSEG1_BASE & 0xFFFF),
/* Pop Binds from tape (in Binds_PadSioStep declaration order) */
load_word(R_PadState, R_TapePtr, O_(Binds_PadSioStep,state0)),
load_word(R_PadStatus, R_TapePtr, O_(Binds_PadSioStep,state1)), /* reserved for port-1 swap */
load_word(R_PadSioBase, R_TapePtr, O_(Binds_PadSioStep,sio_base_addr0)),
load_word(R_PadCountdown, R_TapePtr, O_(Binds_PadSioStep,sio_base_addr1)), /* reserved for port-1 swap */
add_ui_self(R_TapePtr, S_(Binds_PadSioStep)),
/* ============== PORT 0 TRANSACTION ============== */
/* Use R_T0 (byte value / STAT read) + R_T1 (countdown) as scratch.
* R_PadStatus (state1) + R_PadCountdown (sio_base_addr1) are preserved
* through the port-0 body and swapped into R_PadSioBase + R_PadState
* at atom_offset(port1_start, ...) below. */
/* 1. Cleanup: CTRL = 0x0010 (raise /CS, clear stale status) */
add_ui(R_T0, R_0, pad_SIO_CTRL_CLEANUP),
store_half(R_T0, R_PadSioBase, pad_SIO_CTRL_OFFSET),
/* Bounded by pad_SIO_SETTLE_BEFORE_TX = 1000 iterations. */
add_ui(R_T1, R_0, pad_SIO_SETTLE_BEFORE_TX),
atom_label(settle_pre_port0)
nop, /* BD slot */
add_ui_self(R_T1, -1),
branch_ne(R_T1, R_0, atom_offset(settle_pre_port0, settle_pre_port0)),
/* 2. Port-select: CTRL = 0x0003 (TX enable + DTR /CS) for port 0 */
add_ui(R_T0, R_0, pad_SIO_CTRL_TX_ENABLE),
or_i(R_T0, R_T0, pad_SIO_CTRL_DTR_CS), /* set /CS line low */
store_half(R_T0, R_PadSioBase, pad_SIO_CTRL_OFFSET),
/* Bounded by pad_SIO_SETTLE_AFTER_TX = 2000 iterations. */
add_ui(R_T1, R_0, pad_SIO_SETTLE_AFTER_TX),
atom_label(settle_post_port0)
nop,
add_ui_self(R_T1, -1),
branch_ne(R_T1, R_0, atom_offset(settle_post_port0, settle_post_port0)),
/* 3. Address byte (0x01) — send + RX-ready wait + read response + RX-drain confirmation */
add_ui(R_T0, R_0, pad_PROTO_ADDR),
store_byte(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ack0_port0)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_ne(R_T0, R_0, atom_offset(wait_ack0_port0, ack0_received_port0)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ack0_port0)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ack0_port0, wait_ack0_port0)),
/* RX timeout → mark disconnected; skip to port 1 */
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port0_from_ack0)
branch_equal(R_0, R_0, atom_offset(skip_port0_from_ack0, port1_start)),
atom_label(ack0_received_port0)
/* Read open-bus response byte 0 — discard per docs/psx-spx §controllersandmemorycards.md */
load_byte_u(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
/* Confirm RX FIFO drained before sending byte 1. Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ackrel0_port0)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_equal(R_T0, R_0, atom_offset(wait_ackrel0_port0, ack_released_port0)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ackrel0_port0)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ackrel0_port0, wait_ackrel0_port0)),
/* RX-drain timeout → disconnected; skip to port 1 */
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port0_from_ackrel0)
branch_equal(R_0, R_0, atom_offset(skip_port0_from_ackrel0, port1_start)),
atom_label(ack_released_port0)
/* === Byte 1 (port 0): send 0x42 (cmd read) + RX-ready wait + read response + RX-drain confirmation === */
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T0, R_0, pad_PROTO_CMD_READ),
store_byte(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ack1_port0)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_ne(R_T0, R_0, atom_offset(wait_ack1_port0, ack1_received_port0)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ack1_port0)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ack1_port0, wait_ack1_port0)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port0_from_ack1)
branch_equal(R_0, R_0, atom_offset(skip_port0_from_ack1, port1_start)),
atom_label(ack1_received_port0)
/* Read response ID byte — discarded for teaching scope (decode hardcoded). */
load_byte_u(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
/* RX FIFO drain wait. Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ackrel1_port0)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_equal(R_T0, R_0, atom_offset(wait_ackrel1_port0, ack_released1_port0)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ackrel1_port0)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ackrel1_port0, wait_ackrel1_port0)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port0_from_ackrel1)
branch_equal(R_0, R_0, atom_offset(skip_port0_from_ackrel1, port1_start)),
atom_label(ack_released1_port0)
/* === Byte 2 (port 0): send 0x00 + RX-ready wait + read response + RX-drain confirmation === */
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T0, R_0, 0x00),
store_byte(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ack2_port0)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_ne(R_T0, R_0, atom_offset(wait_ack2_port0, ack2_received_port0)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ack2_port0)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ack2_port0, wait_ack2_port0)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port0_from_ack2)
branch_equal(R_0, R_0, atom_offset(skip_port0_from_ack2, port1_start)),
atom_label(ack2_received_port0)
load_byte_u(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ackrel2_port0)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_equal(R_T0, R_0, atom_offset(wait_ackrel2_port0, ack_released2_port0)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ackrel2_port0)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ackrel2_port0, wait_ackrel2_port0)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port0_from_ackrel2)
branch_equal(R_0, R_0, atom_offset(skip_port0_from_ackrel2, port1_start)),
atom_label(ack_released2_port0)
/* === Byte 3 (port 0): send 0x00 + RX-ready wait + read response + RX-drain confirmation === */
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T0, R_0, 0x00),
store_byte(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ack3_port0)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_ne(R_T0, R_0, atom_offset(wait_ack3_port0, ack3_received_port0)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ack3_port0)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ack3_port0, wait_ack3_port0)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port0_from_ack3)
branch_equal(R_0, R_0, atom_offset(skip_port0_from_ack3, port1_start)),
atom_label(ack3_received_port0)
load_byte_u(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ackrel3_port0)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_equal(R_T0, R_0, atom_offset(wait_ackrel3_port0, ack_released3_port0)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ackrel3_port0)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ackrel3_port0, wait_ackrel3_port0)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port0_from_ackrel3)
branch_equal(R_0, R_0, atom_offset(skip_port0_from_ackrel3, port1_start)),
atom_label(ack_released3_port0)
/* === Byte 4 (FINAL, port 0): send 0x00 + RX-not-empty wait + read final byte === */
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T0, R_0, 0x00),
store_byte(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_rx4_port0)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_ne(R_T0, R_0, atom_offset(wait_rx4_port0, rx4_received_port0)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_rx4_port0)
branch_ne(R_T1, R_0, atom_offset(continue_wait_rx4_port0, wait_rx4_port0)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port0_from_rx4)
branch_equal(R_0, R_0, atom_offset(skip_port0_from_rx4, port1_start)),
atom_label(rx4_received_port0)
load_byte_u(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET), /* discard final byte */
/* === RESPONSE DECODE (hardcoded for teaching scope) ===
* Per the plan §"Phase 3 task 3.1" + spec §"Architecture":
* - Full decode (buttons/axes from response bytes) is follow-up scope.
* - Teaching scope: hardcode digital poll response.
* status = PadSioStatus_Digital
* buttons = 0x0000FFFF (no buttons pressed — placeholder)
* axes = 0x80808080 (left_x=0x80, left_y=0x80, right_x=0x80, right_y=0x80)
* attempt = 0
*/
atom_label(decode_port0)
mac_pad_sio_write_pad_state(PadSioStatus_Digital, R_PadState, R_T0),
/* /CS cleanup: raise /CS, clear stale status before exiting port 0. */
add_ui(R_T0, R_0, pad_SIO_CTRL_CLEANUP),
store_half(R_T0, R_PadSioBase, pad_SIO_CTRL_OFFSET),
/* ============== PORT 1 SETUP ============== */
/* Swap: R_PadCountdown holds sio_base_addr1; R_PadStatus holds state1. */
atom_label(port1_start)
add_u(R_PadSioBase, R_0, R_PadCountdown), /* sio_base_addr1 → R_PadSioBase */
add_u(R_PadState, R_0, R_PadStatus), /* state1 → R_PadState */
/* ============== PORT 1 TRANSACTION (mirror of port 0) ============== */
/* R_PadStatus + R_PadCountdown are no longer reserved (port 1 is the
* last transaction); we still use R_T0/R_T1 as scratch to match port 0. */
/* 1. Cleanup: CTRL = 0x0010 (raise /CS, clear stale status) */
add_ui(R_T0, R_0, pad_SIO_CTRL_CLEANUP),
store_half(R_T0, R_PadSioBase, pad_SIO_CTRL_OFFSET),
/* Bounded by pad_SIO_SETTLE_BEFORE_TX = 1000 iterations. */
add_ui(R_T1, R_0, pad_SIO_SETTLE_BEFORE_TX),
atom_label(settle_pre_port1)
nop,
add_ui_self(R_T1, -1),
branch_ne(R_T1, R_0, atom_offset(settle_pre_port1, settle_pre_port1)),
/* 2. Port-select: CTRL = 0x0003 | (1 << 13) (port 1 select) */
add_ui(R_T0, R_0, pad_SIO_CTRL_TX_ENABLE),
or_i(R_T0, R_T0, pad_SIO_CTRL_DTR_CS),
or_i(R_T0, R_T0, 1 << 13), /* port 1 select bit (CTRL bit 13 = port select) */
store_half(R_T0, R_PadSioBase, pad_SIO_CTRL_OFFSET),
/* Bounded by pad_SIO_SETTLE_AFTER_TX = 2000 iterations. */
add_ui(R_T1, R_0, pad_SIO_SETTLE_AFTER_TX),
atom_label(settle_post_port1)
nop,
add_ui_self(R_T1, -1),
branch_ne(R_T1, R_0, atom_offset(settle_post_port1, settle_post_port1)),
/* 3. Address byte (0x01) — send + RX-ready wait + read response + RX-drain confirmation */
add_ui(R_T0, R_0, pad_PROTO_ADDR),
store_byte(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ack0_port1)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_ne(R_T0, R_0, atom_offset(wait_ack0_port1, ack0_received_port1)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ack0_port1)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ack0_port1, wait_ack0_port1)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port1_from_ack0)
branch_equal(R_0, R_0, atom_offset(skip_port1_from_ack0, end_atom)),
atom_label(ack0_received_port1)
load_byte_u(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ackrel0_port1)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_equal(R_T0, R_0, atom_offset(wait_ackrel0_port1, ack_released_port1)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ackrel0_port1)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ackrel0_port1, wait_ackrel0_port1)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port1_from_ackrel0)
branch_equal(R_0, R_0, atom_offset(skip_port1_from_ackrel0, end_atom)),
atom_label(ack_released_port1)
/* === Byte 1 (port 1): send 0x42 (cmd read) + RX-ready wait + read response + RX-drain confirmation === */
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T0, R_0, pad_PROTO_CMD_READ),
store_byte(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ack1_port1)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_ne(R_T0, R_0, atom_offset(wait_ack1_port1, ack1_received_port1)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ack1_port1)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ack1_port1, wait_ack1_port1)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port1_from_ack1)
branch_equal(R_0, R_0, atom_offset(skip_port1_from_ack1, end_atom)),
atom_label(ack1_received_port1)
load_byte_u(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ackrel1_port1)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_equal(R_T0, R_0, atom_offset(wait_ackrel1_port1, ack_released1_port1)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ackrel1_port1)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ackrel1_port1, wait_ackrel1_port1)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port1_from_ackrel1)
branch_equal(R_0, R_0, atom_offset(skip_port1_from_ackrel1, end_atom)),
atom_label(ack_released1_port1)
/* === Byte 2 (port 1): send 0x00 + RX-ready wait + read response + RX-drain confirmation === */
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T0, R_0, 0x00),
store_byte(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ack2_port1)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_ne(R_T0, R_0, atom_offset(wait_ack2_port1, ack2_received_port1)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ack2_port1)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ack2_port1, wait_ack2_port1)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port1_from_ack2)
branch_equal(R_0, R_0, atom_offset(skip_port1_from_ack2, end_atom)),
atom_label(ack2_received_port1)
load_byte_u(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ackrel2_port1)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_equal(R_T0, R_0, atom_offset(wait_ackrel2_port1, ack_released2_port1)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ackrel2_port1)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ackrel2_port1, wait_ackrel2_port1)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port1_from_ackrel2)
branch_equal(R_0, R_0, atom_offset(skip_port1_from_ackrel2, end_atom)),
atom_label(ack_released2_port1)
/* === Byte 3 (port 1): send 0x00 + RX-ready wait + read response + RX-drain confirmation === */
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T0, R_0, 0x00),
store_byte(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ack3_port1)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_ne(R_T0, R_0, atom_offset(wait_ack3_port1, ack3_received_port1)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ack3_port1)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ack3_port1, wait_ack3_port1)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port1_from_ack3)
branch_equal(R_0, R_0, atom_offset(skip_port1_from_ack3, end_atom)),
atom_label(ack3_received_port1)
load_byte_u(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_ackrel3_port1)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_equal(R_T0, R_0, atom_offset(wait_ackrel3_port1, ack_released3_port1)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_ackrel3_port1)
branch_ne(R_T1, R_0, atom_offset(continue_wait_ackrel3_port1, wait_ackrel3_port1)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port1_from_ackrel3)
branch_equal(R_0, R_0, atom_offset(skip_port1_from_ackrel3, end_atom)),
atom_label(ack_released3_port1)
/* === Byte 4 (FINAL, port 1): send 0x00 + RX-not-empty wait + read final byte === */
/* Bounded by pad_SIO_WAIT_BUDGET = 4096 iterations. */
add_ui(R_T0, R_0, 0x00),
store_byte(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(wait_rx4_port1)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_ne(R_T0, R_0, atom_offset(wait_rx4_port1, rx4_received_port1)),
add_ui_self(R_T1, -1),
atom_label(continue_wait_rx4_port1)
branch_ne(R_T1, R_0, atom_offset(continue_wait_rx4_port1, wait_rx4_port1)),
mac_pad_sio_write_pad_state(PadSioStatus_Disconnected, R_PadState, R_T0),
atom_label(skip_port1_from_rx4)
branch_equal(R_0, R_0, atom_offset(skip_port1_from_rx4, end_atom)),
atom_label(rx4_received_port1)
load_byte_u(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET), /* discard final byte */
/* === RESPONSE DECODE (port 1) === */
atom_label(decode_port1)
mac_pad_sio_write_pad_state(PadSioStatus_Digital, R_PadState, R_T0),
/* /CS cleanup: raise /CS, clear stale status before exiting port 1. */
add_ui(R_T0, R_0, pad_SIO_CTRL_CLEANUP),
store_half(R_T0, R_PadSioBase, pad_SIO_CTRL_OFFSET),
atom_label(end_atom)
mac_yield(),
};
#endif /* end pad_sio_step wrap */
/* ----- pad_sio_diag_pin -----
* Per-frame diagnostic counter. The caller binds R_DiagPinScratch to
* scratch_for_atom_diag_pin for temporary gdb verification.
*/
#if 0 /* pad_sio_diag_pin — superseded (raw-SIO phase removed) */
internal MipsAtom_(pad_sio_diag_pin) atom_info(atom_phase(pad_init)
, atom_reads(R_T0, R_T1, R_DiagPinScratch)
, atom_writes(R_T0, R_T1, R_DiagPinScratch)
) {
/* FIX 2026-08-02: explicitly reload R_DiagPinScratch (R_T3 = $t3). Caller-saved
* per O32 ABI; the rgcc binding in main() does not survive tape_run. */
load_upper_i(R_DiagPinScratch, 0x8001),
or_i(R_DiagPinScratch, R_DiagPinScratch, 0xC800),
/* High half = 0xD1A6; low half increments once per atom invocation. */
load_word(R_T1, R_DiagPinScratch, 0),
nop,
add_ui(R_T1, R_T1, 1),
and_i(R_T0, R_T1, 0xFFFF),
load_upper_i(R_T1, 0xD1A6),
or_i(R_T1, R_T1, 0),
or_u(R_T1, R_T1, R_T0),
store_word(R_T1, R_DiagPinScratch, 0),
mac_yield(),
};
#endif /* end pad_sio_diag_pin wrap */
/* ----- pad_sio_diag_byte_exchange -----
* Temporary two-byte wire probe: sends 0x01 and 0x42, then stores the
* open-bus byte and response ID in scratch_for_atom_diag_pin.
*/
#if 0 /* pad_sio_diag_byte_exchange — superseded (raw-SIO phase removed) */
internal MipsAtom_(pad_sio_diag_byte_exchange) atom_info(atom_phase(pad_init)
, atom_reads(R_T0, R_T1, R_T2, R_PadSioBase, R_DiagPinScratch)
, atom_writes(R_T0, R_T1, R_T2, R_PadSioBase, R_DiagPinScratch)
) {
/* FIX 2026-08-02: explicitly reload R_DiagPinScratch (R_T3 = $t3). Caller-saved
* per O32 ABI; the rgcc binding in main() does not survive tape_run. */
load_upper_i(R_DiagPinScratch, 0x8001),
or_i(R_DiagPinScratch, R_DiagPinScratch, 0xC800),
/* FIX 2026-08-02: explicitly load KSEG1 base into R_PadSioBase (R_T6) at the
* top. The rgcc() binding in main() does NOT survive the tape_run call
* because R_T6 is caller-saved per the O32 ABI. */
load_upper_i(R_PadSioBase, pad_IO_KSEG1_BASE >> 16),
or_i(R_PadSioBase, R_PadSioBase, pad_IO_KSEG1_BASE & 0xFFFF),
add_ui(R_T0, R_0, pad_SIO_CTRL_CLEANUP),
store_half(R_T0, R_PadSioBase, pad_SIO_CTRL_OFFSET),
add_ui(R_T0, R_0, pad_SIO_CTRL_TX_ENABLE),
or_i(R_T0, R_T0, pad_SIO_CTRL_DTR_CS),
store_half(R_T0, R_PadSioBase, pad_SIO_CTRL_OFFSET),
add_ui(R_T0, R_0, pad_PROTO_ADDR),
store_byte(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(diag_wait_ack0)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_ne(R_T0, R_0, atom_offset(diag_wait_ack0, diag_ack0_done)),
add_ui_self(R_T1, -1),
branch_ne(R_T1, R_0, atom_offset(diag_wait_ack0, diag_wait_ack0)),
add_ui(R_T0, R_0, 0xDEADAC01),
store_word(R_T0, R_DiagPinScratch, 0),
branch_equal(R_0, R_0, atom_offset(diag_timeout_ack0, diag_timeout)),
atom_label(diag_ack0_done)
load_byte_u(R_T2, R_PadSioBase, pad_SIO_DATA_OFFSET),
add_ui(R_T0, R_0, pad_PROTO_CMD_READ),
store_byte(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
add_ui(R_T1, R_0, pad_SIO_WAIT_BUDGET),
atom_label(diag_wait_ack1)
load_half_u(R_T0, R_PadSioBase, pad_SIO_STAT_OFFSET),
nop,
and_i(R_T0, R_T0, pad_SIO_STAT_RX_NOT_EMPTY),
branch_ne(R_T0, R_0, atom_offset(diag_wait_ack1, diag_ack1_done)),
add_ui_self(R_T1, -1),
branch_ne(R_T1, R_0, atom_offset(diag_wait_ack1, diag_wait_ack1)),
add_ui(R_T0, R_0, 0xDEADAC02),
store_word(R_T0, R_DiagPinScratch, 0),
branch_equal(R_0, R_0, atom_offset(diag_timeout_ack1, diag_timeout)),
atom_label(diag_ack1_done)
load_byte_u(R_T0, R_PadSioBase, pad_SIO_DATA_OFFSET),
nop,
shift_lleft(R_T0, R_T0, 8),
or_u(R_T2, R_T2, R_T0),
store_word(R_T2, R_DiagPinScratch, 0),
atom_label(diag_success)
branch_equal(R_0, R_0, atom_offset(diag_success, diag_done)),
nop,
atom_label(diag_timeout_ack0)
add_ui(R_T0, R_0, 0xDEADAC01),
store_word(R_T0, R_DiagPinScratch, 0),
atom_label(diag_timeout_ack1)
add_ui(R_T0, R_0, 0xDEADAC02),
store_word(R_T0, R_DiagPinScratch, 0),
atom_label(diag_timeout)
add_ui(R_T0, R_0, 0xDEADACFF),
store_word(R_T0, R_DiagPinScratch, 0),
atom_label(diag_done)
add_ui(R_T0, R_0, pad_SIO_CTRL_CLEANUP),
store_half(R_T0, R_PadSioBase, pad_SIO_CTRL_OFFSET),
mac_yield(),
};
#endif /* end pad_sio_diag_byte_exchange wrap */
Binary file not shown.

After

Width:  |  Height:  |  Size: 220 KiB

+29 -7
View File
@@ -6,18 +6,33 @@ A rest from the usual.
## Dependencies
I will be programming from a Windows 11 machine:
![system_info](./docs/assets/system_info.png)
```ps1
# not really used yet for scripts (may never)
scoop install lua
```
I will be programming from a Windows 11 machine (may eventually try this on the Steam Deck...):
[armips](https://github.com/Kingcom/armips)
* Supports doing bare-metal assembly for the ps1
* `scoop install armips` or just clone and build..
* Was used early in the course. Now I just use an macro asm dsl in C11.
[luajit-2.1](https://github.com/LuaJIT/LuaJIT.git)
```
scoop install luajit
```
* Used for lua scripts
* Particularly, ps1_meta.lua which is a staged metaprogram pass for the custom C11 Assembly DSL used in this codebase.
[lpeg](https://github.com/roberto-ieru/LPeg.git)
* Lua is slow (even jitted) so this helps.
[lfs (LuaFileSystem)](https://github.com/lunarmodules/luafilesystem)
* Native directory enumeration + `mkdir` for the build scripts.
* Used by `passes/word_count_eval.lua :: scan_dir` (native walk vs. `dir /b /s` subprocess,
~2ms vs. ~56ms) and by `duffle.lua :: ensure_dir` + `to_absolute_path` (avoids
`cmd.exe mkdir` + `cd` shell spawns, ~50ms each).
[pscx-redux](https://github.com/grumpycoders/pcsx-redux/): A collection of tools, research, hardware design, and libraries aiming at development and reverse engineering on the PlayStation 1.
@@ -57,3 +72,10 @@ scoop install lua
![polys!](./docs/assets/pcsx-redux.main_2025-08-03_20-45-35.png)
![hello_psyq!](./docs/assets/pcsx-redux_2025-08-05_23-01-19.png)
![cube!](./docs/assets/pcsx-redux_2025-10-11_03-04-01.png)
![cube and floor!](./docs/assets/pcsx-redux_2026-07-10_22-47-02.png)
Win 11 machine:
![system_info](./docs/assets/system_info.png)
Still haven't gotten around to trying this on linux...
+30
View File
@@ -0,0 +1,30 @@
-- gte_debug.lua — defensive version + prints error context.
local ok, err = pcall(function()
print("[debug] PCSX exists:", PCSX ~= nil)
print("[debug] PCSX.WebServer exists:", PCSX and PCSX.WebServer ~= nil)
print("[debug] PCSX.WebServer.Handlers exists:", PCSX and PCSX.WebServer and PCSX.WebServer.Handlers ~= nil)
if not PCSX.WebServer then
print("[debug] creating PCSX.WebServer...")
PCSX.WebServer = {}
end
if not PCSX.WebServer.Handlers then
print("[debug] creating PCSX.WebServer.Handlers...")
PCSX.WebServer.Handlers = {}
end
print("[debug] type of Handlers:", type(PCSX.WebServer.Handlers))
PCSX.WebServer.Handlers.gte = function(req)
local r = PCSX.getRegisters()
local out = { "pc=0x" .. string.format("%x", r.pc) }
for i = 0, 31 do
out[#out + 1] = string.format("D[%d]=0x%08x C[%d]=0x%08x",
i, r.CP2D.r[i], i, r.CP2C.r[i])
end
return table.concat(out, "\n")
end
print("[debug] handler registered")
end)
if not ok then
print("[debug] ERROR: " .. tostring(err))
end
View File
+217
View File
@@ -0,0 +1,217 @@
--- audit_lua_nesting.lua — Walk Lua source files and flag any block nesting deeper than 5 levels.
---
--- Usage:
--- luajit scripts/audit_lua_nesting.lua scripts/duffle.lua scripts/ps1_meta.lua
--- luajit scripts/audit_lua_nesting.lua scripts/passes/
---
--- Output: for each file, a list of {line, depth} entries where depth > 5.
--- Returns exit code 1 if any violations found, 0 if clean.
---
--- **Implementation**: a hand-rolled depth tracker that counts:
--- - `do`, `function`, `if`, `for`, `while`, `repeat` -> depth +1
--- - `end`, `until` -> depth -1
--- - `else`, `elseif` -> depth unchanged
---
--- **Caveats**: doesn't fully handle string/comment state (will miscount braces inside multi-line strings or block comments).
--- For our metaprogram files (no embedded code generation), this is acceptable.
local M = {}
local BLOCK_OPEN = {
["do"] = true,
["function"] = true,
["if"] = true,
["for"] = true,
["while"] = true,
["repeat"] = true,
}
local function is_block_close(token) return token == "end" or token == "until" end
-- (internal) Walk one source file and return a list of
-- {line, depth, token} entries where depth > max_nesting.
local function audit_file(path, max_nesting)
local f = io.open(path, "r")
if not f then error("Cannot open " .. path) end
local content = f:read("*a")
f:close()
local violations = {}
local depth = 0
local line = 1
local pos = 1
local src_len = #content
local token_idx = 0
local function read_ident_at(start_pos)
local ident_start = start_pos
if ident_start > src_len then return nil end
local first_ch = content:sub(ident_start, ident_start)
if not (first_ch:match("[%a_]")) then return nil end
local scan = start_pos + 1
while scan <= src_len do
local ch = content:sub(scan, scan)
if not (ch:match("[%w_]")) then break end
scan = scan + 1
end
return content:sub(ident_start, scan - 1), scan
end
-- Skip past a string literal or comment starting at `start_pos`.
-- Returns the position just past the construct, or nil if `start_pos`
-- is not the start of a string/comment.
local function skip_string_or_comment(start_pos)
local ch = content:sub(start_pos, start_pos)
if ch == '"' or ch == "'" then
local scan = start_pos + 1
while scan <= src_len do
local c = content:sub(scan, scan)
if c == "\\" then scan = scan + 2
elseif c == ch then return scan + 1
else scan = scan + 1
end
end
return src_len + 1
elseif ch == "-" and content:sub(start_pos + 1, start_pos + 1) == "-" then
local scan = start_pos + 2
if content:sub(scan, scan + 1) == "[[" and content:sub(scan + 2, scan + 3) == "[" then
-- Long bracket comment [==[ ... ]==]
scan = scan + 2
local eq = ""
while content:sub(scan, scan) == "=" do
eq = eq .. "="
scan = scan + 1
end
local close_marker = "]" .. eq .. "]"
local close_pos = content:find(close_marker, scan, true)
if close_pos then
return close_pos + #close_marker
else
return src_len + 1
end
else
while scan <= src_len and content:sub(scan, scan) ~= "\n" do scan = scan + 1 end
return scan + 1
end
elseif ch == "[" and content:sub(start_pos + 1, start_pos + 1) == "[" then
local scan = start_pos + 2
local eq = ""
while content:sub(scan, scan) == "=" do
eq = eq .. "="
scan = scan + 1
end
local close_marker = "]" .. eq .. "]"
local close_pos = content:find(close_marker, scan, true)
if close_pos then
return close_pos + #close_marker
else
return src_len + 1
end
end
return nil
end
while pos <= src_len do
local ch = content:sub(pos, pos)
if ch == "\n" then line = line + 1 end
local skip_to = skip_string_or_comment(pos)
if skip_to then
for scan = pos, skip_to - 1 do
if content:sub(scan, scan) == "\n" then line = line + 1 end
end
pos = skip_to
elseif ch:match("[%a_]") then
local tok, next_pos = read_ident_at(pos)
token_idx = token_idx + 1
if BLOCK_OPEN[tok] then
depth = depth + 1
if depth > max_nesting then
violations[#violations + 1] = {
line = line,
depth = depth,
token = tok,
}
end
elseif is_block_close(tok) then
depth = depth - 1
end
pos = next_pos
else
pos = pos + 1
end
end
return violations
end
--- Audit one file. Returns nil if clean, else a list of violations.
--- @param path string
--- @param max_nesting integer -- default 5
--- @return table|nil
function M.audit(path, max_nesting)
local violations = audit_file(path, max_nesting or 5)
if #violations == 0 then return nil end
return violations
end
-- Module CLI.
if arg and arg[1] then
local max_nesting = 5
local files = {}
for arg_idx = 1, #arg do
if arg[arg_idx] == "--max" and arg[arg_idx + 1] then
max_nesting = tonumber(arg[arg_idx + 1]) or 5
else
files[#files + 1] = arg[arg_idx]
end
end
-- Accept either a directory or a file path. Directory args are
-- expanded via lfs.dir (native, no subprocess).
local lfs = require("lfs")
local function is_dir(p)
return lfs.attributes(p, "mode") == "directory"
end
local function list_lua(dir)
local out = {}
if not is_dir(dir) then return out end
for entry in lfs.dir(dir) do
if entry:match("%.lua$") then
out[#out + 1] = dir .. "/" .. entry
end
end
return out
end
local to_check = {}
for _, f in ipairs(files) do
if is_dir(f) then
for _, sub in ipairs(list_lua(f)) do to_check[#to_check + 1] = sub end
else
to_check[#to_check + 1] = f
end
end
local total_violations = 0
for _, f in ipairs(to_check) do
local v = M.audit(f, max_nesting)
if v then
io.write(string.format("\n%s\n", f))
for _, x in ipairs(v) do
io.write(string.format(" line %d: depth %d (after '%s')\n", x.line, x.depth, x.token))
end
total_violations = total_violations + #v
end
end
if total_violations == 0 then
io.write("OK: no files exceed max nesting of " .. max_nesting .. "\n")
os.exit(0)
else
io.write(string.format("\n%d nesting violation(s) found.\n", total_violations))
os.exit(1)
end
end
return M
+489 -179
View File
@@ -1,3 +1,13 @@
# --- Parameter Surface (Task 8) -----------------------------------------
# -Reload : After a successful build, invoke reload.ps1 as a child pwsh and propagate its exit code.
# -HelperZipOnly : Skip the build entirely; regenerate the helper zip and exit. Honors -HelperZipOutput for out-of-tree paths.
# -HelperZipOutput: When -HelperZipOnly is set, writes the archive to this path instead of the scripts/pcsx_debug_helper.zip.
param(
[switch]$Reload,
[switch]$HelperZipOnly,
[string]$HelperZipOutput = ''
)
$path_root = split-path -Path $PSScriptRoot -Parent
$path_build = join-path $path_root 'build'
$path_code = join-path $path_root 'code'
@@ -8,6 +18,98 @@ if ((test-path $path_build) -eq $false) {
new-item -itemtype directory -path $path_build
}
# --- HelperZipOnly short-circuit ----------------------------------------
# Must run before any compile/link work.
# Inlines the same logic as Make-HelperZip below to avoid an extra pwsh process spawn (~200 ms).
#The helper zip is small and the BCL call is in-process; cold ~14 ms, warm ~10 ms (assembly load + tiny zip write).
if ($HelperZipOnly) {
$zipDest = if ([string]::IsNullOrEmpty($HelperZipOutput)) {
join-path $path_scripts 'pcsx_debug_helper.zip'
}
else {
$HelperZipOutput
}
$HelperDir = join-path $path_scripts 'pcsx_debug_helper'
$elf32Src = join-path $path_scripts 'elf32.lua'
$elf32Dest = join-path $HelperDir 'elf32.lua'
if (-not (test-path -LiteralPath $HelperDir)) {
write-error "helper dir not found: $HelperDir"
exit 1
}
if (-not (test-path -LiteralPath $elf32Src)) {
write-error "elf32.lua not found at $elf32Src"
exit 1
}
write-host "[build] HelperZipOnly mode -> $zipDest"
# --- Timestamp gate (Fix 1) -------------------------------------------
# PCSX-Redux holds pcsx_debug_helper.zip open via -archive at startup.
# The zip is consumed once at startup; the reload endpoint reads it
# from package.loaded on subsequent calls. Writing it on every build
# is dead work that fights the file lock. Skip the rewrite when the
# three sources (autoexec.lua, reload.lua, elf32.lua) are all older
# than the existing zip.
$sources = @(
(join-path $HelperDir 'autoexec.lua'),
(join-path $HelperDir 'reload.lua'),
$elf32Src
)
$zipMtime = $null
if (test-path -LiteralPath $zipDest) {
$zipMtime = (Get-Item -LiteralPath $zipDest).LastWriteTime
}
$needsRewrite = $false
if ($null -eq $zipMtime) {
$needsRewrite = $true
}
else {
foreach ($s in $sources) {
if (-not (test-path -LiteralPath $s)) { continue }
if ((Get-Item -LiteralPath $s).LastWriteTime -gt $zipMtime) {
$needsRewrite = $true
break
}
}
}
if (-not $needsRewrite) {
$sz = (Get-Item -LiteralPath $zipDest).Length
Write-Host "[build] helper zip up to date: $zipDest ($sz bytes); skipping"
return
}
Copy-Item -LiteralPath $elf32Src -Destination $elf32Dest -Force
try {
# Force the inode release so CreateFromDirectory can write fresh.
# ZipFile.CreateFromDirectory throws if the destination exists.
# If PCSX-Redux holds the file open, Remove-Item raises — fall
# back to writing pcsx_debug_helper.zip.new alongside. The next
# PCSX-Redux restart will read the canonical path; the .new file
# is a hint for the optional launch-script patch in fix 3.
if (test-path -LiteralPath $zipDest) {
try {
# -ErrorAction Stop is required so the catch below fires.
# Remove-Item raises a non-terminating error by default
# (ErrorActionPreference=Continue), which bypasses catch.
Remove-Item -LiteralPath $zipDest -Force -ErrorAction Stop
}
catch {
$zipDest = [System.IO.Path]::ChangeExtension($zipDest, '.zip.new')
Write-Warning "[build] canonical helper zip is locked; writing to $zipDest instead"
}
}
Add-Type -AssemblyName System.IO.Compression.FileSystem
[System.IO.Compression.ZipFile]::CreateFromDirectory(
$HelperDir, $zipDest,
[System.IO.Compression.CompressionLevel]::Optimal, $false) | Out-Null
$sz = (Get-Item -LiteralPath $zipDest).Length
Write-Host "[build] wrote $sz bytes to $zipDest"
}
finally {
if (test-path -LiteralPath $elf32Dest) { Remove-Item -LiteralPath $elf32Dest -Force }
}
return
}
# --- Toolchain Definition ---
# Assumes 'mipsel-none-elf' toolchain is in your system's PATH.
$Prefix = "mipsel-none-elf"
@@ -81,19 +183,6 @@ $path_psyq = join-path $path_toolchain 'psyq-4_7'
$path_psyq_iwyu = join-path $path_toolchain 'psyq_iwyu'
$path_psyq_imyu_inc = join-path $path_psyq_iwyu 'include'
function Get-SourceFiles { param([Parameter(Mandatory=$true)] [string[]]$paths, [Parameter(Mandatory=$true)] [string[]]$extensions)
$files = @()
foreach ($p in $paths) {
if (-not (test-path $p)) { continue }
foreach ($ext in $extensions) {
Get-ChildItem -Path $p -File -Recurse -Filter "*$ext" -ErrorAction SilentlyContinue | ForEach-Object {
$files += $_.FullName
}
}
}
return ($files | Sort-Object -Unique)
}
function assemble-unit { param(
[string] $unit,
[string] $link_module,
@@ -153,7 +242,7 @@ function compile-unit { param(
$f_arch_no_shared,
$f_arch_no_stack_prot
)
# $compile_args += $f_std_c23
$compile_args += $f_std_c11
$compile_args += ($f_include + $path_psyq_imyu_inc)
$compile_args += ($f_include + $path_nugget)
@@ -193,29 +282,18 @@ function link-modules { param([string[]]$link_modules, [string] $elf, [string[]
$link_args += ($f_link_pass_through_prefix + $f_link_mapfile + $map)
$link_args += ($f_link_pass_through_prefix + $f_link_start_group)
# raw_sio_pad_poll_20260802 — Task 5.1c surgical library-list trim.
# The 16 removed entries (c2, card, cd, comb, ds, gs, gun, hmd, math,
# mcrd, mcx, press, sio, snd, spu, tap) had LOAD lines in the map but
# ZERO .o files pulled in — they were unused. The 5 kept libraries
# (api, c, etc, gpu, gte) are required by the C-side calls in
# hello_joypad.c (reset_graph, draw_sync, vsync, etc.).
$libraries = @(
"api",
"c",
"c2",
"card",
"cd",
"comb",
"ds",
"etc",
"gpu",
"gs",
"gte",
"gun",
"hmd",
"math",
"mcrd",
"mcx",
"pad",
"press",
"sio",
"snd",
"spu",
"tap"
"gte"
)
foreach ($lib in $libraries) {
$link_args += ($f_link_lib + $lib)
@@ -238,11 +316,158 @@ function link-modules { param([string[]]$link_modules, [string] $elf, [string[]
function make-binary { param([string]$elf, [string]$exe)
Write-Host "--- Creating Binary ---" -ForegroundColor Cyan
write-host "Converting $elf to PS-EXE -> '$exe'"
$objcopy_args = ($f_objcopy_format + "binary"), $elf, $exe
$objcopy_args = ($f_objcopy_format + "binary"), $elf, $exe
& $Objcopy $objcopy_args
if ($LASTEXITCODE -ne 0) { Write-Error "Objcopy failed. Aborting."; exit 1 }
}
function ps1-meta { param(
[string] $unity_root,
[string[]]$sources,
[Parameter(Mandatory=$true)][string]$metadata,
[string] $out_root = (join-path $path_build 'gen'),
[string[]]$passes = @('--pre-link'),
[string[]]$extra_args = @()
)
# `--unity-root` and `--source` are
# mutually exclusive. Exactly one of `$unity_root` / `$sources` must
# be supplied; the other must be absent.
if ($null -ne $unity_root -and $unity_root -ne '')
{
if ($null -ne $sources -and $sources.Count -gt 0) {
write-error 'ps1-meta: -unity_root and -sources are mutually exclusive'
exit 2
}
}
elseif ($null -eq $sources -or $sources.Count -eq 0) {
write-error 'ps1-meta: either -unity_root <file> or -sources <file...> is required'
exit 2
}
# --- Defensive attribute clear on tracked gen files ------------------------
# Git tracks code/<dir>/gen/*.h files and Windows keeps the Archive bit set
# on them. Combined with transient editor locks or co-running processes,
# this can make io.open(path, "wb") fail with Access Denied / Sharing
# Violation even though Get-ChildItem shows IsReadOnly = False. Clearing
# the Read-only + Archive bits locally is safe; git re-asserts them on
# the next operation but the metaprogram write always wins.
#
# Derived from the caller's parameters: $metadata lives in $path_duffle
# (so its parent is the duffle dir), and $unity_root / $sources[0] lives
# in $path_module (so its parent is the module dir).
$pathToDuffle = split-path -Path $metadata -Parent
$pathToModule = $null
if ($null -ne $unity_root -and $unity_root -ne '') {
$pathToModule = split-path -Path $unity_root -Parent
}
elseif ($null -ne $sources -and $sources.Count -gt 0) {
$pathToModule = split-path -Path $sources[0] -Parent
}
$genFiles = @(
join-path $pathToDuffle 'gen\macs.h'
join-path $pathToDuffle 'gen\offsets.h'
)
if ($null -ne $pathToModule) {
$genFiles += join-path $pathToModule 'gen\macs.h'
$genFiles += join-path $pathToModule 'gen\offsets.h'
}
foreach ($f in $genFiles) {
if (test-path -LiteralPath $f) {
attrib -R $f 2>&1 | Out-Null
attrib -A $f 2>&1 | Out-Null
}
}
$script = join-path $path_scripts 'ps1_meta.lua'
$input_summary = if ($null -ne $unity_root -and $unity_root -ne '') {
"unity=$unity_root"
}
else {
"$($sources.Count) source(s)"
}
write-host "ps1-meta $input_summary, passes=$($passes -join ',')" ` -ForegroundColor Magenta
$arg_list = @($passes) + @('--metadata', $metadata) + @('--out-root', $out_root) + @($extra_args)
if ($null -ne $unity_root -and $unity_root -ne '') {
$arg_list += @('--unity-root', $unity_root)
}
else {
foreach ($s in $sources) { $arg_list += @('--source', $s) }
}
& luajit $script @arg_list
if ($LASTEXITCODE -ne 0) {
write-error "ps1-meta failed (exit $LASTEXITCODE). Aborting."
exit $LASTEXITCODE
}
}
function inject-dwarf { param(
[string]$elf,
[string]$path_gen
)
$base_name = [System.IO.Path]::GetFileNameWithoutExtension($elf)
$path_dwarf_line_bin = join-path $path_gen "$base_name.dwarf_line.bin"
$path_dwarf_aranges_bin = join-path $path_gen "$base_name.dwarf_aranges.bin"
$path_dwarf_rnglists_bin = join-path $path_gen "$base_name.dwarf_rnglists.bin"
$path_dwarf_info_bin = join-path $path_gen "$base_name.dwarf_info.bin"
$path_dwarf_abbrev_bin = join-path $path_gen "$base_name.dwarf_abbrev.bin"
$path_dwarf_str_bin = join-path $path_gen "$base_name.dwarf_str.bin"
$path_dwarf_loc_bin = join-path $path_gen "$base_name.dwarf_loc.bin"
$path_dwarf_loclists_bin = join-path $path_gen "$base_name.dwarf_loclists.bin"
$path_inject_elf = join-path $path_build "$base_name.dwarf-injected.elf"
if (-not (Test-Path $path_dwarf_line_bin)) { return }
if (-not (Test-Path $path_dwarf_aranges_bin)) { return }
if (-not (Test-Path $path_dwarf_rnglists_bin)) { return }
Write-Host "[build] DWARF-injecting $elf -> $path_inject_elf"
Copy-Item -LiteralPath $elf -Destination $path_inject_elf -Force
# Objcopy call 1: 3x --update-section for the PC-mapping tables (line, aranges, rnglists).
$objcopy_args_dwarf_pc = @(
"--update-section=.debug_line=$path_dwarf_line_bin",
"--update-section=.debug_aranges=$path_dwarf_aranges_bin",
"--update-section=.debug_rnglists=$path_dwarf_rnglists_bin"
)
& $Objcopy @objcopy_args_dwarf_pc $path_inject_elf 2>&1 | Out-Null
if ($LASTEXITCODE -ne 0) {
Write-Warning "[build] objcopy dwarf-pc splice failed (exit $LASTEXITCODE); removing $path_inject_elf"
Remove-Item -LiteralPath $path_inject_elf -ErrorAction SilentlyContinue
return
}
# Objcopy call 2: 3x --update-section + 2x --add-section for the debug-data tables (info, abbrev, str, loc, loclists).
$objcopy_args_dwarf_info = @(
"--update-section=.debug_info=$path_dwarf_info_bin",
"--update-section=.debug_abbrev=$path_dwarf_abbrev_bin",
"--update-section=.debug_str=$path_dwarf_str_bin",
"--add-section=.debug_loc=$path_dwarf_loc_bin",
"--add-section=.debug_loclists=$path_dwarf_loclists_bin"
)
& $Objcopy @objcopy_args_dwarf_info $path_inject_elf 2>&1 | Out-Null
if ($LASTEXITCODE -ne 0) {
Write-Warning "[build] objcopy dwarf-info splice failed (exit $LASTEXITCODE); removing $path_inject_elf"
Remove-Item -LiteralPath $path_inject_elf -ErrorAction SilentlyContinue
return
}
# Baked atoms execute from RAM but are emitted as C data arrays, so their ELF sections lack SHF_EXECINSTR.
# GDB discards line rows for non-code sections. Mark only the debug-copy sections executable.
# The original ELF and PS-EXE remain byte/flag unchanged.
& $Objcopy `
--set-section-flags ".rodata=alloc,load,readonly,code,contents" `
--set-section-flags ".data=alloc,load,data,code,contents" `
$path_inject_elf 2>&1 | Out-Null
if ($LASTEXITCODE -ne 0) {
Write-Warning "[build] atom-section flag update failed (exit $LASTEXITCODE); removing $path_inject_elf"
Remove-Item -LiteralPath $path_inject_elf -ErrorAction SilentlyContinue
}
else {
Write-Host "[build] DWARF-injected ELF: $path_inject_elf"
}
}
# inject-dwarf
function build-hello_psyqo {
$includes += @()
@@ -287,7 +512,7 @@ function build-graphis_hello {
$src_asm_crt = join-path $path_nugget_common 'crt0/crt0.s'
$module_asm_crt = join-path $path_build 'crt0.o'
# assemble-unit $src_asm_crt $module_asm_crt $includes $assemble_args
assemble-unit $src_asm_crt $module_asm_crt $includes $assemble_args
$src_asm = join-path $path_module 'hello_gpu.s'
$module_asm = join-path $path_build 'hello_gpu.o'
@@ -317,114 +542,16 @@ function build-graphis_hello {
}
# build-graphis_hello
function generate-TapeAtomOffsets {param([Parameter(Mandatory=$true)] [string[]]$sources, [Parameter(Mandatory=$true)] [string]$metadata)
$gen_atom_offsets_script = join-path $path_scripts 'tape_atom.offset_gen.meta.lua'
$any_stale = $false
foreach ($src in $sources) {
$basename = [System.IO.Path]::GetFileNameWithoutExtension($src)
$dir = split-path -Path $src -Parent
$gen_dir = join-path $dir 'gen'
$out = join-path $gen_dir "$basename.offsets.h"
if (-not (test-path $out)) { $any_stale = $true; break }
$src_mtime = (get-item $src).LastWriteTimeUtc
$out_mtime = (get-item $out).LastWriteTimeUtc
$meta_mtime = (get-item $metadata).LastWriteTimeUtc
if (($src_mtime -gt $out_mtime) -or ($meta_mtime -gt $out_mtime)) {
$any_stale = $true
break
}
}
if (-not $any_stale) {
write-host "AtomOffsets all $($sources.Count) source(s) up-to-date" -ForegroundColor DarkGray
return
}
write-host "AtomOffsets $($sources.Count) source(s)" -ForegroundColor Magenta
& lua $gen_atom_offsets_script $metadata @sources
if ($LASTEXITCODE -ne 0) {
write-error "Atom offset generation failed. Aborting."
exit 1
}
}
function generate-TapeAtomAnnotations {param([Parameter(Mandatory=$true)] [string[]]$sources, [Parameter(Mandatory=$true)] [string]$metadata)
# Sibling to generate-TapeAtomOffsets. Validates TAPE_ATOM_* / TAPE_WORDS
# annotations against the metadata manifest. Emits gen/<basename>.errors.h
# containing #error directives for the C build to fail on annotation drift.
$gen_atom_annot_script = join-path $path_scripts 'tape_atom_annotation_pass.lua'
$any_stale = $false
foreach ($src in $sources) {
$basename = [System.IO.Path]::GetFileNameWithoutExtension($src)
$dir = split-path -Path $src -Parent
$gen_dir = join-path $dir 'gen'
$out_txt = join-path $gen_dir "$basename.annotations.txt"
$out_err = join-path $gen_dir "$basename.errors.h"
if (-not (test-path $out_txt) -or -not (test-path $out_err)) { $any_stale = $true; break }
$src_mtime = (get-item $src).LastWriteTimeUtc
$out_txt_mtime = (get-item $out_txt).LastWriteTimeUtc
$out_err_mtime = (get-item $out_err).LastWriteTimeUtc
$out_mtime = if ($out_txt_mtime -gt $out_err_mtime) { $out_txt_mtime } else { $out_err_mtime }
$meta_mtime = (get-item $metadata).LastWriteTimeUtc
if (($src_mtime -gt $out_mtime) -or ($meta_mtime -gt $out_mtime)) {
$any_stale = $true
break
}
}
if (-not $any_stale) {
write-host "AtomAnnotations all $($sources.Count) source(s) up-to-date" -ForegroundColor DarkGray
return
}
write-host "AtomAnnotations $($sources.Count) source(s)" -ForegroundColor Magenta
& lua $gen_atom_annot_script $metadata @sources
if ($LASTEXITCODE -ne 0) {
write-error "Atom annotation generation failed. Aborting."
exit 1
}
# If any source produced annotation errors, surface them now and halt the
# build. The errors.h files are also #include'd via -include below, so
# the C build would fail at preprocessing time anyway — failing here gives
# a more readable error in the build log.
$err_count = 0
foreach ($src in $sources) {
$basename = [System.IO.Path]::GetFileNameWithoutExtension($src)
$dir = split-path -Path $src -Parent
$gen_dir = join-path $dir 'gen'
$ann_txt = join-path $gen_dir "$basename.annotations.txt"
$err_h = join-path $gen_dir "$basename.errors.h"
if ((test-path $ann_txt) -and (test-path $err_h)) {
$txt = get-content $ann_txt -raw
if ($txt -match 'Errors:\s+([1-9]\d*)') {
$err_count += [int]$Matches[1]
write-warning "Annotation errors in $src — see $ann_txt"
}
}
}
if ($err_count -gt 0) {
write-error "Annotation pass failed: $err_count error(s) across $($sources.Count) source(s). Aborting."
exit 1
}
}
function build-gte_hello {
function build-hello_gte {
$includes += @()
$path_module = join-path $path_code 'gte_hello'
$path_module = join-path $path_code 'hello_gte'
$path_duffle = join-path $path_code 'duffle'
$path_atom_metadata = join-path $path_module 'tape_atom.metadata.h'
$path_atom_metadata = join-path $path_duffle 'word_count.metadata.h'
$path_build_gen = join-path $path_build 'gen'
$source_dirs = @($path_duffle, $path_module)
$atom_sources = Get-SourceFiles -paths $source_dirs -extensions @('.h', '.c')
generate-TapeAtomAnnotations -sources $atom_sources -metadata $path_atom_metadata
generate-TapeAtomOffsets -sources $atom_sources -metadata $path_atom_metadata
$src_c = join-path $path_module 'hello_gte.c'
ps1-meta -unity_root $src_c -metadata $path_atom_metadata -out_root $path_build_gen
$assemble_args = @()
$assemble_args += $f_debug
@@ -433,14 +560,13 @@ function build-gte_hello {
$src_asm_crt = join-path $path_nugget_common 'crt0/crt0.s'
$module_asm_crt = join-path $path_build 'crt0.o'
# assemble-unit $src_asm_crt $module_asm_crt $includes $assemble_args
assemble-unit $src_asm_crt $module_asm_crt $includes $assemble_args
# $src_asm = join-path $path_module 'hello_gte.s'
# $module_asm = join-path $path_build 'hello_gte.o'
# assemble-unit $src_asm $module_asm $includes $assemble_args
$src_c = join-path $path_module 'hello_gte.c'
$module_c = join-path $path_build 'hello_gte_c.o'
$compile_args = @()
@@ -458,54 +584,238 @@ function build-gte_hello {
$link_args = @()
$link_args += $f_debug
# $link_args += $f_optimize_size
link-modules @($module_asm_crt, $module_c) $elf $link_args
$link_modules = @(
$module_asm_crt,
$module_c
)
link-modules $link_modules $elf $link_args
make-binary $elf $exe
# Post-link: gdb-runtime + dwarf-injection in a single Lua invocation (one luajit cold start).
ps1-meta -unity_root $src_c -metadata $path_atom_metadata -out_root $path_build_gen -passes @('--post-link') ` -extra_args @('--elf', $elf)
inject-dwarf $elf $path_build_gen
}
build-gte_hello
# build-hello_gte
function build-hello_joypad {
$includes += @()
# NO idea if this works yet...
function Send-ToEmulator { param(
[string]$exePath
)
$uri = "http://localhost:8080/api/v1/load-exec"
$path_module = join-path $path_code 'hello_joypad'
$path_duffle = join-path $path_code 'duffle'
$path_atom_metadata = join-path $path_duffle 'word_count.metadata.h'
$path_build_gen = join-path $path_build 'gen'
# Absolute path is safest for the emulator web server
$absolutePath = [System.IO.Path]::GetFullPath($exePath)
$src_c = join-path $path_module 'hello_joypad.c'
ps1-meta -unity_root $src_c -metadata $path_atom_metadata -out_root $path_build_gen
# Create JSON payload pointing to your compiled .ps-exe
$body = @{
filename = $absolutePath
} | ConvertTo-Json
$assemble_args = @()
$assemble_args += $f_debug
$assemble_args += $f_optimize_none
$assemble_args += ($f_include + $path_code)
Write-Host "Pushing hot-reload to PCSX-Redux..." -ForegroundColor Magenta
try {
$response = Invoke-RestMethod -Uri $uri -Method Post -Body $body -ContentType "application/json"
Write-Host "Hot-reload successful!" -ForegroundColor Green
} catch {
Write-Warning "Could not connect to PCSX-Redux web server. Ensure the emulator is running and Web Server is enabled."
}
$src_asm_crt = join-path $path_nugget_common 'crt0/crt0.s'
$module_asm_crt = join-path $path_build 'crt0.o'
assemble-unit $src_asm_crt $module_asm_crt $includes $assemble_args
$module_c = join-path $path_build 'hello_joypad_c.o'
$compile_args = @()
$compile_args += $f_debug
$compile_args += $f_optimize_none
# $compile_args += $f_optimize_intrinsics
# $compile_args += $f_optimize_size
# $compile_args += $f_optimize_debug
$compile_args += ($f_include + $path_code)
compile-unit $src_c $module_c $includes $compile_args
$elf = join-path $path_build 'hello_joypad.elf'
$exe = join-path $path_build 'hello_joypad.ps-exe'
$link_args = @()
$link_args += $f_debug
# $link_args += $f_optimize_size
$link_modules = @(
$module_asm_crt,
$module_c
)
link-modules $link_modules $elf $link_args
make-binary $elf $exe
# Post-link: gdb-runtime + dwarf-injection in a single Lua invocation (one luajit cold start).
ps1-meta -unity_root $src_c -metadata $path_atom_metadata -out_root $path_build_gen -passes @('--post-link') ` -extra_args @('--elf', $elf)
inject-dwarf $elf $path_build_gen
}
# build-hello_joypad
function build-hello_camera {
$includes += @()
$path_module = join-path $path_code 'hello_camera'
$path_duffle = join-path $path_code 'duffle'
$path_atom_metadata = join-path $path_duffle 'word_count.metadata.h'
$path_build_gen = join-path $path_build 'gen'
$src_c = join-path $path_module 'hello_camera.c'
ps1-meta -unity_root $src_c -metadata $path_atom_metadata -out_root $path_build_gen
$assemble_args = @()
$assemble_args += $f_debug
$assemble_args += $f_optimize_none
$assemble_args += ($f_include + $path_code)
$src_asm_crt = join-path $path_nugget_common 'crt0/crt0.s'
$module_asm_crt = join-path $path_build 'crt0.o'
assemble-unit $src_asm_crt $module_asm_crt $includes $assemble_args
$module_c = join-path $path_build 'hello_camera_c.o'
$compile_args = @()
$compile_args += $f_debug
$compile_args += $f_optimize_none
# $compile_args += $f_optimize_intrinsics
# $compile_args += $f_optimize_size
# $compile_args += $f_optimize_debug
$compile_args += ($f_include + $path_code)
compile-unit $src_c $module_c $includes $compile_args
$elf = join-path $path_build 'hello_camera.elf'
$exe = join-path $path_build 'hello_camera.ps-exe'
$link_args = @()
$link_args += $f_debug
# $link_args += $f_optimize_size
$link_modules = @(
$module_asm_crt,
$module_c
)
link-modules $link_modules $elf $link_args
make-binary $elf $exe
# Post-link: gdb-runtime + dwarf-injection in a single Lua invocation (one luajit cold start).
ps1-meta -unity_root $src_c -metadata $path_atom_metadata -out_root $path_build_gen -passes @('--post-link') ` -extra_args @('--elf', $elf)
inject-dwarf $elf $path_build_gen
}
build-hello_camera
# ── Helper-zip + reload helpers (Task 8) ──
# Defined right after the final build-hello_camera function so they're in scope for the post-build calls below.
# The Make-HelperZip function is also reused by the -HelperZipOnly short-circuit at the top of this script.
# Both call the in-process BCL CreateFromDirectory rather than spawning a child pwsh to avoid the ~200 ms process-spawn overhead.
function Make-HelperZip {
param([string]$OutputPath = '')
$dest = if ([string]::IsNullOrEmpty($OutputPath)) {
join-path $path_scripts 'pcsx_debug_helper.zip'
}
else {
$OutputPath
}
$HelperDir = join-path $path_scripts 'pcsx_debug_helper'
$elf32Src = join-path $path_scripts 'elf32.lua'
$elf32Dest = join-path $HelperDir 'elf32.lua'
if (-not (test-path -LiteralPath $HelperDir)) {
write-warning "[build] helper dir not found: $HelperDir; skipping helper zip"
return
}
if (-not (test-path -LiteralPath $elf32Src)) {
write-warning "[build] elf32.lua not found at $elf32Src; skipping helper zip"
return
}
# --- Timestamp gate (Fix 1) -------------------------------------------
# PCSX-Redux holds pcsx_debug_helper.zip open via -archive at startup.
# The zip is consumed once at startup; the reload endpoint reads it
# from package.loaded on subsequent calls. Writing it on every build
# is dead work that fights the file lock. Skip the rewrite when the
# three sources (autoexec.lua, reload.lua, elf32.lua) are all older
# than the existing zip.
$sources = @(
(join-path $HelperDir 'autoexec.lua'),
(join-path $HelperDir 'reload.lua'),
$elf32Src
)
$zipMtime = $null
if (test-path -LiteralPath $dest) {
$zipMtime = (Get-Item -LiteralPath $dest).LastWriteTime
}
$needsRewrite = $false
if ($null -eq $zipMtime) {
$needsRewrite = $true
}
else {
foreach ($s in $sources) {
if (-not (test-path -LiteralPath $s)) { continue }
if ((Get-Item -LiteralPath $s).LastWriteTime -gt $zipMtime) {
$needsRewrite = $true
break
}
}
}
if (-not $needsRewrite) {
$sz = (Get-Item -LiteralPath $dest).Length
Write-Host "[build] helper zip up to date: $dest ($sz bytes); skipping"
return
}
write-host "[build] regenerating helper zip -> $dest"
Copy-Item -LiteralPath $elf32Src -Destination $elf32Dest -Force
try {
# Force the inode release so CreateFromDirectory can write fresh.
# ZipFile.CreateFromDirectory throws if the destination exists.
# If PCSX-Redux holds the file open, Remove-Item raises — fall
# back to writing pcsx_debug_helper.zip.new alongside. The next
# PCSX-Redux restart will read the canonical path; the .new file
# is a hint for the optional launch-script patch in fix 3.
if (test-path -LiteralPath $dest) {
try {
# -ErrorAction Stop is required so the catch below fires.
# Remove-Item raises a non-terminating error by default
# (ErrorActionPreference=Continue), which bypasses catch.
Remove-Item -LiteralPath $dest -Force -ErrorAction Stop
}
catch {
$dest = [System.IO.Path]::ChangeExtension($dest, '.zip.new')
Write-Warning "[build] canonical helper zip is locked; writing to $dest instead"
}
}
Add-Type -AssemblyName System.IO.Compression.FileSystem
[System.IO.Compression.ZipFile]::CreateFromDirectory(
$HelperDir, $dest,
[System.IO.Compression.CompressionLevel]::Optimal, $false) | Out-Null
$sz = (Get-Item -LiteralPath $dest).Length
Write-Host "[build] wrote $sz bytes to $dest"
}
finally {
if (test-path -LiteralPath $elf32Dest) { Remove-Item -LiteralPath $elf32Dest -Force }
}
}
# # Automatically hot-reloads it into the running emulator
# Send-ToEmulator (join-path $path_build 'hello_gte.ps-exe')
# Invokes reload.ps1 as a child pwsh instead of POSTing to the nonexistent /api/v1/load-exec endpoint.
# Exit code is propagated so the build fails loud if the reload fails.
function Send-ToEmulator {
param([string]$ElfPath = (join-path $path_build 'hello_camera.elf'))
# --- Hot Reload via PCSX-Redux Web Server ---
# $exe_path = join-path $path_build 'hello_gte.ps-exe'
# $absolute_path = [System.IO.Path]::GetFullPath($exe_path)
$reloadScript = join-path $path_scripts 'reload.ps1'
if (-not (test-path -LiteralPath $reloadScript)) {
write-error "[build] reload.ps1 not found at $reloadScript"
exit 1
}
# PCSX-Redux expects the file location in the URL query string?
# We URL-encode the path to ensure backslashes and spaces don't break the HTTP request?
# $encoded_path = [uri]::EscapeDataString($absolute_path)
# $uri = "http://localhost:8080/api/v1/load-exec?path=$encoded_path"
write-host "[build] hot-reloading $ElfPath via reload.ps1" -ForegroundColor Magenta
& pwsh -NoProfile -File $reloadScript -Mode elf -Target hello_camera -ElfPath $ElfPath
if ($LASTEXITCODE -ne 0) {
write-error "[build] reload.ps1 failed (exit $LASTEXITCODE)"
exit $LASTEXITCODE
}
}
# Write-Host "Pushing hot-reload to PCSX-Redux..." -ForegroundColor Magenta
# try {
# # Send the request with the query string included
# Invoke-RestMethod -Uri $uri -Method Post
# Write-Host "Hot-reload successful!" -ForegroundColor Green
# } catch {
# Write-Host "Failed to hot-reload." -ForegroundColor Red
# # This will print the *actual* HTTP error instead of our generic warning
# Write-Host $_.Exception.Message -ForegroundColor Yellow
# }
# Post-build: Regenerate the helper zip (canonical output) and, if -Reload was passed, kick a hot-reload against the just-built ELF.
# Any future targets compiled by this script should add their own Make-HelperZip call after their build step; today's only target is hello_camera.
Make-HelperZip
if ($Reload) {
Send-ToEmulator
}
+2615
View File
File diff suppressed because it is too large Load Diff
+92
View File
@@ -0,0 +1,92 @@
--- duffle_paths.lua — Single-line bootstrap helper for the tape-atom Lua scripts.
---
--- Each entry script (ps1_meta.lua + the 7 passes/*.lua files) starts with one of:
--- ```lua
--- -- Entry script (ps1_meta.lua — `arg[0]` is set):
--- local duffle = dofile((arg[0]:match("(.*[/\\])") or "./") .. "duffle_paths.lua")
---
--- -- Pass module (debug.getinfo path resolution; works both standalone and when require'd):
--- local _bootstrap_dir = debug.getinfo(1, "S").source:match("^@?(.*[/\\])") or "./"
--- local duffle = dofile(_bootstrap_dir .. "../duffle_paths.lua")
--- ```
---
--- That small bootstrap: (a) locates this helper via `arg[0]` / `debug.getinfo`,
--- (b) loads it (which sets `package.path` + `package.cpath`),
--- (c) at the bottom calls `require("duffle")` (now resolvable since `package.path` was just set) and returns the duffle M.
--- Net effect: the caller gets the duffle module in one statement; no separate `dofile(...)` + `require("duffle")` dance.
---
local M = {}
-- Cache key for the repo root. Stored in `package.loaded` (process-global) so all 8 entry scripts + passes scripts share one resolution.
local CACHE_KEY = "__duffle_repo_root__"
--- Resolve the repo root from this script's own path. Zero shell spawn.
--- `duffle_paths.lua` always lives at `<repo>/scripts/duffle_paths.lua`, so the repo root is the
--- parent of the directory containing this script. We derive it directly from `debug.getinfo(1, "S").source`
--- (returns `@<path>` for the currently-running chunk).
---
--- If `debug.getinfo` can't parse this script's path (shouldn't happen — dofile always populates source),
--- return nil and let `M.setup()` fail loud.
--- @return string|nil
local function find_repo_root()
if package.loaded[CACHE_KEY] then return package.loaded[CACHE_KEY] end
local source = debug.getinfo(1, "S").source
-- Strip the leading `@` (Lua's dofile marker) and the trailing `/duffle_paths.lua` filename.
-- What remains is the directory containing this script, i.e. `<repo>/scripts/`.
local scripts_dir = source and source:match("^@?(.*)[/\\]duffle_paths%.lua$")
if not scripts_dir then return nil end
-- The repo root is the parent of `scripts/`. Strip the trailing `scripts/` (with or without trailing slash).
local root = scripts_dir:gsub("scripts[\\/]?$", "")
root = root:gsub("\\", "/")
if root == "" then root = "./" end
if not root:match("/$") then root = root .. "/" end
package.loaded[CACHE_KEY] = root
return root
end
--- Set `package.path` (for `require("duffle")` + `require("passes.X")`) and
--- `package.cpath` (for `lpeg.dll`).
---
--- This script does NOT touch the OS environment: no `os.setenv`, no `os.putenv`, no `$PATH` mods.
--- It just sets `package.path` and `package.cpath` (the standard Lua way to register module search dirs).
--- lpeg is built by `update_deps.ps1` to `toolchain/lpeg/`,
--- which we wire into `package.cpath` here (so `require("lpeg")` from `duffle.lua` resolves without any global state).
function M.setup()
local repo_root = find_repo_root()
if not repo_root then
-- Unreachable in practice: find_repo_root() derives the repo root from this script's
-- own source path via debug.getinfo(1, "S").source (no subprocess, no git CLI, <1ms).
-- A nil return means the source path did not match the expected
-- <repo>/scripts/duffle_paths.lua layout — a packaging bug, not a "missing git repo"
-- condition. os.exit(2) is retained so a real failure surfaces loud rather than
-- silently producing an unconfigured module table.
os.exit(2)
end
local scripts_dir = repo_root .. "scripts/"
local passes_dir = repo_root .. "scripts/passes/"
package.path = scripts_dir .. "?.lua;"
.. scripts_dir .. "?/init.lua;"
.. passes_dir .. "?.lua;"
.. passes_dir .. "?/init.lua;"
.. package.path
-- lpeg: built by `update_deps.ps1` to `toolchain/lpeg/lpeg.dll`.
-- lfs: compiled from pcsx-redux's vendored luafilesystem source to `toolchain/lfs/lfs.dll`.
-- Wire both directories into cpath so `require("lpeg")` and `require("lfs")` resolve.
local lpeg_dir = repo_root .. "toolchain/lpeg/"
local lfs_dir = repo_root .. "toolchain/lfs/"
package.cpath = lpeg_dir .. "?.dll;"
.. lfs_dir .. "?.dll;"
.. package.cpath
end
-- Run the setup as a side effect.
M.setup()
-- Now that package.path includes scripts/, `require("duffle")` resolves. Return the duffle module
-- so callers can do `local duffle = dofile(...duffle_paths.lua)` in one line.
return require("duffle")
File diff suppressed because it is too large Load Diff
+105
View File
@@ -0,0 +1,105 @@
# scripts/gdb/gdb_tape_atoms.gdb
#
# Wrapper for the tape-atom step-debug helpers.
# The 9 user commands are defined here as STUBS (degraded-state messages).
# The real implementations + the per-atom data tables are emitted by `passes/atoms_source_map.lua`
# (post-link invocation: `ps1_meta.lua --atoms-source-map --gdb-runtime --elf <elf>`) into `build/gdb_tape_atoms_runtime.gdb`.
# Sourcing that file RE-DEFINES the commands with real implementations.
#
# If `build/gdb_tape_atoms_runtime.gdb` is missing or stale, the stubs remain (E1: no source map).
# The user just needs to re-run `build_psyq.ps1` to regenerate.
# ?? Stub commands (defined here so they're always present, even if the runtime file is missing). The runtime file overrides these if sourced. ??
define tape_atoms
echo "[gdb_tape_atoms] STUB: runtime file build/gdb_tape_atoms_runtime.gdb not found."
echo "[gdb_tape_atoms] STUB: run .\\build_psyq.ps1 to regenerate, then re-source this file."
end
document tape_atoms
List every tape atom symbol in the loaded ELF (code_<name>) with its .rodata address and word count.
STUB state: runtime file not sourced. Run build_psyq.ps1 to regenerate.
end
define break_atom
echo "[gdb_tape_atoms] STUB: build/gdb_tape_atoms_runtime.gdb not sourced. Run build_psyq.ps1."
end
document break_atom
Set a breakpoint at the start of tape atom <name>. STUB state.
end
define step_atom
echo "[gdb_tape_atoms] STUB: build/gdb_tape_atoms_runtime.gdb not sourced. Run build_psyq.ps1."
end
document step_atom
Resume execution until the next atom boundary. STUB state.
end
define next_atom
echo "[gdb_tape_atoms] STUB: build/gdb_tape_atoms_runtime.gdb not sourced. Run build_psyq.ps1."
end
document next_atom
Alias for step_atom. STUB state.
end
define where_in_atom
echo "[gdb_tape_atoms] STUB: build/gdb_tape_atoms_runtime.gdb not sourced. Run build_psyq.ps1."
end
document where_in_atom
Report current atom name, .rodata addr, word offset, and source line (if known). STUB state.
end
define stepi_inside_atom
echo "[gdb_tape_atoms] STUB: build/gdb_tape_atoms_runtime.gdb not sourced. Run build_psyq.ps1."
end
document stepi_inside_atom
One MIPS-instruction step, then where_in_atom. STUB state.
end
define show_c2
printf "C2[ 0] 0x%08x\n", $c2_data[0]
printf "C2[ 7] 0x%08x [otz]\n", $c2_data[7]
printf "C2[12] 0x%08x [sxy0]\n", $c2_data[12]
printf "C2[13] 0x%08x [sxy1]\n", $c2_data[13]
printf "C2[14] 0x%08x [sxy2]\n", $c2_data[14]
printf "C2[24] 0x%08x [mac0]\n", $c2_data[24]
printf "...\n"
echo "(STUB state: only 7 representative regs shown. Run build_psyq.ps1 for full dump.)"
end
document show_c2
Pretty-print all 32 C2 data registers as hex + named alias. STUB state (7 reg subset).
end
define show_c2ctl
printf "C2CTL[ 0] 0x%08x\n", $c2_control[0]
printf "...\n"
echo "(STUB state: only 1 reg shown. Run build_psyq.ps1 for full dump.)"
end
document show_c2ctl
Pretty-print all 32 C2 control registers. STUB state (1 reg subset).
end
define wave_ctx
printf "$t4 = R_FaceCursor 0x%08x\n", $t4
printf "$t5 = R_VertBase 0x%08x\n", $t5
printf "$t6 = R_OtBase 0x%08x\n", $t6
printf "$t7 = R_PrimCursor 0x%08x\n", $t7
end
document wave_ctx
Pretty-print the 4 wave-context GPRs ($t4..$t7). (wave_ctx works in stub state too.)
end
# ?? Source the runtime file (re-defines commands with real impls + data). ??
# Try to source from project-root-relative path first (the typical case).
# If the user is in a different CWD, the source will fail and stubs remain.
# The runtime file path is computed relative to the ELF's source map convention (build/gdb_tape_atoms_runtime.gdb).
echo [gdb_tape_atoms] Wrapper loaded. Sourcing runtime file...
# Suppress the "Redefine command" prompts that would otherwise appear when the runtime file overrides the 9 stub commands defined above.
# The runtime's `define` blocks are intended to overwrite ? there's no ambiguity to confirm.
set confirm off
# Source the runtime file (re-defines commands with real impls + data).
source build/gdb_tape_atoms_runtime.gdb
set confirm on
echo [gdb_tape_atoms] Runtime sourced successfully (9 commands now have real implementations).
+174
View File
@@ -0,0 +1,174 @@
# scripts/launch_pcsx_debug.ps1
#
# One-shot launcher for debug sessions: starts pcsx-redux with the .ps-exe
# loaded, the gdb stub enabled, the web server enabled, AND the
# pcsx_debug_helper Lua plugin loaded so external CLI tools can drive
# reloads via http://localhost:8080/api/v1/lua/reload.
#
# After launch:
# - gdb: target remote localhost:3333
# - web: POST http://localhost:8080/api/v1/lua/reload?mode=prime&...
#
# usage:
# .\scripts\launch_pcsx_debug.ps1
# .\scripts\launch_pcsx_debug.ps1 -ExePath build\hello_camera.ps-exe
# .\scripts\launch_pcsx_debug.ps1 -Cpu dynarec
# .\scripts\launch_pcsx_debug.ps1 -ElfPath build\hello_camera.elf
#
# Companion: scripts/debug_psyq.ps1 (bare launch — no .ps-exe, no helper).
[CmdletBinding()]
param(
[string]$PcsxPath = (Join-Path $PSScriptRoot '..\toolchain\pcsx-redux\vsprojects\x64\Release\pcsx-redux.exe'),
[string]$ExePath = (Join-Path $PSScriptRoot '..\build\hello_camera.ps-exe'),
[string]$ElfPath = '',
[string]$HelperZip = (Join-Path $PSScriptRoot 'pcsx_debug_helper.zip'),
[int] $GdbPort = 3333,
[int] $WebPort = 8080,
[ValidateSet('interpreter', 'dynarec')][string]$Cpu = 'interpreter'
)
$ErrorActionPreference = 'Stop'
# ── Derive -ElfPath when absent ──
# Convention: the .elf sits beside the .ps-exe with the same stem.
if ([string]::IsNullOrEmpty($ElfPath)) {
$exeFull = [System.IO.Path]::GetFullPath($ExePath)
$stem = [System.IO.Path]::GetFileNameWithoutExtension($exeFull)
$exeDir = [System.IO.Path]::GetDirectoryName($exeFull)
$ElfPath = Join-Path $exeDir "$stem.elf"
}
# ── Pre-checks ──
foreach ($p in @($PcsxPath, $ExePath, $ElfPath, $HelperZip)) {
if (-not (Test-Path -LiteralPath $p)) {
Write-Error "Missing: $p"
exit 1
}
}
# ── Reject a stale helper zip (Task 8) ──
# The helper zip must be newer than every .lua source that contributes
# to it. A stale zip means the running plugin does not match the on-disk
# source, which makes the reload contract meaningless.
$helperDir = Join-Path $PSScriptRoot 'pcsx_debug_helper'
$elf32Src = Join-Path $PSScriptRoot 'elf32.lua'
$sourceLuas = @(
(Join-Path $helperDir 'autoexec.lua'),
(Join-Path $helperDir 'reload.lua'),
$elf32Src
) | Where-Object { Test-Path -LiteralPath $_ }
$zipTime = (Get-Item -LiteralPath $HelperZip).LastWriteTime
$stale = $false
foreach ($src in $sourceLuas) {
$srcTime = (Get-Item -LiteralPath $src).LastWriteTime
if ($srcTime -gt $zipTime) {
Write-Error "helper zip is older than source: $src (zip=$($zipTime.ToString('o')) src=$($srcTime.ToString('o')); rerun build_psyq.ps1 to regenerate."
$stale = $true
}
}
if ($stale) {
exit 1
}
# Kill any existing pcsx-redux so the archive file isn't locked.
Get-Process pcsx-redux -ErrorAction SilentlyContinue | Stop-Process -Force
Start-Sleep -Seconds 2
# ── Launch ──
$absExe = [System.IO.Path]::GetFullPath($ExePath)
$absZip = [System.IO.Path]::GetFullPath($HelperZip)
$cpuFlag = if ($Cpu -eq 'dynarec') { '-dynarec' } else { '-interpreter' }
$args = @(
'-gdb', '-run'
'-loadexe', "`"$absExe`""
'-archive', "`"$absZip`""
'-webserver'
$cpuFlag
)
Write-Host "Launching pcsx-redux..." -ForegroundColor Cyan
Write-Host " ps-exe : $absExe"
Write-Host " elf : $ElfPath"
Write-Host " helper zip: $absZip"
Write-Host " gdb : localhost:$GdbPort"
Write-Host " web : localhost:$WebPort/api/v1/lua/reload"
Write-Host " cpu : $Cpu ($cpuFlag)"
Write-Host ""
Start-Process -FilePath $PcsxPath -ArgumentList $args | Out-Null
# ── Wait for both endpoints to come up ──
$deadline = (Get-Date).AddSeconds(15)
while ((Get-Date) -lt $deadline) {
$gdbUp = $false
$webUp = $false
try {
$tcp = New-Object System.Net.Sockets.TcpClient
$tcp.BeginConnect('localhost', $GdbPort, $null, $null) | Out-Null
Start-Sleep -Milliseconds 100
$gdbUp = $tcp.Connected
$tcp.Close()
} catch { $gdbUp = $false }
try {
$r = Invoke-WebRequest -Uri "http://localhost:$WebPort/" -UseBasicParsing -TimeoutSec 1 -ErrorAction SilentlyContinue
$webUp = $r.StatusCode -ne 0
} catch { $webUp = $false }
if ($gdbUp -and $webUp) { break }
Start-Sleep -Milliseconds 500
}
# ── Smoke-test the gte handler ──
try {
$r = Invoke-WebRequest -Uri "http://localhost:$WebPort/api/v1/lua/gte" -UseBasicParsing -TimeoutSec 5
$firstLine = ([System.Text.Encoding]::UTF8.GetString($r.Content) -split "`n")[0]
Write-Host "GTE handler OK: $firstLine" -ForegroundColor Green
} catch {
Write-Warning "GTE handler NOT responding: $_"
Write-Host "Check the pcsx-redux Lua Console for debug cli messages." -ForegroundColor Yellow
}
# ── Prime the reload handler (Task 8) ──
# The reload handler keeps an internal ACTIVE manifest of the running
# ELF; reload requests fail with reload_not_primed until prime succeeds.
# We retry until the response carries ok=true or the launch deadline
# expires — the helper may not have finished registering handlers in the
# first web-poll cycle after the gte handler comes up.
$absElf = [System.IO.Path]::GetFullPath($ElfPath)
$encodedPath = [uri]::EscapeDataString($absElf)
$primeUri = "http://localhost:${WebPort}/api/v1/lua/reload?mode=prime&target=hello_camera&path=${encodedPath}"
Write-Host "Priming reload handler: $primeUri" -ForegroundColor Cyan
$primeDeadline = (Get-Date).AddSeconds(15)
$primeOk = $false
while ((Get-Date) -lt $primeDeadline) {
try {
$resp = Invoke-WebRequest -Method Post -Uri $primeUri -UseBasicParsing -TimeoutSec 5
$body = if ($resp.Content -is [byte[]]) {
[System.Text.Encoding]::UTF8.GetString([byte[]]$resp.Content)
} else {
[string]$resp.Content
}
$obj = $body | ConvertFrom-Json
if ($obj.ok) {
Write-Host "Prime OK: $(($obj | ConvertTo-Json -Compress))" -ForegroundColor Green
$primeOk = $true
break
} else {
Write-Host "Prime not yet ready: error=$($obj.error)" -ForegroundColor Yellow
}
} catch {
Write-Host "Prime request failed: $($_.Exception.Message)" -ForegroundColor Yellow
}
Start-Sleep -Milliseconds 500
}
if (-not $primeOk) {
Write-Warning "Prime did not return ok=true before the launch deadline. Reload requests will fail until the user primes manually."
}
Write-Host ""
Write-Host "pcsx-redux running. PIDs:" -ForegroundColor Cyan
Get-Process pcsx-redux | Select-Object Id, ProcessName | Format-Table
+82
View File
@@ -0,0 +1,82 @@
# make_helper_zip.ps1
#
# Regenerate scripts/pcsx_debug_helper.zip from scripts/pcsx_debug_helper/.
# The archive contains exactly three entries at archive root:
#
# autoexec.lua
# elf32.lua (copied in from scripts/elf32.lua before packaging)
# reload.lua
#
# Determinism: CreateFromDirectory on the same set of files produces
# identical bytes. Verified by running the same command twice and
# asserting SHA-256 equality (see plan.md Task 6 Step 4).
#
# Performance: the implementation uses System.IO.Compression.ZipFile
# (BCL, in-process). Benchmarked: ~2 ms cold, ~2 ms warm on this
# workstation. Compress-Archive is rejected because its first call
# takes ~200 ms (assembly load) and subsequent calls take ~16 ms
# (process spawn per invocation). The 50 ms budget documented in
# plan.md Task 8 Step 3 excludes the compiler/assembler toolchain.
#
# Usage:
# pwsh -NoProfile -File scripts\make_helper_zip.ps1
#
# Optional -OutputPath switches the destination. Default is
# scripts/pcsx_debug_helper.zip next to the helper dir.
#
# Companion: scripts/pcsx_debug_helper/{autoexec,elf32,reload}.lua
# tests/reload_helper_zip_regen.ps1 (planned Task 8 verifier)
[CmdletBinding()]
param(
[string]$HelperDir = (Join-Path $PSScriptRoot 'pcsx_debug_helper'),
[string]$SourcesDir = $PSScriptRoot,
[string]$OutputPath = (Join-Path $PSScriptRoot 'pcsx_debug_helper.zip')
)
$ErrorActionPreference = 'Stop'
if (-not (Test-Path -LiteralPath $HelperDir)) {
throw "helper dir not found: $HelperDir"
}
# Stage elf32.lua into the helper dir so the in-process ZipFile walker
# picks it up alongside the helper-local files. elf32.lua is the shared
# ELF32 byte reader; the production reload.lua loads it through
# Support.extra.dofile("elf32.lua") at runtime.
$elf32Src = Join-Path $SourcesDir 'elf32.lua'
$elf32Dest = Join-Path $HelperDir 'elf32.lua'
if (-not (Test-Path -LiteralPath $elf32Src)) {
throw "elf32.lua not found at $elf32Src"
}
Copy-Item -LiteralPath $elf32Src -Destination $elf32Dest -Force
try {
# Remove any existing archive so CreateFromDirectory can write fresh.
# ZipFile.CreateFromDirectory throws if the destination exists.
if (Test-Path -LiteralPath $OutputPath) {
Remove-Item -LiteralPath $OutputPath -Force
}
# In-process zip; ~2 ms cold, ~2 ms warm. BCL compression matches
# Compress-Archive at CompressionLevel Optimal for these small files.
# Assembly is loaded once per pwsh.exe; the first run pays ~14 ms,
# subsequent runs pay ~0.2 ms.
Add-Type -AssemblyName System.IO.Compression.FileSystem
[System.IO.Compression.ZipFile]::CreateFromDirectory(
$HelperDir, $OutputPath,
[System.IO.Compression.CompressionLevel]::Optimal,
$false) | Out-Null
$sha = (Get-FileHash -LiteralPath $OutputPath -Algorithm SHA256).Hash
Write-Output ("[make_helper_zip] wrote {0} bytes, sha256={1}" -f `
(Get-Item -LiteralPath $OutputPath).Length, $sha)
Write-Output "[make_helper_zip] entries: autoexec.lua, elf32.lua, reload.lua"
}
finally {
# Remove the staged elf32.lua so the helper directory only contains
# the files the user expects to see there.
if (Test-Path -LiteralPath $elf32Dest) {
Remove-Item -LiteralPath $elf32Dest -Force
}
}
+639
View File
@@ -0,0 +1,639 @@
--- passes/annotation.lua — Atom-annotation DSL validator.
---
--- Validates `MipsAtom_(name) atom_info(atom_bind(Binds_X), atom_reads(...), atom_writes(...)) { ... }` declarations in source files.
--- Also reads `Binds_*` struct declarations (`typedef Struct_(Binds_X) { ... };`).
---
--- `duffle.scan_source()` scans each source once upstream; `ps1_meta.lua` stores that result in `src.scan`.
---
--- Ownership: the canonical `ctx.shared.corpus` supplies cross-source registries, while each `src.scan` supplies its source's declarations and bodies.
--- A context without `ctx.shared.corpus` is rejected with an explicit canonical-corpus message.
-- Bootstrap follows the entry scripts; `scripts/duffle_paths.lua` sets package.path and package.cpath. See `ps1_meta.lua` for the rationale.
-- `debug.getinfo(1, "S").source` locates this file for standalone and orchestrated runs, then `duffle_paths.lua` returns the loaded `duffle` module.
local _bootstrap_dir = debug.getinfo(1, "S").source:match("^@?(.*[/\\])") or "./"
local duffle = dofile(_bootstrap_dir .. "../duffle_paths.lua")
-- The annotation pass reads the source-derived registries from scan_source:
-- * pipe_ctx.register_alias_registry — for atom_dbg_reg_default(R_X, ...) and atom_reg_types(R_X, ...) member-identity checks
-- * pipe_ctx.type_name_registry — for atom_dbg_reg_default(<T>, ...) and atom_reg_types(<T>, ...) type-identity checks
-- ════════════════════════════════════════════════════════════════════════════
-- Type declarations
-- ════════════════════════════════════════════════════════════════════════════
--- @class SourceFile
--- @field path string -- Absolute path to the source file
--- @field text string -- Full source text
--- @field dir string -- Directory containing the source
--- @field basename string -- Filename without extension
--- @field scan table -- Pre-scanned SourceScan payload (from duffle.scan_source)
--- @class PassCtx
--- @field sources SourceFile[]
--- @field metadata_path string
--- @field shared table
--- @field shared.word_counts table<string, integer>
--- @field out_root string
--- @field project_root string
--- @field upstream table<string, table>
--- @field flags table
--- @field verbose boolean
--- @class PassResult
--- @field outputs table[]
--- @field errors table[]
--- @field warnings table[]
--- @class AtomAnnotation
--- @field line integer -- Source line of the atom_info call
--- @field macro string -- Macro name (always "atom_info" in the new shape)
--- @field name string -- Atom name
--- @field kind string -- Always "info"
--- @field binds string|nil -- Binds_X name if any
--- @field reads string[] -- R_* names (read targets)
--- @field writes string[] -- R_* names (write targets)
--- @field errors string[]|nil -- Parse-time errors from scan_source (atom_info body malformed)
--- @class DebugSkipMarker -- Sub-shape of scan_source.lua's @class DebugSkipMarker
--- @field marker_kind string -- Exact marker ident read from source. Only "atom_dbg_skip" (bare) is positive.
--- @field marker_line integer
--- @field args string|nil -- Trimmed text inside the parens (nil when has_parens is false)
--- @field has_parens boolean
--- @field is_bare boolean -- true iff marker_kind == "atom_dbg_skip" AND has_parens == false (the only positive form)
--- @field pending boolean -- true while awaiting the following declaration
--- @field superseded_by_marker_line integer|nil -- Set on a marker that was bumped out of the pending slot
--- @field target_kind string|nil -- "atom" | "comp_bare" | "comp_proc" | "unrelated" once observed
--- @class Finding
--- @field line integer -- Source line (or 0 for pass-level)
--- @field msg string -- Finding message
--- @class Findings
--- @field errors Finding[]
--- @field warnings Finding[]
--- @field info Finding[]
--- @class PipeCtx
--- @field atom_index table<string, AtomAnnotation> -- Name -> AtomAnnotation (only kind=="atom")
--- @field binds_index table<string, BindsStruct> -- Name -> BindsStruct
--- @field annot_counts table<string, integer> -- Name -> annotation count (for unique_annotation check)
--- @field types table<string, RegTypeDefault> -- From scan_source
--- @field atom_views table<string, AtomViewEntry> -- From scan_source
--- @field seen_defaults table<string, integer> -- Duplicate atom_dbg_reg_default detection
--- @field seen_field table<string, integer> -- Binds_* -> count of fields (set/checked by check_binds_no_duplicate_fields)
--- @field _scan SourceScan -- Full scan payload (typed-view sub-calls live here)
--- @class AnnotatedResult
--- @field atoms AtomEntry[]
--- @field annots AtomAnnotation[]
--- @field macros MacroEntry[]
--- @field binds BindsEntry[]
--- @field errors Finding[]
--- @field warnings Finding[]
--- @field info Finding[]
-- ════════════════════════════════════════════════════════════════════════════
-- Per-check functions (the CHECK_RULES table's payload)
-- ════════════════════════════════════════════════════════════════════════════
--- The dispatcher in `validate()` routes each result by convention: existence checks write errors[] and shape checks write warnings[].
--- `macro_word_drift` writes errors[] for missing or mismatched metadata and info[] for a match.
--- Check: Every annotated atom must have a matching MipsAtom_(name) declaration.
--- @param a AtomAnnotation
--- @param pipe_ctx PipeCtx
--- @param findings Findings
local function check_atom_decl_exists(a, pipe_ctx, findings)
if not pipe_ctx.atom_index[a.name] then
findings.errors[#findings.errors + 1] = {
line = a.line,
msg = string.format("annotation for '%s' has no matching MipsAtom_(%s) { ... }", a.name, a.name),
}
end
end
--- Check: Every atom may have AT MOST ONE annotation.
--- Post-loop: Needs full-corpus `annot_counts` from pipe_ctx.
--- @param pipe_ctx PipeCtx
--- @param findings Findings
local function check_unique_annotation(pipe_ctx, findings)
for name, n in pairs(pipe_ctx.annot_counts) do
if n > 1 then
findings.errors[#findings.errors + 1] = {
line = pipe_ctx.atom_index[name] and pipe_ctx.atom_index[name].line or 0,
msg = string.format("MipsAtom_(%s) has %d annotations (expected at most 1)", name, n),
}
end
end
end
--- Check: BIND atoms must reference a real Binds_* struct.
--- I keep this as a warning so the annotation pass can report the common test-fixture case; `check_abi_handoff` in static analysis supplies the build-stopping error.
--- @param a AtomAnnotation
--- @param pipe_ctx PipeCtx
--- @param findings Findings
local function check_binds_struct_exists(a, pipe_ctx, findings)
if not a.binds then return end
if pipe_ctx.binds_index[a.binds] then return end
findings.warnings[#findings.warnings + 1] = {
line = a.line,
msg = string.format("'%s' binds '%s' but no Struct_(%s) { ... } "
.. "declaration found (also flagged as an error by check_abi_handoff in the static-analysis pass)"
, a.name, a.binds, a.binds),
}
end
--- Check: TAPE_WORDS(mac_X, N) ↔ WORD_COUNT(mac_X, N) drift.
--- Three outcomes: missing (error), mismatch (error), match (info).
--- @param m MacroEntry
--- @param wc table<string, integer> -- Shared word-count table (from ctx.shared.word_counts)
--- @param findings Findings
local function check_macro_word_drift(m, wc, findings)
local declared = wc[m.name]
if not declared then
findings.errors[#findings.errors + 1] = {
line = m.line,
msg = string.format("TAPE_WORDS(%s, %d) but '%s' is not in metadata.h", m.name, m.words, m.name),
}
return
end
if declared ~= m.words then
findings.errors[#findings.errors + 1] = {
line = m.line,
msg = string.format("DRIFT: TAPE_WORDS(%s, %d) but metadata.h declares WORD_COUNT(%s, %d)", m.name, m.words, m.name, declared),
}
return
end
findings.info[#findings.info + 1] = {
line = m.line,
msg = string.format("OK: %s = %d words", m.name, m.words),
}
end
--- Check: atom_dbg_reg_default(R_X, <type>) targets an alias in `pipe_ctx.register_alias_registry` and a type in `pipe_ctx.type_name_registry`.
--- Pointer depth remains bounded to 0 or 1, and duplicate defaults remain errors.
--- @param _src SourceFile -- unused (kept for the per_source shape)
--- @param pipe_ctx PipeCtx
--- @param findings Findings
local function check_semantic_reg_defaults(_src, pipe_ctx, findings)
-- Detect duplicate defaults using the ordered occurrence list (the out.types hash only retains the last declaration).
local seen_first_line = {}
for _, occ in ipairs(pipe_ctx.type_occurrences or {}) do
if seen_first_line[occ.reg] == nil then
seen_first_line[occ.reg] = occ.source_line
else
findings.errors[#findings.errors + 1] = {
line = occ.source_line,
msg = string.format(
"duplicate atom_dbg_reg_default for %q at line %d (first declared at line %d); one default per register",
occ.reg, occ.source_line, seen_first_line[occ.reg]),
}
end
end
local reg_registry = pipe_ctx.register_alias_registry or {}
local type_registry = pipe_ctx.type_name_registry or {}
for reg, def in pairs(pipe_ctx.types or {}) do
if not reg_registry[reg] then
findings.errors[#findings.errors + 1] = {
line = def.source_line,
msg = string.format(
"atom_dbg_reg_default at line %d references unknown register %q (not in register_alias_registry)",
def.source_line, reg),
}
end
if def.pointer_depth == nil or def.pointer_depth < 0 or def.pointer_depth > 1 then
findings.errors[#findings.errors + 1] = {
line = def.source_line,
msg = string.format(
"atom_dbg_reg_default at line %d for %q has unsupported pointer depth %d (expected 0 or 1)",
def.source_line, reg, def.pointer_depth or -1),
}
end
if not def.type_name or not type_registry[def.type_name] then
findings.errors[#findings.errors + 1] = {
line = def.source_line,
msg = string.format(
"atom_dbg_reg_default at line %d for %q uses unknown type %q (not in type_name_registry)",
def.source_line, reg, tostring(def.type_name)),
}
end
end
end
--- Check: atom_reg_types(R_X, <type>) entries target an alias in `pipe_ctx.register_alias_registry` and a type in `pipe_ctx.type_name_registry`.
--- A bare `atom_reg` marker opts the `R_<n>` alias into GPR identity; references to R_T0..R_T3 require the same explicit marker.
--- @param _src SourceFile
--- @param pipe_ctx PipeCtx
--- @param findings Findings
local function check_atom_reg_types(_src, pipe_ctx, findings)
local reg_registry = pipe_ctx.register_alias_registry or {}
local type_registry = pipe_ctx.type_name_registry or {}
for _, ai in ipairs(pipe_ctx.atom_infos_list or {}) do
if ai.reg_type_overrides then
for reg, ov in pairs(ai.reg_type_overrides) do
if not reg_registry[reg] then
findings.errors[#findings.errors + 1] = {
line = ai.info_line,
msg = string.format(
"atom '%s' has atom_reg_types for %q; compute-register types are restricted to opt-in aliases (%q not in register_alias_registry)",
ai.atom_name, reg, reg),
}
end
if not ov.type_name or not type_registry[ov.type_name] then
findings.errors[#findings.errors + 1] = {
line = ai.info_line,
msg = string.format(
"atom '%s' atom_reg_types for %q uses unknown compute type %q (not in type_name_registry)",
ai.atom_name, reg, tostring(ov.type_name)),
}
end
end
end
end
end
--- Check: atom_view(Binds_X) entries reference a Binds_* struct with at least one field.
--- @param _src SourceFile
--- @param pipe_ctx PipeCtx
--- @param findings Findings
local function check_atom_view_layout(_src, pipe_ctx, findings)
for atom_name, view in pairs(pipe_ctx.atom_views or {}) do
if not view.binds_name then
-- The atom had atom_reg_types but no atom_view; no layout check needed.
else
local bs = pipe_ctx.binds_index[view.binds_name]
if not bs then
findings.errors[#findings.errors + 1] = {
line = view.info_line,
msg = string.format(
"atom '%s' has atom_view(%s) but no Struct_(%s) { ... } declaration was found",
atom_name, view.binds_name, view.binds_name),
}
elseif not bs.fields or #bs.fields == 0 then
findings.errors[#findings.errors + 1] = {
line = bs.line,
msg = string.format(
"atom '%s' has atom_view(%s) but that struct declares zero typed fields",
atom_name, view.binds_name),
}
end
end
end
end
--- Check: Binds_* structs require unique field names because atom_view uses those names for typed-field lookup in gdb.
--- @param _src SourceFile
--- @param pipe_ctx PipeCtx
--- @param findings Findings
local function check_binds_no_duplicate_fields(_src, pipe_ctx, findings)
for _, bs in ipairs(pipe_ctx.binds_list or {}) do
local seen = {}
for _, f in ipairs(bs.fields or {}) do
seen[f.name] = (seen[f.name] or 0) + 1
end
for name, count in pairs(seen) do
if count > 1 then
findings.errors[#findings.errors + 1] = {
line = bs.line,
msg = string.format(
"%s has duplicate field name %q (count %d); the typed-view contract requires unique field names",
bs.name, name, count),
}
end
end
end
end
-- Check: Debug-skip markers must satisfy shape + placement constraints.
--- Walks the priority list once; each marker produces at most one error, so one source defect yields one finding.
--- Priority order (first defect wins):
--- 1. marker_kind ~= "atom_dbg_skip" -> legacy/renamed spelling (use `atom_dbg_skip`)
--- 2. marker_kind == "atom_dbg_skip" AND has_parens -> parenthesized form (the marker is bare-only)
--- 3. args ~= "" -> takes no arguments
--- 4. superseded_by_marker_line -> duplicate marker (cite superseding line)
--- 5. pending + no target_kind -> dangling (no following declaration)
--- 6. unsupported target_kind -> marker precedes an unrelated declaration
--- Valid markers stamp `debug_skip` on whole-atom, bare-component, and proc-component declaration records in scan_source.lua.
--- @param marker DebugSkipMarker
--- @param _pipe_ctx PipeCtx -- Unused; kept for consistency with per_annot // TODO(Ed): Remove?
--- @param findings Findings
local function check_skip_marker(marker, _pipe_ctx, findings)
local kind = marker.marker_kind
local line = marker.marker_line
-- Left `scan.debug_skip_markers` with production records for `atom_dbg_skip` only; other identifiers take the walker's unrelated branch.
if marker.has_parens then
findings.errors[#findings.errors + 1] = {
line = line,
msg = string.format("%s marker at line %d must be bare; the parenthesized form is no longer accepted (use `atom_dbg_skip MipsAtom_(name) { ... }`)",
kind, line),
}
return
end
if marker.args ~= nil and marker.args ~= "" then
findings.errors[#findings.errors + 1] = {
line = line,
msg = string.format("%s marker at line %d takes no arguments; found %q", kind, line, marker.args),
}
return
end
if marker.superseded_by_marker_line then
findings.errors[#findings.errors + 1] = {
line = line,
msg = string.format("duplicate %s marker at line %d; superseded by another %s marker at line %d"
, kind, line, kind, marker.superseded_by_marker_line),
}
return
end
if marker.pending and not marker.target_kind then
findings.errors[#findings.errors + 1] = {
line = line,
msg = string.format("dangling %s marker at line %d: no following MipsAtom_/MipsAtomComp_/MipsAtomComp_Proc_ declaration"
, kind, line),
}
return
end
if marker.target_kind
and marker.target_kind ~= "atom"
and marker.target_kind ~= "comp_bare"
and marker.target_kind ~= "comp_proc" then
findings.errors[#findings.errors + 1] = {
line = line,
msg = string.format("%s marker at line %d must precede MipsAtom_, MipsAtomComp_, or MipsAtomComp_Proc_; found an unrelated declaration"
, kind, line),
}
end
end
--- Warn when a source references an unregistered alias.
--- When a source uses an unregistered R_X, this check emits one pass-level info entry for that source and directs C-ABI register names to explicit alias registration.
--- @param _src SourceFile
--- @param pipe_ctx PipeCtx
--- @param findings Findings
local function check_wave_context_migration(_src, pipe_ctx, findings)
if not (pipe_ctx.types and next(pipe_ctx.types)) then return end
if not (pipe_ctx.atom_infos_list) then return end
local reg_registry = pipe_ctx.register_alias_registry or {}
for _, ai in ipairs(pipe_ctx.atom_infos_list) do
if ai.reg_type_overrides then
for reg, _ in pairs(ai.reg_type_overrides) do
if not reg_registry[reg] then
findings.warnings[#findings.warnings + 1] = {
line = 0,
msg = "wave-context removed; opt in via #define atom_reg in mips.h "
.. "(every R_<alias> that should be visible to the annotation pass "
.. "must be enum-declared with the bare atom_reg marker)",
}
return
end
end
end
end
end
-- ════════════════════════════════════════════════════════════════════════════
-- CHECK_RULES — data-driven check dispatch (the plex pattern)
-- ════════════════════════════════════════════════════════════════════════════
--
-- Each rule entry picks one of four "shapes" of dispatch:
-- per_annot(annot, pipe_ctx, findings) -- runs once per AtomAnnotation
-- post(pipe_ctx, findings) -- runs once after all per_annot calls complete (full-corpus aggregation)
-- per_macro(macro, wc, findings) -- runs once per TAPE_WORDS / _Pragma macro declaration
-- per_skip_marker(marker, pipe_ctx, findings) -- runs once per src.scan.debug_skip_markers entry
--
-- Adding a new check = 1 row here + 1 function above. The `validate()` dispatch loop never needs editing.
local CHECK_RULES = {
{ name = "atom_decl_exists", per_annot = check_atom_decl_exists },
{ name = "binds_struct_exists", per_annot = check_binds_struct_exists },
{ name = "unique_annotation", post = check_unique_annotation },
{ name = "macro_word_drift", per_macro = check_macro_word_drift },
{ name = "skip_marker_validation", per_skip_marker = check_skip_marker },
{ name = "semantic_reg_defaults", per_source = check_semantic_reg_defaults },
{ name = "atom_reg_types", per_source = check_atom_reg_types },
{ name = "atom_view_layout", per_source = check_atom_view_layout },
{ name = "binds_no_duplicate_fields", per_source = check_binds_no_duplicate_fields },
{ name = "wave_context_migration", per_source = check_wave_context_migration },
}
-- ════════════════════════════════════════════════════════════════════════════
-- Validation
-- ════════════════════════════════════════════════════════════════════════════
-- Pure check: Read from src.scan, run validations, emit findings. The scan was done once upstream.
--- Builds one pass-wide pipe_ctx from the merged `corpus.*` registries and source-ordered `corpus.atom_infos`; per-source declarations and bodies remain in `src.scan`.
--- The module ownership contract above requires callers to construct `ctx.shared.corpus` through `build_ctx`; the error message below enforces that gate.
--- @param ctx PassCtx
--- @return PipeCtx
local function build_corpus_pipe_ctx(ctx)
local corpus = ctx.shared and ctx.shared.corpus
if not corpus then
error("annotation requires ctx.shared.corpus "
.. "(the canonical corpus is the source of truth; "
.. "no per-source fallback is supported)", 0)
end
-- `corpus.atom_infos` preserves source order and duplicates; I precompute counts here for `check_unique_annotation` and the per-source checks.
local annot_counts = {}
for _, info in ipairs(corpus.atom_infos or {}) do
if info and info.atom_name then
annot_counts[info.atom_name] = (annot_counts[info.atom_name] or 0) + 1
end
end
-- Every consumer of these fields observes mutations via the canonical corpus without independently mutable registry construction.
return {
-- Cross-source lookup tables from corpus.
register_alias_registry = corpus.register_alias_registry or {},
type_name_registry = corpus.type_name_registry or {},
atom_views = corpus.atom_views or {},
atom_ctxs = corpus.atom_ctxs or {},
atom_phases = corpus.atom_phases or {},
binds_by_name = corpus.binds_by_name or {},
atoms_by_name = corpus.atoms_by_name or {},
-- Corpus-wide ordered list of atom_info records (source-order + duplicates).
atom_infos_list = corpus.atom_infos or {},
-- Corpus-wide annotation count aggregation (post-rule consumes this).
annot_counts = annot_counts,
-- Corpus-wide collisions (recorded by scan_source.merge_corpus_registries).
collisions = corpus.collisions or {},
-- `check_macro_word_drift` reads `corpus.word_counts`, populated by word_count_eval.run.
word_counts = corpus.word_counts or {},
}
end
--- Validate one source against its pre-scanned SourceScan payload + the corpus-wide pipe_ctx.
--- @param ctx PassCtx
--- @param src SourceFile
--- @param corpus_pipe_ctx PipeCtx|nil -- Built once per pass from corpus registries; nil builds the same projection here.
--- @return AnnotatedResult
local function validate(ctx, src, corpus_pipe_ctx)
corpus_pipe_ctx = corpus_pipe_ctx or build_corpus_pipe_ctx(ctx)
local scan = src.scan
-- Project the pre-scanned atoms to the AtomEntry shape this pass needs.
local atoms = {}
for _, a in ipairs(scan.atoms) do
if a.kind == "atom" then
atoms[#atoms + 1] = { line = a.line, name = a.raw_name }
end
end
-- Project the pre-scanned atom_infos to AtomAnnotation shape.
local annots = {}
for _, info in ipairs(scan.atom_infos) do
annots[#annots + 1] = {
line = info.info_line,
macro = "atom_info",
name = info.atom_name,
kind = "info",
binds = info.binds,
reads = info.reads or {},
writes = info.writes or {},
errors = info.errors,
}
end
-- Build a per-source pipe_ctx: shared lookups come from `corpus_pipe_ctx`, while declarations, bodies, types, views, defaults, and occurrences come from `src.scan`.
local seen_defaults = {}; for reg, _ in pairs (scan.types or {}) do seen_defaults[reg] = (seen_defaults[reg] or 0) + 1 end
local atom_infos_list = {}; for _, ai in ipairs(scan.atom_infos or {}) do atom_infos_list[#atom_infos_list + 1] = ai end
local pipe_ctx = {
atom_index = {},
binds_index = {},
annot_counts = corpus_pipe_ctx.annot_counts,
types = scan.types or {},
type_occurrences = scan.type_occurrences or {},
atom_views = scan.atom_views or {},
seen_defaults = seen_defaults,
atom_infos_list = atom_infos_list,
binds_list = scan.binds or {},
-- See the module ownership contract; these shared lookup tables come from corpus_pipe_ctx.
register_alias_registry = corpus_pipe_ctx.register_alias_registry,
type_name_registry = corpus_pipe_ctx.type_name_registry,
}
for _, a in ipairs(atoms) do pipe_ctx.atom_index [a.name] = a end
for _, b in ipairs(scan.binds) do pipe_ctx.binds_index[b.name] = b end
-- Findings live in a single struct with three lists (errors / warnings / info).
-- Each check writes to the list appropriate for its severity.
local findings = { errors = {}, warnings = {}, info = {} }
-- Lift parse-time errors already recorded in scan_source's atom_info payload into this pass's findings list.
for _, a in ipairs(annots) do
if a.errors then
for _, msg in ipairs(a.errors) do
findings.errors[#findings.errors + 1] = {
line = a.line,
msg = string.format("'%s': %s", a.name, msg),
}
end
end
end
-- THE per-annotation pipeline. ONE loop. CHECK_RULES dispatches per_annot rules.
for _, a in ipairs(annots) do
for _, rule in ipairs(CHECK_RULES) do
if rule.per_annot then rule.per_annot(a, pipe_ctx, findings) end
end
end
-- Post-loop rules (one-shot checks that need full-corpus aggregation in pipe_ctx).
for _, rule in ipairs(CHECK_RULES) do
if rule.post then rule.post(pipe_ctx, findings) end
end
-- scan_source records each marker in scan.debug_skip_markers; this loop validates each record independently and emits at most one error per marker.
-- Valid markers stamp `debug_skip = true` on the following atom or component declaration, which downstream consumers read directly.
local skip_markers = scan.debug_skip_markers or {}
for _, marker in ipairs(skip_markers) do
for _, rule in ipairs(CHECK_RULES) do
if rule.per_skip_marker then rule.per_skip_marker(marker, pipe_ctx, findings) end
end
end
-- Per-macro rules (TAPE_WORDS vs WORD_COUNT drift).
local wc = corpus_pipe_ctx.word_counts
for _, m in ipairs(scan.macros) do
for _, rule in ipairs(CHECK_RULES) do
if rule.per_macro then rule.per_macro(m, wc, findings) end
end
end
-- Per-source rules (reg defaults, atom_view layout, compute-register type overrides, Binds_* field uniqueness).
-- Each per_source rule sees the full scan payload via pipe_ctx.
for _, rule in ipairs(CHECK_RULES) do
if rule.per_source then rule.per_source(src, pipe_ctx, findings) end
end
-- Information summary (always emitted).
findings.info[#findings.info + 1] = {
line = 0,
msg = string.format("scanned: %d atom(s), %d annotation(s), %d macro-word-decl(s), %d binds struct(s)"
, #atoms, #annots, #scan.macros, #scan.binds),
}
return {
atoms = atoms,
annots = annots,
macros = scan.macros,
binds = scan.binds,
errors = findings.errors,
warnings = findings.warnings,
info = findings.info,
}
end
-- ════════════════════════════════════════════════════════════════════════════
-- M.run — orchestrator entry
-- ════════════════════════════════════════════════════════════════════════════
--- @class M
local M = {}
-- Expose `validate` for downstream passes (e.g. report.lua) that need to re-render the per-source results into a per-MODULE report.
M.validate = validate
--- @param ctx PassCtx
--- @return PassResult
function M.run(ctx)
local outputs = {}
local errors = {}
local warnings = {}
-- Build the shared pipe_ctx once for this run; every validate() call sees the same cross-source registries.
-- The corpus owns the canonical cross-source registries; per-source scans retain body / declaration ownership.
local corpus_pipe_ctx = build_corpus_pipe_ctx(ctx)
local corpus = ctx.shared.corpus
-- Group `corpus.sources_by_dir` by module, validate every source in each bucket, and emit one errors.h per directory.
local by_dir = (corpus and corpus.sources_by_dir) or {}
for dir, dir_sources in pairs(by_dir) do
local dir_basename = dir:match("([^/\\]+)$") or dir
local dir_atoms = 0
local dir_errors = {}
local dir_warnings = {}
for _, src in ipairs(dir_sources) do
local result = validate(ctx, src, corpus_pipe_ctx)
result.source = src.path -- tag for downstream rendering
dir_atoms = dir_atoms + #result.atoms
for _, e in ipairs(result.errors) do
dir_errors[#dir_errors + 1] = { line = e.line, msg = e.msg, source = src.path }
errors [#errors + 1] = { line = e.line, msg = e.msg }
end
for _, w in ipairs(result.warnings) do
dir_warnings[#dir_warnings + 1] = { line = w.line, msg = w.msg }
warnings [#warnings + 1] = { line = w.line, msg = w.msg }
end
end
end
return { outputs = outputs, errors = errors, warnings = warnings }
end
return M
+568
View File
@@ -0,0 +1,568 @@
--- passes/atoms_source_map.lua — Per-.word source-line map emitter for tape atoms.
---
--- Writer: this pass, given `atom.paths` (the per-atom mutable surface owned by `emission_model`). Readers:
--- `passes/dwarf_injection.lua` (synthesizes DW_TAG_inlined_subroutine + per-word line program rows) and
--- the gdb-runtime wrapper at `scripts/gdb/gdb_tape_atoms.gdb` (loads the source map via `source <path>`).
---
--- Inputs from `atom.paths`: the ordered `items` stream, dense `word_events`, `invocations` views. Outputs:
--- one `WORD N LINE L TEXT T` line per emitted `.word`, plus the per-word provenance form that DWARF synthesis consumes.
---
--- Two output forms:
--- 1. Markdown form: Handled by `passes/report.lua` (writes `<module>.atoms.md`).
--- The render functions `render_source_map` + `render_provenance` are exported for `report.lua` to call directly.
--- Compile artifacts (`*.macs.h`, `*.offsets.h`) stay in `<source_dir>/gen/`.
--- 2. `gdb_tape_atoms_runtime.gdb`: Post-link opt-in (`ctx.flags.gdb_runtime`),
--- so the gdb wrapper script + the generated runtime script share the same canonical location.
--- Triggered by `--post-link` or `--gdb-runtime`.
---
--- Output forma (sourcemap.txt form):
--- ```
--- # FORMAT_VERSION 1
--- # auto-generated by ps1_meta.lua (passes/atoms_source_map.lua) — DO NOT EDIT
--- ATOM <name> "<abs-source-path>" <total_words>
--- WORD 0 LINE 49 TEXT load_half_u(R_T0, R_FaceCursor, 0 * S_(S2)),
--- WORD 1 LINE 49 TEXT load_half_u(R_T0, R_FaceCursor, 0 * S_(S2)),
--- ... (one WORD line per .word emitted by the atom body) ...
--- ENDATOM
--- ATOM <next-name> "<abs-source-path>" <total_words>
--- ...
--- ENDATOM
--- ```
--- Marker records are zero-width in `atom.paths.items`, so they emit no WORD rows in the dense word view.
-- ════════════════════════════════════════════════════════════════════════════
-- Module-scope requires + package.path setup
-- ════════════════════════════════════════════════════════════════════════════
-- Bootstrap: load `duffle_paths.lua` via `debug.getinfo(1, "S").source`
-- (works both standalone + when require'd). `duffle_paths.lua` sets package.path then returns `require("duffle")`
-- at the bottom, so the dofile value IS the duffle module.
local _bootstrap_dir = debug.getinfo(1, "S").source:match("^@?(.*[/\\])") or "./"
local duffle = dofile(_bootstrap_dir .. "../duffle_paths.lua")
local elf_dwarf = require("elf_dwarf")
-- ════════════════════════════════════════════════════════════════════════════
-- Constants
-- ════════════════════════════════════════════════════════════════════════════
-- Format version emitted as the first line. Bump + add a migration test if the format changes;
-- the gdb runtime loader rejects mismatches (E2).
local FORMAT_VERSION = 1
-- ════════════════════════════════════════════════════════════════════════════
-- Type declarations
-- ════════════════════════════════════════════════════════════════════════════
--- @class AtomSourceMapCtx
--- @field shared table -- `ctx.shared`
--- @field shared.corpus table -- source-order registry; single writer is build_ctx
--- @field shared.word_counts table
--- @field out_root string -- output root (e.g. "build/gen")
--- @field flags table -- `ctx.flags`; reads `flags.gdb_runtime` + `flags.elf_path`
-- ════════════════════════════════════════════════════════════════════════════
-- Atom-path renderers
-- ════════════════════════════════════════════════════════════════════════════
--- Join word boundaries (from `items`) to per-word call text + source lines (from `word_events`).
--- @param atom table
--- @return table[], integer
local function canonical_word_entries(atom)
local paths = atom.paths or {}
local events = paths.word_events or {}
local word_items = {}
for _, item in ipairs(paths.items or {}) do
if item.kind == "word" then word_items[#word_items + 1] = item end
end
local entries = {}
for index, event in ipairs(events) do
local item = word_items[index] or {}
entries[#entries + 1] = {
pos = event.i or (index - 1),
line = event.call_line or item.line or 0,
text = event.call_text or item.call_text or "",
body_line = event.body_line or item.body_line or item.line or 0,
invocation = (event.outermost_invocation_id
and paths.invocations
and paths.invocations[event.outermost_invocation_id]) or nil,
}
end
return entries, #events
end
--- Render one atom's provenance stanza. Format 1 line shapes:
--- `WORD N CALL <src-path>:<src-line> MACRO <name> "<def-path>:<def-line>" BODY <line>` (component invocation)
--- `WORD N CALL <src-path>:<src-line> RAW` (raw `.word` outside any mac_* component)
--- Component identity comes from the outermost invocation record; the count-table lookup confirms the component was declared in `corpus.word_counts`
--- (populated by word_count_eval + components passes).
--- @param src table
--- @param atom table
--- @param wc table -- identity alias of corpus.word_counts
--- @return string[], integer
local function emit_provenance_stanza(src, atom, wc)
local lines = {}
local rel_path = src.path:gsub("\\\\", "/")
local entries, total = canonical_word_entries(atom)
lines[#lines + 1] = string.format('ATOM %s "%s" 0', atom.raw_name or atom.name, rel_path)
for _, entry in ipairs(entries) do
local inv = entry.invocation
local macro_count = inv and wc["mac_" .. inv.component_name]
if inv and macro_count ~= nil then
lines[#lines + 1] = string.format('WORD %d CALL %s:%d MACRO %s "%s:%d" BODY %d'
, entry.pos, rel_path, entry.line, inv.component_name
, inv.def_path or "", inv.def_line or 0, entry.body_line)
else
lines[#lines + 1] = string.format("WORD %d CALL %s:%d RAW", entry.pos, rel_path, entry.line)
end
end
lines[1] = lines[1]:gsub(" 0$", " " .. tostring(total))
lines[#lines + 1] = "ENDATOM"
return lines, total
end
--- Render the full provenance file content for one source.
--- @param src table
--- @param wc table
--- @return string
local function render_provenance(src, wc)
local lines = {}
lines[#lines + 1] = "# FORMAT_VERSION 1"
lines[#lines + 1] = "# auto-generated by ps1_meta.lua (passes/atoms_source_map.lua) — DO NOT EDIT"
lines[#lines + 1] = "# Per-.word provenance: maps each emitted .word to its call site (atom body"
lines[#lines + 1] = "# file:line) and, when the word was emitted by a `mac_X(...)` component invocation,"
lines[#lines + 1] = "# the component's definition file:line + the per-word BODY line. Used by"
lines[#lines + 1] = "# dwarf_injection to synthesize DW_TAG_inlined_subroutine instances + per-word"
lines[#lines + 1] = "# line program rows for native source-level step into component bodies."
local function append(atom)
local stanza = emit_provenance_stanza(src, atom, wc)
for _, line in ipairs(stanza) do lines[#lines + 1] = line end
end
for _, atom in ipairs(src.scan.atoms or {}) do
if atom.paths then append(atom) end
end
for _, atom in ipairs(src.scan.raw_atoms or {}) do
if atom.paths then append(atom) end
end
return table.concat(lines, "\n") .. "\n"
end
--- Render one atom's stanza for the sourcemap.txt form (ATOM header line, N WORD lines, ENDATOM marker).
--- Returns (lines, total_words).
--- @param src table
--- @param atom table
--- @param wc table
--- @return string[], integer
local function emit_atom_stanza(src, atom)
local lines = {}
local rel_path = src.path:gsub("\\\\", "/")
local entries, total = canonical_word_entries(atom)
lines[#lines + 1] = string.format('ATOM %s "%s" 0', atom.raw_name or atom.name, rel_path)
for _, entry in ipairs(entries) do
lines[#lines + 1] = string.format("WORD %d LINE %d TEXT %s",
entry.pos, entry.line, entry.text)
end
lines[1] = lines[1]:gsub(" 0$", " " .. tostring(total))
lines[#lines + 1] = "ENDATOM"
return lines, total
end
--- Render the full source map file content for one source (one .atoms.sourcemap.txt per source). Mirrors offsets.lua's
--- `project_atoms` shape: scan.atoms + scan.raw_atoms, no kind filter.
--- @param src table
--- @param wc table
--- @return string
local function render_source_map(src)
local lines = {}
lines[#lines + 1] = "# FORMAT_VERSION " .. FORMAT_VERSION
lines[#lines + 1] = "# auto-generated by ps1_meta.lua (passes/atoms_source_map.lua) — DO NOT EDIT"
local function append(atom)
local stanza = emit_atom_stanza(src, atom)
for _, line in ipairs(stanza) do lines[#lines + 1] = line end
end
for _, atom in ipairs(src.scan.atoms or {}) do
if atom.paths then append(atom) end
end
for _, atom in ipairs(src.scan.raw_atoms or {}) do
if atom.paths then append(atom) end
end
return table.concat(lines, "\n") .. "\n"
end
-- ════════════════════════════════════════════════════════════════════════════
-- gdb-runtime emission (post-link, addresses via nm)
-- ════════════════════════════════════════════════════════════════════════════
--- Escape a string for embedding in a gdb `set $var = "..."` literal.
--- gdb uses C-style escaping; we escape `\` and `"` (newlines were flattened earlier).
--- @param s string
--- @return string
local function gdb_escape(s)
return (s:gsub("\\", "\\\\"):gsub('"', '\\"'))
end
--- Build the list of atoms with addresses + word entries. Shared helper for the gdb-runtime file emission.
--- @param ctx PassCtx
--- @return table[] -- list of {idx, name, src_path, file_base, addr, size_bytes, words, entries}
local function build_atom_table(ctx)
local addrs = elf_dwarf.read_nm(ctx.flags.elf_path)
local corpus = ctx.shared and ctx.shared.corpus
local matched = {}
for _, src in ipairs(corpus.source_order or {}) do
local file_base = src.path:match("([^/\\\\]+)$") or src.path
local function append(atom)
if not atom.paths then return end
local name = atom.raw_name or atom.name
local info = addrs[name]
if not info then return end
local entries, total = canonical_word_entries(atom)
matched[#matched + 1] = {
name = name,
src_path = src.path,
file_base = file_base,
addr = info[1],
size_bytes = info[2],
words = total,
entries = entries,
}
end
for _, atom in ipairs((src.scan or {}).atoms or {}) do append(atom) end
for _, atom in ipairs((src.scan or {}).raw_atoms or {}) do append(atom) end
end
-- Deterministic order: sort by address (matches `nm` output ordering).
table.sort(matched, function(a, b) return a.addr < b.addr end)
for i, a in ipairs(matched) do a.idx = i - 1 end
return matched
end
--- Append the 9 gdb command definitions to `lines`. Pure gdb scripting — addresses come from `nm`,
--- the convenience vars set in `emit_gdb_runtime` provide printf args, and
--- each command is a static sequence of `printf` / `tbreak` / `if ... end` blocks.
--- The Lua pass emits N atoms' worth of lines; runtime iteration is gdb's job.
---
--- Why hardcoded per-atom: gdb's `$` substitution doesn't concat inside var names — `$__atom_name_$__i` in a `while`
--- loop resolves to one literal identifier, not `name_i`. Compile-time emission is the only path.
--- @param lines table -- output line buffer (mutated in place)
--- @param matched table -- list of atom records from `build_atom_table`
local function append_gdb_commands(lines, matched)
-- ── tape_atoms ──
-- Hardcoded one printf per atom. No loop.
lines[#lines + 1] = "define tape_atoms"
for _, a in ipairs(matched) do
-- gdb 12.1 quirk: literals in printf args require an attached target.
-- Use the per-atom convenience vars set above as printf args.
lines[#lines + 1] = string.format(' printf " code_%%-32s @ 0x%%08x %%4d words\\n", $__atom_name_%d, $__atom_addr_%d, $__atom_words_%d',
a.idx, a.idx, a.idx)
end
lines[#lines + 1] = "end"
lines[#lines + 1] = "document tape_atoms"
lines[#lines + 1] = " List every tape atom symbol in the loaded ELF (code_<name>) with .rodata addr + word count."
lines[#lines + 1] = "end"
lines[#lines + 1] = ""
-- ── break_atom (generic) + per-atom break_atom_X ──
lines[#lines + 1] = "define break_atom"
lines[#lines + 1] = ' echo "Usage: break_atom_<exact_name> (pick from the list below)"'
for _, a in ipairs(matched) do
lines[#lines + 1] = string.format(' printf " break_atom_%%-32s\\n", $__atom_name_%d', a.idx)
end
lines[#lines + 1] = "end"
lines[#lines + 1] = "document break_atom"
lines[#lines + 1] = " Generic help: lists the per-atom break_atom_<name> commands."
lines[#lines + 1] = "end"
lines[#lines + 1] = ""
for _, a in ipairs(matched) do
lines[#lines + 1] = string.format("define break_atom_%s", a.name)
lines[#lines + 1] = string.format(" break *$__atom_addr_%d", a.idx)
lines[#lines + 1] = string.format(' printf " Breakpoint set at code_%s (0x%%08x)\\n", $__atom_addr_%d', a.name, a.idx)
lines[#lines + 1] = "end"
lines[#lines + 1] = string.format("document break_atom_%s", a.name)
lines[#lines + 1] = string.format(" Set a breakpoint at code_%s.", a.name)
lines[#lines + 1] = "end"
lines[#lines + 1] = ""
end
-- ── step_atom / next_atom ──
-- Hardcoded one tbreak per atom. No loop.
lines[#lines + 1] = "define step_atom"
for _, a in ipairs(matched) do
lines[#lines + 1] = string.format(" tbreak *$__atom_addr_%d", a.idx)
end
lines[#lines + 1] = " continue"
lines[#lines + 1] = "end"
lines[#lines + 1] = "document step_atom"
lines[#lines + 1] = " Set one-shot BPs at every atom + continue. Stops at the next atom boundary."
lines[#lines + 1] = "end"
lines[#lines + 1] = ""
lines[#lines + 1] = "define next_atom"
lines[#lines + 1] = " step_atom"
lines[#lines + 1] = "end"
lines[#lines + 1] = "document next_atom"
lines[#lines + 1] = " Alias for step_atom."
lines[#lines + 1] = "end"
lines[#lines + 1] = ""
-- ── where_in_atom ──
-- Hardcoded one outer-if per atom; inside, one inner-if per WORD entry.
lines[#lines + 1] = "define where_in_atom"
lines[#lines + 1] = " set $__pc = (unsigned int)$pc"
lines[#lines + 1] = " set $__matched = 0"
for _, a in ipairs(matched) do
-- Precompute end_addr (gdb 12.1's expression evaluator chokes on `addr + words*4`).
lines[#lines + 1] = string.format(" set $__end_%d = $__atom_addr_%d + $__atom_words_%d * 4", a.idx, a.idx, a.idx)
lines[#lines + 1] = string.format(" if $__pc >= $__atom_addr_%d && $__pc < $__end_%d", a.idx, a.idx)
lines[#lines + 1] = string.format(' printf "atom: code_%%s\\n", $__atom_name_%d', a.idx)
lines[#lines + 1] = ' printf "addr: 0x%08x\\n", $__pc'
lines[#lines + 1] = string.format(" set $__word = ($__pc - $__atom_addr_%d) / 4", a.idx)
lines[#lines + 1] = string.format(' printf "word: %%d/%%d\\n", $__word, $__atom_words_%d', a.idx)
-- One inner-if per WORD entry. Each word's line + text hardcoded.
for _, we in ipairs(a.entries) do
lines[#lines + 1] = string.format(" if $__word == %d", we.pos)
-- Escape TEXT for printf format string.
local escaped_text = we.text:gsub("%%", "%%%%"):gsub('"', '\\"')
lines[#lines + 1] = string.format(' printf "source: %%s:%%d %%s\\n", $__atom_file_%d, %d, "%s"', a.idx, we.line, escaped_text)
lines[#lines + 1] = " end"
end
-- Fallback for words beyond the source map (shouldn't happen if nm matches).
local max_word = 0
if #a.entries > 0 then max_word = a.entries[#a.entries].pos end
lines[#lines + 1] = string.format(' if $__word > %d', max_word)
lines[#lines + 1] = ' printf "source: (no source-map entry for word %%d; map may be stale)\\n", $__word'
lines[#lines + 1] = " end"
lines[#lines + 1] = " set $__matched = 1"
lines[#lines + 1] = " end"
end
lines[#lines + 1] = " if !$__matched"
lines[#lines + 1] = ' echo PC is not inside any known atom (in .text or unmapped region).'
lines[#lines + 1] = " end"
lines[#lines + 1] = "end"
lines[#lines + 1] = "document where_in_atom"
lines[#lines + 1] = " Report current atom name, .rodata addr, word offset, and source line."
lines[#lines + 1] = "end"
lines[#lines + 1] = ""
-- ── stepi_inside_atom ──
-- Hardcoded one if-containment-check per atom (no loop).
-- Precompute end_addr in Lua so we don't ask gdb to evaluate `addr + words*4` inside the if condition
-- (gdb 12.1's expression evaluator chokes on the `*` and emits a misleading 'function malloc' error in some gdb builds).
lines[#lines + 1] = "define stepi_inside_atom"
lines[#lines + 1] = " set $__in_atom = 0"
lines[#lines + 1] = " set $__did_step = 0"
lines[#lines + 1] = " set $__pc = (unsigned int)$pc"
for _, a in ipairs(matched) do
-- Precompute end_addr in the convenience var (single expression gdb handles).
lines[#lines + 1] = string.format(" set $__end_%d = $__atom_addr_%d + $__atom_words_%d * 4", a.idx, a.idx, a.idx)
lines[#lines + 1] = string.format(" if $__pc >= $__atom_addr_%d && $__pc < $__end_%d", a.idx, a.idx)
lines[#lines + 1] = " set $__in_atom = 1"
lines[#lines + 1] = " stepi"
lines[#lines + 1] = " set $__did_step = 1"
lines[#lines + 1] = " end"
end
lines[#lines + 1] = " if !$__did_step"
lines[#lines + 1] = ' echo [gdb_tape_atoms] stepi_inside_atom: PC is not inside any atom; refusing to step.'
lines[#lines + 1] = " end"
lines[#lines + 1] = " where_in_atom"
lines[#lines + 1] = "end"
lines[#lines + 1] = "document stepi_inside_atom"
lines[#lines + 1] = " One MIPS-instruction step, then where_in_atom. The step-and-see-source-line workflow."
lines[#lines + 1] = "end"
lines[#lines + 1] = ""
-- ── wave_ctx ──
lines[#lines + 1] = "define wave_ctx"
lines[#lines + 1] = ' printf "$t4 = R_FaceCursor 0x%08x\\n", $t4'
lines[#lines + 1] = ' printf "$t5 = R_VertBase 0x%08x\\n", $t5'
lines[#lines + 1] = ' printf "$t6 = R_OtBase 0x%08x\\n", $t6'
lines[#lines + 1] = ' printf "$t7 = R_PrimCursor 0x%08x\\n", $t7'
lines[#lines + 1] = "end"
lines[#lines + 1] = "document wave_ctx"
lines[#lines + 1] = " Pretty-print the 4 wave-context GPRs ($t4=R_FaceCursor, $t5=R_VertBase, $t6=R_OtBase, $t7=R_PrimCursor). Requires target attached."
lines[#lines + 1] = "end"
end
--- Emit the gdb-runtime file (post-link). Pure gdb scripting — addresses come from `mipsel-none-elf-nm -S`, get embedded
--- in `<ctx.out_root>/gdb_tape_atoms_runtime.gdb`, and load via `set $var = ...` + `define ... end` blocks at gdb source-time.
--- @param ctx PassCtx
local function emit_gdb_runtime(ctx)
if not (ctx.flags and ctx.flags.gdb_runtime) then return end
local elf_path = ctx.flags.elf_path
if not elf_path or elf_path == "" then
io.stderr:write("[atoms_source_map] --gdb-runtime requires --elf <elf>\n")
return
end
if lfs.attributes(elf_path, "mode") ~= "file" then
io.stderr:write(string.format(
"[atoms_source_map] --gdb-runtime: ELF not found at %s\n", elf_path))
return
end
local matched = build_atom_table(ctx)
if #matched == 0 then
io.stderr:write("[atoms_source_map] --gdb-runtime: no atoms matched against nm symbols (stale scan?).\n")
return
end
local lines = {}
lines[#lines + 1] = "# Auto-generated by ps1_meta.lua (passes/atoms_source_map.lua)"
lines[#lines + 1] = "# DO NOT EDIT — re-run ps1_meta.lua --atoms-source-map --gdb-runtime to regenerate"
lines[#lines + 1] = "# Sourced by scripts/gdb/gdb_tape_atoms.gdb (the wrapper)."
lines[#lines + 1] = "# Pure gdb scripting — no Python, no Tcl, no Guile required."
lines[#lines + 1] = "# Commands are FULLY HARDCODED per-atom because gdb doesn't do nested"
lines[#lines + 1] = "# `$` substitution in var names (`$foo_$i` is one literal identifier)."
lines[#lines + 1] = "# Per-atom convenience vars ($__atom_name_<i> etc.) are set so gdb's"
lines[#lines + 1] = "# `printf` has valid expression args (gdb 12.1 quirks: literals in"
lines[#lines + 1] = "# printf args require an attached target; convenience-var args do not)."
lines[#lines + 1] = string.format("# %d atoms from ELF: %s", #matched, elf_path)
lines[#lines + 1] = ""
-- Format version + count + ELF path (the latter is referenced by the load-line).
lines[#lines + 1] = "set $__atom_format_version = " .. FORMAT_VERSION
lines[#lines + 1] = string.format("set $__atom_count = %d", #matched)
lines[#lines + 1] = string.format('set $__elf_path = "%s"', gdb_escape(elf_path))
lines[#lines + 1] = ""
-- Per-atom convenience vars (used as printf args; literals aren't accepted
-- without an attached target on gdb 12.1).
for _, a in ipairs(matched) do
lines[#lines + 1] = string.format('set $__atom_name_%d = "%s"', a.idx, gdb_escape(a.name))
lines[#lines + 1] = string.format("set $__atom_addr_%d = 0x%x", a.idx, a.addr)
lines[#lines + 1] = string.format("set $__atom_words_%d = %d", a.idx, a.words)
lines[#lines + 1] = string.format('set $__atom_file_%d = "%s"', a.idx, gdb_escape(a.file_base))
end
lines[#lines + 1] = ""
-- The 9 commands (each `define ... end` overrides the wrapper's stub).
lines[#lines + 1] = "# ── 9 user commands (overrides wrapper stubs) ──"
append_gdb_commands(lines, matched)
lines[#lines + 1] = ""
-- Confirmation line for the source operator.
lines[#lines + 1] = 'printf "[gdb_tape_atoms] runtime loaded %d atoms from %s\\n", $__atom_count, $__elf_path'
local out_path
-- Move out of `<out_root>/gdb_tape_atoms_runtime.gdb` to `<out_root>/../gdb_tape_atoms_runtime.gdb` when the conventional `<out_root>` is `<build>/gen`
-- (any equivalent spelling — relative, absolute backslash, absolute forward-slash, trailing-separator variants).
-- This puts the gdb runtime alongside the ELF at `build/` rather than under the report subdir.
local function ends_with_gen_dir(p)
if type(p) ~= "string" then return false end
return p:match("[/\\]gen[/\\]?$") ~= nil or p == "build/gen" or p == "build\\gen"
end
if ends_with_gen_dir(ctx.out_root) then
-- Strip the trailing `/gen` segment, then write the runtime script under `build/`.
-- e.g. "C:/projects/Pikuma/ps1/build/gen" -> "C:/projects/Pikuma/ps1/build".
local parent = ctx.out_root:gsub("[/\\]gen[/\\]?$", "")
out_path = parent .. "/gdb_tape_atoms_runtime.gdb"
else
out_path = ctx.out_root .. "/gdb_tape_atoms_runtime.gdb"
end
duffle.ensure_dir(duffle.dirname(out_path))
duffle.write_file_lf(out_path, table.concat(lines, "\n") .. "\n")
-- io.stderr:write(string.format("[atoms_source_map] wrote %s (%d atoms)\n", out_path, #matched))
end
-- ════════════════════════════════════════════════════════════════════════════
-- M — module exports
-- ════════════════════════════════════════════════════════════════════════════
local M = {}
-- Expose the pure render functions so `report.lua` and the focused tests can call them directly without triggering the file-emit path.
M.render_source_map = render_source_map
M.render_provenance = render_provenance
--- Render ONE atom's sourcemap stanza.
--- @param atom table -- atom record (must have `atom.paths` populated)
--- @return string
function M.render_atom_source_map(atom)
assert(type(atom) == "table", "render_atom_source_map: atom must be a table")
assert(type(atom.paths) == "table", "render_atom_source_map: atom.paths must be a table")
local entries, total = canonical_word_entries(atom)
local lines = {}
lines[#lines + 1] = string.format("ATOM %s %d", (atom.raw_name or atom.name), total)
for _, entry in ipairs(entries) do
lines[#lines + 1] = string.format("WORD %d LINE %d TEXT %s",
entry.pos, entry.line, entry.text)
end
lines[#lines + 1] = "ENDATOM"
return table.concat(lines, "\n") .. "\n"
end
--- Render ONE atom's provenance stanza — no per-file format header, no enumeration of other atoms.
---
--- `rel_path` is the source path (forward-slashes) embedded in every `CALL` line.
--- The .md caller (report.lua) is expected to derive this once per `## <source>` heading and pass it down for each atom in that source.
--- @param atom table -- atom record (must have `atom.paths` populated)
--- @param wc table -- identity alias of `corpus.word_counts`
--- @param rel_path string -- source path (forward-slashes) for `CALL` fields
--- @return string
function M.render_atom_provenance(atom, wc, rel_path)
assert(type(atom) == "table", "render_atom_provenance: atom must be a table")
assert(type(atom.paths) == "table", "render_atom_provenance: atom.paths must be a table")
assert(type(rel_path) == "string", "render_atom_provenance: rel_path must be a string")
local entries, total = canonical_word_entries(atom)
local lines = {}
lines[#lines + 1] = string.format("ATOM %s %d", (atom.raw_name or atom.name), total)
for _, entry in ipairs(entries) do
local inv = entry.invocation
local macro_count = inv and wc and wc["mac_" .. inv.component_name]
if inv and macro_count ~= nil then
lines[#lines + 1] = string.format(
'WORD %d CALL %s:%d MACRO %s "%s:%d" BODY %d',
entry.pos, rel_path, entry.line, inv.component_name,
inv.def_path or "", inv.def_line or 0, entry.body_line)
else
lines[#lines + 1] = string.format(
"WORD %d CALL %s:%d RAW", entry.pos, rel_path, entry.line)
end
end
return table.concat(lines, "\n") .. "\n"
end
--- Pass entry. For each source that declares at least one `MipsAtom_(name)` / `MipsCode code_<name>`,
--- emit two files in `<out_root>/`: `<basename>.atoms.sourcemap.txt` (per-word call-site map) and `<basename>.atoms.provenance.txt`
--- (per-word definition + body line, resolved via the outermost `mac_X(...)` invocation).
--- When `ctx.flags.gdb_runtime` is true and `ctx.flags.elf_path` exists, also emit the post-link gdb script `<ctx.out_root>/gdb_tape_atoms_runtime.gdb`.
--- @param ctx PassCtx
--- @return PassResult
function M.run(ctx)
local outputs = {}
local errors = {}
local warnings = {}
local corpus = ctx.shared and ctx.shared.corpus
if type(corpus) ~= "table" or type(corpus.source_order) ~= "table" then
error("atoms_source_map.run requires ctx.shared.corpus.source_order (canonical corpus).", 0)
end
-- Word counts come from `corpus.word_counts` (populated by word_count_eval + components passes).
local wc = corpus.word_counts or {}
if not next(wc) then
warnings[#warnings + 1] = {
line = 0,
msg = "atoms_source_map: corpus.word_counts is empty; the word-counts + components passes may not have populated it. Check the PASSES dep edges.",
}
end
-- atoms.sourcemap.txt + atoms.provenance.txt content moved to report.lua via `<module>.atoms.md` markdown file.
-- This pass emits only the post-link gdb_runtime artifact (see emit_gdb_runtime below).
-- Optionally emit the gdb-runtime form (post-link, one file per build).
if ctx.flags and ctx.flags.gdb_runtime then
emit_gdb_runtime(ctx)
end
return { outputs = outputs, errors = errors, warnings = warnings }
end
return M
+786
View File
@@ -0,0 +1,786 @@
--- passes/components.lua — Component-macro header generator.
---
--- Ownership: `corpus.word_counts`, `corpus.components`, and `corpus.component_body_index`.
--- Scanner owns `declaration_comment` and `debug_skip` on each declaration record; this pass projects both forward.
---
--- Reads the pre-scanned SourceScan payload from `duffle.scan_source` for `MipsAtomComp_(ac_X)` and `MipsAtomComp_Proc_(ac_X, { body })` declarations,
--- then resolves the function-args string from the preceding `FI_ Slice_MipsCode ac_X(...)` declaration via a backward walk.
---
--- Emits one `gen/macs.h` per *immediate source directory* with `#define mac_X(sig) \` macros plus `WORD_COUNT(mac_X, N)` entries for downstream offset computation.
--- All sources inside the same directory contribute to the same file (per-directory aggregation).
--- The directory itself is the namespace, so the filename does not repeat the module name.
-- ════════════════════════════════════════════════════════════════════════════
-- Module-scope requires + package.path setup
-- ════════════════════════════════════════════════════════════════════════════
-- Bootstrap: same as entry scripts. See `ps1_meta.lua` for the rationale.
-- Bootstrap: load `scripts/duffle_paths.lua` (sets package.path + package.cpath).
-- Uses `debug.getinfo` to find this file's own directory, so it works both standalone and when require'd from the orchestrator.
-- Bootstrap: load `duffle_paths.lua` via `debug.getinfo(1, "S").source` (works both standalone + when require'd).
-- duffle_paths.lua sets package.path then returns `require("duffle")` at the bottom, so the dofile value IS the duffle module.
local _bootstrap_dir = debug.getinfo(1, "S").source:match("^@?(.*[/\\])") or "./"
local duffle = dofile(_bootstrap_dir .. "../duffle_paths.lua")
-- ════════════════════════════════════════════════════════════════════════════
-- Constants
-- ════════════════════════════════════════════════════════════════════════════
-- Atom component declaration identifiers.
local ATOM_COMP_PROC = "MipsAtomComp_Proc_"
local MIPS_ATOM = "Slice_MipsCode" -- prefix on the function declaration that wraps an AtomComp_Proc_
-- Component-name prefixes.
local AC_PREFIX = "ac_" -- arg to MipsAtomComp_(ac_X); the X is the atom name
local AC_PREFIX_LEN = 3
local MAC_PREFIX = "mac_" -- prefix on generated macros; the rest is the atom name
local MAC_PREFIX_LEN = 4
-- ASCII byte values used in tokenization.
local BYTE_NEWLINE = 10
local BYTE_SLASH = 47
-- Output gen subdirectory + filename (per-directory aggregation; the directory name is the namespace).
local GEN_SUBDIR = "gen"
local MACS_FILENAME = "macs.h"
-- ════════════════════════════════════════════════════════════════════════════
-- Type declarations
-- ════════════════════════════════════════════════════════════════════════════
--- @class SourceFile
--- @field path string -- Absolute path to the source file
--- @field text string -- Full source text
--- @field dir string -- Directory containing the source
--- @field basename string -- Filename without extension
--- @field scan table -- Pre-scanned SourceScan payload (from duffle.scan_source)
--- @class PassCtx
--- @field sources SourceFile[] -- All source files in the build
--- @field metadata_path string -- Path to word_count.metadata.h
--- @field shared table -- Cross-pass shared state
--- @field out_root string -- Output root (e.g. "build/gen")
--- @field project_root string -- Project root (e.g. "code/")
--- @field upstream table<string, table> -- Per-pass upstream outputs
--- @field flags table -- CLI flags
--- @field verbose boolean -- Log diagnostic info
--- @class PassResult
--- @field outputs table[] -- {kind=, path=} entries describing emit files
--- @field errors table[] -- {line=, msg=} entries; build-stops
--- @field warnings table[] -- {line=, msg=} entries; build-succeeds
--- @class Component
--- @field name string -- Atom name (without `ac_` prefix)
--- @field body string -- Brace-delimited body (without the braces)
--- @field args string|nil -- Function-args string (function form only)
--- @field line integer -- Source line of the declaration
--- @field comment string|nil -- Scanner-owned `declaration_comment`; the components pass reads it from the scanner record
--- @field kind string -- "comp_bare" | "comp_proc"
--- @field debug_skip boolean -- Mirror of `a.debug_skip` (scanner-owned); true iff a bare `atom_dbg_skip` marker immediately preceded the declaration
-- ════════════════════════════════════════════════════════════════════════════
-- Local helpers (file I/O + path normalization)
-- ════════════════════════════════════════════════════════════════════════════
local M = {}
-- ════════════════════════════════════════════════════════════════════════════
-- Back-walk helpers (composed into the entry point below: find_function_args_for)
--
-- Only the function-args lookup for proc components occurs here.
-- The preceding-comment walk occur in `scan_source.lua` — `a.declaration_comment` carries the resolved comment,
-- so this file reads it forward rather than re-walking the source.
-- ════════════════════════════════════════════════════════════════════════════
--- Find the args of the function declaration that immediately precedes a `MipsAtomComp_Proc_` invocation of the given name.
--- Returns the args string (e.g., `"U4 off, U4 code, U1 r, U1 g, U1 b"`) or nil if no function declaration is found.
---
--- Convention: function form is
--- `FI_ Slice_MipsCode ac_X(args) MipsAtomComp_Proc_(ac_X, { body })`
--- We find the LAST occurrence of `"ac_X("` before `before_pos` and extract the args from inside the parens.
--- We then verify the preceding context ends with `Slice_MipsCode`
--- (the function-decl keyword with possible qualifiers between).
---
--- @param source string
--- @param name string
--- @param before_pos integer
--- @return string|nil
local function find_function_args_for(source, name, before_pos)
-- Find the LAST occurrence of `name + "("` in `source[1..before_pos]`.
local name_open = name .. "("
local last_idx = nil
local scan_pos = 1
while true do
-- Pass `before_pos + 1` so string.find only returns positions < before_pos + 1
-- (string.find's 4th arg `plain` is true; we use the 3rd arg `init` for the upper bound).
local found = source:find(name_open, scan_pos, true)
if not found or found >= before_pos then break end
last_idx = found
scan_pos = found + #name_open
end
if not last_idx then return nil end
-- Verify the preceding context ends with "MipsAtom" (with possible qualifiers between).
local before = source:sub(1, last_idx - 1)
local trimmed = duffle.trim(before)
if trimmed:sub(-#MIPS_ATOM) ~= MIPS_ATOM then
-- Preceding context is not a function declaration.
return nil
end
local open_paren = last_idx + #name -- position of "("
-- scan: MipsAtom ac_X(
local inner = duffle.read_parens(source, open_paren)
-- scan: MipsAtom ac_X(<args>)
if not inner then return nil end
return inner
end
-- ════════════════════════════════════════════════════════════════════════════
-- Argument-name extraction
-- ════════════════════════════════════════════════════════════════════════════
--- Extract just the parameter NAMES from a function-args string (stripping type annotations). E.g.,
--- `"U4 off, U4 code, U1 r, U1 g, U1 b"` -> `{"off", "code", "r", "g", "b"}`
--- `"U4 *ptr"` -> `{"ptr"}`
--- `""` -> nil
--- @param args_str string|nil
--- @return string[]|nil
local function extract_arg_names(args_str)
if not args_str or args_str == "" then return nil end
local names = {}
local tokens = duffle.split_top_level_commas(args_str)
for _, tok in ipairs(tokens) do
local trimmed = duffle.trim(tok)
if trimmed ~= "" then
-- Find the identifier at the end: walk back over trailers (whitespace + `*` + `[]`),
-- then walk back over the identifier chars (alnum + `_`).
local ident_end = #trimmed
while ident_end > 0 do
local ch = trimmed:sub(ident_end, ident_end)
if ch == " " or ch == "\t" or ch == "*" or ch == "]" or ch == "[" then
ident_end = ident_end - 1
else
break
end
end
local ident_start = ident_end
while ident_start > 0 do
local ch = trimmed:sub(ident_start, ident_start)
if duffle.is_alnum_byte(string.byte(ch)) or ch == "_" then
ident_start = ident_start - 1
else
break
end
end
ident_start = ident_start + 1
local name = trimmed:sub(ident_start, ident_end)
if name ~= "" then names[#names + 1] = name end
end
end
if #names == 0 then return nil end
return names
end
-- ════════════════════════════════════════════════════════════════════════════
-- Component projection (read from pre-scanned SourceScan)
-- ════════════════════════════════════════════════════════════════════════════
--- Project pre-scanned MipsAtomComp_ / MipsAtomComp_Proc_ entries into Component shape.
--- Reads the scanner-owned `declaration_comment` (resolved by scan_source.lua, skipping backward across an associated bare `atom_dbg_skip` marker when present).
--- Per-source backward lookups remain in place only for the function `args` of proc components.
--- That lookup is unique to components.lua and stays separate from the declaration-comment walk.
--- Carries `body_tokens` forward from scan-source so word_count_rec reads from the precomputed table instead of calling duffle.tokenize_body again.
--- Carries the scanner-owned `debug_skip` flag forward so the generated projection can emit `/* atom_dbg_skip */`
--- before the authored comment and so `update_canonical_components` can mirror the same field onto `corpus.components[name]`.
--- @param source string -- the full source text (needed for backward lookups)
--- @param scan table -- SourceScan from duffle.scan_source
--- @return Component[]
local function project_components(source, scan)
local out = {}
for _, a in ipairs(scan.atoms) do
if a.kind == "comp_bare" or a.kind == "comp_proc" then
local args = find_function_args_for(source, a.raw_name, a.ident_pos)
-- Comment ownership: scan_source.lua stamps `declaration_comment` on the record by walking backward past any associated bare marker.
-- The pass reads `declaration_comment` directly.
local comment = a.declaration_comment or ""
out[#out + 1] = {
line = a.line,
name = a.name,
body = a.body,
body_off = a.body_off,
body_tokens = a.body_tokens,
args = args,
comment = comment,
kind = a.kind, -- "comp_bare" | "comp_proc"; provenance emitter reads this.
debug_skip = a.debug_skip == true,
}
end
end
return out
end
-- ════════════════════════════════════════════════════════════════════════════
-- Line-comment → block-comment conversion
-- ════════════════════════════════════════════════════════════════════════════
-- Convert `//` line comments to `/* */` block comments in a token.
-- C macros use `\` line-continuations; a `//` comment before `\` would consume the continuation,
-- breaking the macro. We convert `//` to `/* */` so the multi-line macro structure is preserved.
--
-- Skips `//` sequences that are inside string or character literals
-- (a rough heuristic — sufficient for component bodies which don't have those constructs).
--- @param s string
--- @return string
local function convert_line_comments_to_block(s)
local result = s
local pos = 1
local len = #result
while pos <= len do
local is_double_slash = result:byte(pos) == BYTE_SLASH
and pos + 1 <= len and result:byte(pos + 1) == BYTE_SLASH
if not is_double_slash then
pos = pos + 1
else
-- Find end of line.
local eol = pos
while eol <= len and result:byte(eol) ~= BYTE_NEWLINE do
eol = eol + 1
end
local before = result:sub(1, pos - 1)
local comment = result:sub(pos + 2, eol - 1) -- skip the `//`
local after
if eol <= len and result:byte(eol) == BYTE_NEWLINE then
after = " */" .. result:sub(eol) -- keep the newline
else
after = " */"
end
result = before .. "/*" .. comment .. after
pos = #before + 2 + #comment + 3 -- skip past converted comment
end
end
return result
end
-- ════════════════════════════════════════════════════════════════════════════
-- Word-count computation (memoized recursive lookup)
-- ════════════════════════════════════════════════════════════════════════════
--- Strip the `mac_` prefix from a component-call ident so we can look it up against the components-by-name table.
--- Returns the ident unchanged if it doesn't start with the prefix
--- (so a non-component ident like `mask_upper` falls through to the wc-table branch).
--- @param ident string|nil
--- @return string|nil
local function strip_mac_prefix(ident)
if not ident then return nil end
if ident:sub(1, MAC_PREFIX_LEN) == MAC_PREFIX then
return ident:sub(MAC_PREFIX_LEN + 1)
end
return ident
end
--- (internal) Recursive word-count lookup. `cache` is the memoization table shared across all components
--- in a single source's `count_all_components` pass; the in-progress -1 sentinel detects cycles (A -> B -> A).
--- @param name string -- the component name (without `mac_`)
--- @param comp_by_name table<string, Component>
--- @param wc table<string, integer>
--- @param cache table<string, integer>
--- @return integer
local function word_count_rec(name, comp_by_name, wc, cache)
if cache[name] ~= nil then return cache[name] end
cache[name] = -1 -- mark in-progress (cycle detection)
local cc = comp_by_name[name]
local n
if cc then
n = 0
local tokens = cc.body_tokens
for _, t in ipairs(tokens) do
local trimmed = t.tok
if trimmed ~= "" then
local lookup = strip_mac_prefix(duffle.read_ident(trimmed, 1))
if lookup and comp_by_name[lookup] then
-- It's a `mac_X(...)` call. Recurse.
n = n + word_count_rec(lookup, comp_by_name, wc, cache)
elseif lookup and wc and wc[lookup] then
-- Encoding macro or pseudo-instruction (e.g. mask_upper = 2, nop2 = 2).
n = n + wc[lookup]
else
-- Unrecognized token. Fall back to 1 word.
n = n + 1
end
end
end
else
-- Not a known component: assume 1 word (regular instruction).
n = 1
end
cache[name] = n
return n
end
--- Compute word counts for every component in `components` in a single pass.
--- The name-lookup table + memoization cache are built ONCE (per source) instead of per-component,
--- so the cache survives across siblings and a component's recursive `mac_Y(...)`
--- references hit memoized values instead of re-walking the body.
--- Cycle detection (A -> B -> A) is preserved via the in-progress `-1` sentinel in `cache`.
--- @param components Component[]
--- @param wc table<string, integer>
--- @return table<string, integer> -- map of component name (without `mac_`) -> word count
local function count_all_components(components, wc)
local comp_by_name = {}
for _, cc in ipairs(components) do comp_by_name[cc.name] = cc end
local cache = {}
local counts = {}
for _, c in ipairs(components) do
counts[c.name] = word_count_rec(c.name, comp_by_name, wc, cache)
end
return counts
end
-- ═══════════════════════════════════════════
-- Per-component metadata derivation (replaces the hardcoded `M.GP0_MACRO_CONTRIB` + `M.INSTRUCTION_LATENCY[mac_*]` tables that previously lived in `duffle.lua`).
--
-- Each `MipsAtomComp_(ac_X) { body }` definition in `code/duffle/lottes_tape.h` is the canonical source.
-- The `mac_X(...)` macros are GENERATED from these definitions by `emit_component_macros_h` for tape-side composition;
-- the metaprogram must NEVER walk the generated variants to derive metadata.
-- Always walk the original `MipsAtomComp_` body via `cc.body_tokens`.
-- ═══════════════════════════════════════════
--- (internal) Recursive cycle-cost derivation. Sum `latency[ident]` per emitted instruction in the component body,
--- recursing through nested `mac_*` calls (so `mac_format_g4_color`'s cost = 4 × `mac_pack_color_word`'s cost).
--- Special rule: `mac_yield`'s cost = 0 (per `lottes_tape.h:125-130` "the runtime cost lands in the next atom's prologue").
--- @param name string -- component bare name (e.g. "yield", "pack_color_word")
--- @param comp_by_name table<string, Component>
--- @param latency table<string, integer>
--- @param cache table<string, integer> -- shared memoization; `-1` sentinel detects cycles
--- @return integer
local function cycle_cost_rec(name, comp_by_name, latency, cache)
if cache[name] ~= nil then return cache[name] end
cache[name] = -1
local cc = comp_by_name[name]
local n
if cc then
if name == "yield" then
-- mac_yield's cost is 0 by convention (the runtime cost lands in the next atom's prologue).
n = 0
else
n = 0
local tokens = cc.body_tokens
for _, t in ipairs(tokens) do
local trimmed = t.tok
if trimmed ~= "" then
local ident = duffle.read_ident(trimmed, 1)
if ident and ident:sub(1, MAC_PREFIX_LEN) == MAC_PREFIX then
-- Nested `mac_X(...)` call: recurse.
local nested = ident:sub(MAC_PREFIX_LEN + 1)
n = n + cycle_cost_rec(nested, comp_by_name, latency, cache)
else
-- Leaf instruction or pseudo-macro. Look up in INSTRUCTION_LATENCY; default 1.
n = n + (latency[ident] or 1)
end
end
end
end
else
n = 1
end
cache[name] = n
return n
end
--- (internal) Recursive GP0 prim-buffer contribution. Count `store_word` / `store_half` / `store_byte`
--- calls in the component body that target `R_PrimCursor` (these are the
--- RAM-side prim-buffer words the macro contributes), recursing through nested `mac_*` calls.
--- Only `R_PrimCursor`-targeting stores count. Stores targeting other registers (e.g. `R_OtBase`, heap pointers) are not prim-buffer contributions.
--- @param name string
--- @param comp_by_name table<string, Component>
--- @param cache table<string, integer>
--- @return integer
local function gp0_contrib_rec(name, comp_by_name, cache)
if cache[name] ~= nil then return cache[name] end
cache[name] = -1
local cc = comp_by_name[name]
local n
if cc then
n = 0
local tokens = cc.body_tokens
for _, t in ipairs(tokens) do
local trimmed = t.tok
if trimmed ~= "" then
local ident = duffle.read_ident(trimmed, 1)
if ident and ident:sub(1, MAC_PREFIX_LEN) == MAC_PREFIX then
-- Nested `mac_X(...)` call: recurse.
local nested = ident:sub(MAC_PREFIX_LEN + 1)
n = n + gp0_contrib_rec(nested, comp_by_name, cache)
elseif ident == "store_word" or ident == "store_half" or ident == "store_byte" then
if trimmed:find("R_PrimCursor", 1, true) then
n = n + 1
end
end
end
end
else
n = 0
end
cache[name] = n
return n
end
--- Compute `cycle_cost` + `gp0_contrib` for every component in `components` in a single pass.
--- Memoization cache is built ONCE (per source) and shared across both helpers so that
--- a nested `mac_Y` reference inside a `mac_X` body computes its values once.
--- @param components Component[]
--- @param latency table<string, integer>
--- @return table<string, {cycle_cost=integer, gp0_contrib=integer}>
local function compute_components_metadata(components, latency)
local comp_by_name = {}
for _, cc in ipairs(components) do comp_by_name[cc.name] = cc end
local cc_cache = {}
local gc_cache = {}
local out = {}
for _, c in ipairs(components) do
out[c.name] = {
cycle_cost = cycle_cost_rec(c.name, comp_by_name, latency, cc_cache),
gp0_contrib = gp0_contrib_rec(c.name, comp_by_name, gc_cache),
}
end
return out
end
-- ════════════════════════════════════════════════════════════════════════════
-- Per-component emit logic
-- ════════════════════════════════════════════════════════════════════════════
--- Split a (possibly multi-line) comment into per-line entries.
--- Hand-rolled (no regex patterns used).
--- @param s string
--- @return string[]
local function split_comment_lines(s)
local out = {}
local pos = 1
local s_len = #s
while pos <= s_len do
local nl = s:find("\n", pos, true)
if not nl then
out[#out + 1] = s:sub(pos)
break
end
out[#out + 1] = s:sub(pos, nl - 1)
pos = nl + 1
end
return out
end
--- Determine the macro signature: function-args list (function form) or variadic-ignored (bare form).
--- @param args_str string|nil
--- @return string
local function signature_from_args(args_str)
local arg_names = extract_arg_names(args_str)
if arg_names and #arg_names > 0 then
return table.concat(arg_names, ", ")
end
return "..."
end
--- Strip the trailing `" \"` (space + backslash) line continuation from the last body line.
--- The last 2 chars are always that pair.
local function strip_trailing_continuation(lines)
local last = lines[#lines]
if last:sub(-2) == " \\" then
lines[#lines] = last:sub(1, -3)
end
end
--- Emit the `#define mac_X(sig) \<newline>\t<tok1> \<newline>,\t<tok2> ...` block.
--- Converts `//` line comments to `/* */` block comments in each token so they don't break the C macro `\` line continuations.
local function emit_macro_body(lines, c, sig, tokens)
for tok_idx = 1, #tokens do
tokens[tok_idx] = convert_line_comments_to_block(tokens[tok_idx])
end
lines[#lines + 1] = "#define mac_" .. c.name .. "(" .. sig .. ") \\"
lines[#lines + 1] = "\t" .. tokens[1] .. " \\"
for tok_idx = 2, #tokens do
lines[#lines + 1] = ",\t" .. tokens[tok_idx] .. " \\"
end
strip_trailing_continuation(lines)
end
--- Build the list of lines for one component
--- (signature comment, `#define mac_X(...)` line with backslash-continued tokens, then `WORD_COUNT(mac_X, N)` entry).
--- For skipped components, a `/* atom_dbg_skip */` marker comment is emitted immediately before the authored comment block.
--- The marker is a single line, the comment comes next, and the `#define` line follows. The `debug_skip` stamp is scanner-owned
--- (`a.debug_skip == true` on the declaration record); the components pass projects it directly.
--- @param c Component
--- @param components Component[]
--- @param wc table<string, integer>
--- @return string[] -- list of lines for this component
local function build_component_lines(c, counts)
local lines = {}
-- Marker comment: emitted once for every skipped component.
-- The marker is scanner-owned (declared by `atom_dbg_skip` immediately before the declaration in the source);
-- the components pass projects `c.debug_skip` and emits the marker as a generated comment.
if c.debug_skip then
lines[#lines + 1] = "/* atom_dbg_skip */"
end
if c.comment and c.comment ~= "" then
for _, line in ipairs(split_comment_lines(c.comment)) do
lines[#lines + 1] = line
end
end
local tokens = duffle.split_top_level_commas(c.body)
for i = 1, #tokens do tokens[i] = duffle.trim(tokens[i]) end
local sig = signature_from_args(c.args)
-- Direct lookup against the per-source precomputed `counts` table (built once by count_all_components).
local n = counts[c.name]
if n > 0 then
emit_macro_body(lines, c, sig, tokens)
end
-- Emit the WORD_COUNT(mac_<X>, N) entry.
lines[#lines + 1] = "WORD_COUNT(mac_" .. c.name .. ", " .. n .. ")"
lines[#lines + 1] = ""
return lines
end
-- ════════════════════════════════════════════════════════════════════════════
-- Per-source emit logic
-- ════════════════════════════════════════════════════════════════════════════
--- Build the boilerplate header lines (the `#ifdef INTELLISENSE_DIRECTIVES` block,
--- the `// Auto-generated` comment, the `// Source:` line, and the self-contained `WORD_COUNT` macro definition).
--- @param dir string -- the absolute source directory
--- @param sources SourceFile[] -- sources contributing to this directory (for the header comment)
--- @return string[]
local function header_boilerplate(dir, sources)
local source_lines = { "// Directory: " .. duffle.to_absolute_path(dir) .. "/" }
for _, src in ipairs(sources) do
source_lines[#source_lines + 1] = "// source: " .. duffle.to_absolute_path(src.path)
end
local source_blob = table.concat(source_lines, "\n")
return {
-- #pragma once wrapped in #ifdef INTELLISENSE_DIRECTIVES, matching the convention in lottes_tape.h.
-- The build does manual unity includes (the user controls include order), so the pragma is only active for IDE/tooling.
"#ifdef INTELLISENSE_DIRECTIVES",
"#pragma once",
"#endif",
"// Auto-generated by ps1_meta.lua — DO NOT EDIT",
source_blob,
"// Component atoms (MipsAtomComp_(ac_*)) -> macro variants (mac_*)",
"",
-- Self-contained: define WORD_COUNT if not already defined.
-- We use the same definition here so the auto-generated entries below expand
-- to compile-time constants whether the metadata file is included first or not.
"#ifndef WORD_COUNT",
"#define WORD_COUNT(name, count) enum { words_##name = (count) };",
"#endif",
"",
}
end
--- Compute the per-directory output path for `.macs.h`.
--- e.g. any source in `code/duffle/` produces `code/duffle/gen/macs.h` regardless of source filename.
--- The directory name is the namespace; the filename does not repeat it.
--- @param dir string -- the absolute source directory
--- @return string -- the output directory
--- @return string -- the full output path
local function compute_macs_h_path(dir)
local out_dir = dir .. "/" .. GEN_SUBDIR
local out_path = out_dir .. "/" .. MACS_FILENAME
return out_dir, out_path
end
--- Emit a per-directory `.macs.h` header with the aggregated `mac_X` macros + `WORD_COUNT` entries.
--- Writes in BINARY mode so LF line endings are preserved (the git blob is LF; Windows text-mode would emit CRLF and break the byte-identical diff).
--- @param ctx PassCtx
--- @param dir string -- the absolute source directory
--- @param sources SourceFile[] -- sources contributing to this directory (for the header comment)
--- @param components Component[] -- aggregated components from all sources in this directory
--- @param counts table<string, integer> -- precomputed word counts (from count_all_components)
--- @return string|nil -- path to the written file (nil if no components)
local function emit_component_macros_h(ctx, dir, sources, components, counts)
if #components == 0 then return nil end
local out_dir, out_path = compute_macs_h_path(dir)
local lines = header_boilerplate(dir, sources)
for _, c in ipairs(components) do
for _, l in ipairs(build_component_lines(c, counts)) do
lines[#lines + 1] = l
end
end
local content = table.concat(lines, "\n") .. "\n"
duffle.ensure_dir(out_dir)
duffle.write_file_lf(out_path, content)
print(string.format(" -> %s", out_path))
return out_path
end
-- ════════════════════════════════════════════════════════════════════════════
-- Pass entry
-- ════════════════════════════════════════════════════════════════════════════
--- (internal) Extend `corpus.word_counts` with this source's component macros so offsets sees them without re-reading the file.
--- First declaration wins: a later caller's count is dropped (the existing entry from the first source is preserved).
--- @param corpus table -- the corpus
--- @param components Component[]
--- @param counts table<string, integer> -- precomputed word counts (from count_all_components)
local function update_canonical_word_counts(corpus, components, counts)
local wc = corpus.word_counts
for _, c in ipairs(components) do
local key = "mac_" .. c.name
if wc[key] == nil then
wc[key] = counts[c.name]
end
end
end
--- @class ComponentDef
--- @field name string -- bare name (without ac_/mac_ prefix)
--- @field line integer -- definition source line (line of `MipsAtomComp_(ac_X)` / `MipsAtomComp_Proc_(ac_X, ...)`)
--- @field path string -- absolute source path of the definition
--- @field kind string -- "comp_bare" | "comp_proc"
--- @field debug_skip boolean -- mirror of the scanner-owned `a.debug_skip`; consumers read this directly
--- (internal) Populate `corpus.components` with this source's components-by-name map.
--- First declaration wins; later declarations of the same bare name are dropped and recorded as a collision via `corpus.collisions` (kind = "component").
--- The pass does NOT write to `ctx.shared.components`.
--- No parallel skip map is built here; consumers that need the per-component skip state read `corpus.components[name].debug_skip` directly.
--- The `cycle_cost` + `gp0_contrib` fields are populated from `metadata[c.name]` (computed by `compute_components_metadata` against the original `MipsAtomComp_` body).
--- @param corpus table -- the corpus
--- @param src SourceFile
--- @param components Component[]
--- @param metadata table<string, {cycle_cost=integer, gp0_contrib=integer}>
local function update_canonical_components(corpus, src, components, metadata)
local rel_path = src.path:gsub("\\", "/")
for _, c in ipairs(components) do
-- Keyed by bare name (e.g. `yield`, `load_tri_indices`).
-- The atoms_source_map pass looks up components by bare name from the corpus;
-- `mac_` prefix lives at the call-site identifier and is stripped before lookup.
local m = metadata and metadata[c.name] or nil
if corpus.components[c.name] == nil then
corpus.components[c.name] = {
name = c.name,
line = c.line,
path = rel_path,
kind = c.kind or "comp_bare",
debug_skip = c.debug_skip == true,
cycle_cost = m and m.cycle_cost or nil,
gp0_contrib = m and m.gp0_contrib or nil,
}
else
-- A second declaration of the same bare name: record a typed collision so static-analysis + the report can surface it.
-- Identical-shape declarations (same path + line) reuse the first-wins entry without a collision record.
local existing = corpus.components[c.name]
if existing.path ~= rel_path or existing.line ~= c.line then
local kind = c.kind or "comp_bare"
local first_kind = existing.kind or "comp_bare"
corpus.collisions[#corpus.collisions + 1] = {
kind = "component",
name = c.name,
first_site = { path = existing.path, line = existing.line },
conflicting_site = { path = rel_path, line = c.line },
first_shape = "kind=" .. first_kind,
conflicting_shape = "kind=" .. kind,
}
end
end
end
end
--- (internal) Populate `corpus.component_body_index` with this source's body index entries.
--- First declaration wins; later declarations are dropped (no separate collision record: the components collision is already surfaced by `update_canonical_components`).
--- The pass writes to `corpus.component_body_index` only (the corpus owns this projection).
--- @param corpus table -- the corpus
--- @param src SourceFile
--- @param components Component[]
--- @param scan table -- the SourceScan payload (for line_of)
local function update_canonical_component_body_index(corpus, src, components, scan)
local line_of = scan and scan.line_of
for _, c in ipairs(components) do
if corpus.component_body_index[c.name] == nil then
corpus.component_body_index[c.name] = {
body_tokens = c.body_tokens,
body_off = c.body_off,
line_of = line_of,
source = src.path,
declaration = c.line,
kind = c.kind,
}
end
end
end
--- @param ctx PassCtx
--- @return PassResult
function M.run(ctx)
local outputs = {}
local errors = {}
local warnings = {}
-- Corpus ownership gate.
local corpus = ctx.shared and ctx.shared.corpus
if type(corpus) ~= "table" then
error("components.run requires ctx.shared.corpus.", 0)
end
if type(corpus.source_order) ~= "table" then
error("components.run requires ctx.shared.corpus.source_order.", 0)
end
if type(corpus.word_counts) ~= "table" then
error("components.run requires ctx.shared.corpus.word_counts; "
.. "word_count_eval.run must run before components.run "
.. "(see PASSES deps).", 0)
end
-- Projection ownership:
-- * `corpus.word_counts["mac_"..name]` — current component count
-- * `corpus.components[name]` — bare-name component definition
-- * `corpus.component_body_index[name]` — body / line_of / source index
-- The pass writes to the corpus only; consumers read from the corpus directly.
-- Per-directory aggregation: every source in the same directory contributes to one `gen/macs.h`.
-- The directory itself is the namespace. `corpus.sources_by_dir` preserves source-order within each bucket (matches `corpus.source_order`).
local sources_by_dir = corpus.sources_by_dir or duffle.group_sources_by_dir(corpus.source_order)
for dir, sources in pairs(sources_by_dir) do
-- Aggregate components from every source in this directory.
-- `project_components` returns nil for sources with no `MipsAtomComp_` declarations; we skip those.
local aggregated_components = {}
local metadata_per_source = {}
for _, src in ipairs(sources) do
local per_source = project_components(src.text, src.scan) or {}
for _, c in ipairs(per_source) do
aggregated_components[#aggregated_components + 1] = c
end
if #per_source > 0 then
metadata_per_source[src] = compute_components_metadata(per_source, duffle.INSTRUCTION_LATENCY)
end
end
if #aggregated_components > 0 then
-- Compute word counts across the aggregated set. `corpus.word_counts` carries the
-- same-source + prior-directory entries so the recursive lookup sees both.
local counts = count_all_components(aggregated_components, corpus.word_counts)
local macs_path = emit_component_macros_h(ctx, dir, sources, aggregated_components, counts)
if macs_path then
outputs[#outputs + 1] = { macs_h = macs_path }
-- Populate the projections AFTER disk emission (byte-identical `.macs.h` contract).
update_canonical_word_counts(corpus, aggregated_components, counts)
for _, src in ipairs(sources) do
local per_source = project_components(src.text, src.scan) or {}
if #per_source > 0 then
update_canonical_components(corpus, src, per_source, metadata_per_source[src])
update_canonical_component_body_index(corpus, src, per_source, src.scan)
end
end
end
end
end
return { outputs = outputs, errors = errors, warnings = warnings }
end
return M
File diff suppressed because it is too large Load Diff
+237
View File
@@ -0,0 +1,237 @@
--- passes/emission_model.lua: Per-atom emission projection.
---
--- The `emission-model` pass owns `atom.paths`, the canonical per-atom mutable surface for atoms and raw atoms with bodies in `ctx.shared.corpus.source_order`.
--- For each atom, the pass invokes `duffle.project_emission(body_text, component_index, word_counts, components)`.
--- It stores the ordered `items` stream plus the dense `word_events` / `markers` / `invocations` views on `atom.paths`.
---
--- Public boundary:
--- * `M.run(ctx)` is the only entry point.
--- * The pass returns `{outputs = {}, errors = ..., warnings = ...}`.
--- Pass kind = `validation` → `PASS_KIND_STOP_ON_ERROR.validation` preserves the existing build-stopping policy.
---
--- Source-order discipline:
--- * `corpus.source_order` sets the source-record order.
--- * Within each source, the pass visits `src.scan.atoms` and `src.scan.raw_atoms` in declaration order.
---
--- Per-atom projection fields on `atom.paths`:
--- `tokens`, `line_in_body`, `items`, `word_events`, `markers`, `invocations`, `errors`, `warnings`.
--- The construction walk appends `items` and derives each dense view from that ordered stream.
---
--- Component expansion and construction validation:
--- * known `mac_X(...)` calls recursively expand component bodies;
--- * invocation records retain monotonic IDs, parent IDs, immediate call text, and the immutable outermost root call text;
--- * invocation construction stamps `debug_skip` from `corpus.components[name].debug_skip` at the construction site (no second pass, no source parse, no parallel lookup);
--- * component cycles close balanced invocation boundaries and emit a `cycle` construction error at the recursive edge;
--- * declared-vs-measured component word counts emit `count_mismatch` construction errors; opaque uncounted macros emit warnings.
---
--- `passes.scan_source` strips its private `_code_macros` / `_code_macro_bodies` tables before this pass runs.
local M = {}
-- ─────────────────────────────────────────────────────────────────────────
-- Bootstrap: load `duffle_paths.lua` via debug.getinfo so the module works standalone (run as `luajit passes/emission_model.lua`) and when require'd from the orchestrator.
-- ─────────────────────────────────────────────────────────────────────────
local _bootstrap_dir = debug.getinfo(1, "S").source:match("^@?(.*[/\\])") or "./"
local duffle = dofile(_bootstrap_dir .. "../duffle_paths.lua")
-- ─────────────────────────────────────────────────────────────────────────
-- Helpers
-- ─────────────────────────────────────────────────────────────────────────
-- Convert the recursive walk's body-relative line numbers into physical source lines once.
-- The walker builds `line_of` from `body_text` and stamps body-relative line numbers (1..N) into `item.line` and `invocation.call_line`.
-- This function converts those values to physical source lines at the close site with the forwarded source `line_of` closure.
-- `call_line` discipline:
-- * ROOT invocations (`inv.parent_id == 0`) receive body-relative `call_line` values directly from `M.LineIndex(body_text)` in the walker.
-- The source `line_of` closure supplies physical lines at the close site, so this function converts each root value exactly once.
-- * INNER invocations (`inv.parent_id ~= 0`) receive physical `call_line` values directly from the COMPONENT's `line_of` in the walker.
-- Recursive descent forwards that closure through `corpus.component_body_index[name].line_of`; those values arrive physical and remain unchanged.
--
-- After this function, every `inv.call_line` is physical. DWARF and provenance output read it directly.
-- The word-event loop forwards the already-physical `outer_inv.call_line` into `we.call_line` for words inside an invocation.
local function stamp_root_provenance(projection, atom_record, src, corpus)
local root_line_of = src.scan and src.scan.line_of
assert(type(root_line_of) == "function"
, "emission_model: src.scan.line_of is required (canonical LineIndex closure over the source text) to stamp physical provenance")
assert(type(atom_record.body_off) == "number"
, "emission_model: atom_record.body_off (byte offset of the body's first byte in source) is required to derive `root_body_line`. The scanner must populate body_off for every atom record.")
-- `root_body_line` is the physical source line of the ATOM HEADER byte containing the opening `{`; that byte is one byte BEFORE `atom_record.body_off`.
-- The walker assigns line 2 to the body's first content line because line 1 is the trailing `\n` after `{`. Body-text line k therefore maps to `root_body_line + (k - 1)`.
-- `body_off - 1` points at the opening `{`, whose line index identifies the header line. `body_off` points after `{` and would shift every word row forward by one line.
local root_body_line = root_line_of(atom_record.body_off - 1) or atom_record.line or 0
local component_index = corpus.component_body_index or {}
local word_items = {}
for _, item in ipairs(projection.items) do
if item.kind == "word" then word_items[#word_items + 1] = item end
end
-- Resolve one word's physical body line, where the byte containing that word appears in source.
-- * Component expansions carry `invocation_ids`; the component's full-file `line_of` leaves `item.line` physical.
-- * Raw tokens in the root atom body carry an empty `invocation_ids` list and a body-relative `item.line`; convert them here.
local function body_line_for(event, item)
local ids = event.invocation_ids or {}
-- The innermost open invocation identifies which line index the walker used.
-- A component `line_of` makes `item.line` physical; the atom's `body_text` line index makes it body-relative.
if ids and #ids > 0 then
local inner_id = ids[#ids]
local inner_inv = inner_id and projection.invocations[inner_id]
if inner_inv then
local component = component_index[inner_inv.component_name]
if component and component.line_of then
-- Walker used `comp.line_of`, which is the source's physical LineIndex. item.line is already physical.
return item.line or 0
end
end
end
-- RAW root-body word: item.line is body-text's 1-based line number (the first content line is line 2 because line 1 is the trailing `\n` after `{`).
-- Convert body-text-relative → physical using `root_body_line + (item.line - 1)`.
return (root_body_line or 0) + (item.line or 1) - 1
end
-- Stamp the root source path onto invocation records whose `call_path` the walker left empty.
-- The walker passes `body_entry.source` to `emit_invoke_begin`; `M.project_emission` creates the root `body_entry` with source `""`, leaving its `call_path` empty.
-- This stamp gives every invocation a physical `call_path` matching `passes/atoms_source_map.lua`'s in-memory provenance projection.
local root_path = src.path or ""
for _, inv in ipairs(projection.invocations) do
if inv.call_path == nil or inv.call_path == "" then
inv.call_path = root_path
end
end
-- Normalize `inv.call_line` to a physical source line.
-- * ROOT invocations (`parent_id == 0`) carry body-relative `call_line` values from `M.LineIndex(body_text)`; convert them once with `root_body_line`.
-- * INNER invocations (`parent_id ~= 0`) carry physical `call_line` values from the component's `line_of`; retain them unchanged.
for _, inv in ipairs(projection.invocations) do
if inv.parent_id == 0 then
inv.call_line = (root_body_line or 0) + (inv.call_line or 1) - 1
end
end
-- Build `body_lines` for each invocation.
-- `atoms_source_map` and `dwarf_injection` read `inv.body_lines[k]` directly from the invocation record created here.
-- Component words already carry physical `item.line` values from the walker's COMPONENT line index, so `body_line_for` returns them unchanged.
for _, inv in ipairs(projection.invocations) do
local sw = inv.start_word
local ew = inv.end_word
local bls = {}
for i = sw, ew do
local it = projection.items and projection.items[i]
if it and it.kind == "word" then
local fake_event = { invocation_ids = { inv.id } }
bls[#bls + 1] = body_line_for(fake_event, it) or 0
end
end
inv.body_lines = bls
end
-- Resolve each `word_event`'s physical `body_line` and `call_line`.
-- For words inside an invocation, `we.call_line` identifies the OUTER atom source line containing the `mac_X(...)` token that triggered expansion.
-- The root-invocation conversion above makes every `inv.call_line` physical; forward it directly and use each raw word's `body_line` as the fallback.
for index, we in ipairs(projection.word_events) do
local item = word_items[index] or {}
local body_line = body_line_for(we, item)
item.line = body_line
we.body_line = body_line
local call_line = body_line
local outer_id = we.outermost_invocation_id or 0
local outer_inv = projection.invocations[outer_id]
if outer_inv then
-- `outer_inv.call_line` is physical after the conversion loop above, so use it directly.
call_line = outer_inv.call_line
end
we.call_line = call_line
if we.def_path == nil or we.def_path == "" then we.def_path = src.path or "" end
if we.def_line == nil or we.def_line == 0 then we.def_line = atom_record.line or 0 end
if we.call_path == nil or we.call_path == "" then we.call_path = src.path or "" end
end
end
-- Project one atom record into `atom.paths`.
-- Mutates the atom record in-place and returns the projection (for pass-level error/warning accumulation).
local function project_atom(atom_record, src, corpus)
local body = atom_record.body or ""
local wc = corpus.word_counts or {}
local cbi = corpus.component_body_index or {}
-- That construction site stamps `invocation.debug_skip` while appending each record to `proj.invocations`.
local proj = duffle.project_emission(body, cbi, wc, corpus.components)
local paths = {
tokens = atom_record.body_tokens or {},
line_in_body = duffle.build_body_line_index(body),
items = proj.items,
word_events = proj.word_events,
markers = proj.markers,
invocations = proj.invocations,
errors = proj.errors,
warnings = proj.warnings,
}
stamp_root_provenance(proj, atom_record, src, corpus)
atom_record.paths = paths
return proj
end
-- ─────────────────────────────────────────────────────────────────────────
-- Run the emission-model pass.
-- ─────────────────────────────────────────────────────────────────────────
--- @param ctx PassCtx -- { shared = { corpus = ... }, out_root, ... }
--- @return PassResult
function M.run(ctx)
local outputs = {}
local errors = {}
local warnings = {}
local corpus = ctx and ctx.shared and ctx.shared.corpus
if type(corpus) ~= "table" then error("emission_model: ctx.shared.corpus is required (canonical projection)", 0) end
if type(corpus.source_order) ~= "table" then error("emission_model: ctx.shared.corpus.source_order is required", 0) end
-- Project once, collect errors + warnings for one atom.
-- Kind must be one of: atom | raw_atom | comp_bare | comp_proc.
local function process_atom(atom, src)
if not (atom and atom.body) then return end
local kind = atom.kind
if kind ~= "atom" and kind ~= "raw_atom" and kind ~= "comp_bare" and kind ~= "comp_proc" then
return
end
local proj = project_atom(atom, src, corpus)
for _, e in ipairs(proj.errors) do
-- Preserve `kind` (cycle / count_mismatch / unbalanced) so readers dispatch on the diagnostic class and leave the message string as display text.
errors[#errors + 1] = {
kind = e.kind,
line = e.line,
msg = e.msg,
source = e.source or src.path,
}
end
for _, w in ipairs(proj.warnings) do
warnings[#warnings + 1] = {
kind = w.kind,
line = w.line,
msg = w.msg,
}
end
end
-- Walk `corpus.source_order`; within each source, visit atoms followed by raw_atoms.
-- Recognized kinds (atom | raw_atom | comp_bare | comp_proc) each receive the atom.paths projection via duffle.project_emission.
-- Components are macros inlined into atom bodies; focused tests and isolated component analyses consume atom.paths directly.
for _, src in ipairs(corpus.source_order) do
local scan = src.scan or {}
for _, atom in ipairs(scan.atoms or {}) do
process_atom(atom, src)
end
for _, atom in ipairs(scan.raw_atoms or {}) do
process_atom(atom, src)
end
end
return {
outputs = outputs,
errors = errors,
warnings = warnings,
}
end
return M
+289
View File
@@ -0,0 +1,289 @@
--- passes/offsets.lua — Branch-offset generator.
---
--- Reads the pre-scanned SourceScan payload (produced once upstream by `duffle.scan_source`)
--- for `MipsAtom_(name)` and `MipsCode code_<name>` declarations, computes the word offset
--- from each `atom_offset(F, T)` marker to its target `atom_label(T)` declaration, and emits
--- `gen/offsets.h` with one `#define _atom_offset_F_T = N` per branch.
--- Per-directory aggregation: every source in the same directory contributes to the same `gen/offsets.h`.
--- The directory itself is the namespace; the filename does not repeat the module name.
---
--- The offset is `target_word - branch_word - 1` (the standard MIPS branch-immediate encoding: branch_offset = relative_pc_in_words - 1).
-- ════════════════════════════════════════════════════════════════════════════
-- Module-scope requires + package.path setup
-- ════════════════════════════════════════════════════════════════════════════
-- Bootstrap: same as entry scripts. See `ps1_meta.lua` for the rationale.
-- Bootstrap: load `scripts/duffle_paths.lua` (sets package.path + package.cpath).
-- Uses `debug.getinfo` to find this file's own directory, so it works both standalone and when require'd from the orchestrator.
-- Bootstrap: load `duffle_paths.lua` via `debug.getinfo(1, "S").source` (works both standalone + when require'd).
-- duffle_paths.lua sets package.path then returns `require("duffle")` at the bottom, so the dofile value IS the duffle module.
local _bootstrap_dir = debug.getinfo(1, "S").source:match("^@?(.*[/\\])") or "./"
local duffle = dofile(_bootstrap_dir .. "../duffle_paths.lua")
-- ════════════════════════════════════════════════════════════════════════════
-- Constants
-- ════════════════════════════════════════════════════════════════════════════
-- Offset macro/enum naming prefixes (the emitted header uses these).
local OFFSET_MACRO_PREFIX = "_atom_offset_"
local OFFSET_ENUM_PREFIX = "atom_offset_"
-- Column width for the `#define _atom_offset_F_T = N` alignment.
local OFFSET_MACRO_COL = 44
-- ════════════════════════════════════════════════════════════════════════════
-- Type declarations
-- ════════════════════════════════════════════════════════════════════════════
--- @class SourceFile
--- @field path string -- Absolute path to the source file
--- @field text string -- Full source text
--- @field dir string -- Directory containing the source
--- @field basename string -- Filename without extension
--- @field scan table -- Pre-scanned SourceScan payload (from duffle.scan_source)
--- @class PassCtx
--- @field shared table -- Cross-pass shared state
--- @field shared.corpus table -- Corpus projection
--- @field shared.word_counts table
--- @field out_root string -- Output root (e.g. "build/gen")
--- @class PassResult
--- @field outputs table[] -- {kind=, path=} entries describing emit files
--- @field errors table[] -- {line=, msg=} entries; build-stops
--- @field warnings table[] -- {line=, msg=} entries; build-succeeds
--- @class BranchOffset
--- @field tag string -- Marker tag (e.g. "F" in `atom_offset(F, T)`)
--- @field target string -- Target label name (e.g. "T" in `atom_offset(F, T)`)
--- @field branch_word integer -- Branch word position within the atom body
--- @field offset integer -- Computed per consuming instruction (see `compute_offsets`)
--- @field consuming_encoder string|nil -- Instruction consuming the offset (e.g. "branch_le_zero", "jump", "call_addr")
--- @field consuming_arg_pos integer|nil -- 1-based arg position within the consuming instruction's arg list
--- @class AtomData
--- @field name string -- Atom name
--- @field total_words integer -- Total word count of the atom body
--- @field offsets BranchOffset[] -- Per-branch offset list
-- ════════════════════════════════════════════════════════════════════════════
-- Canonical marker projection
-- ════════════════════════════════════════════════════════════════════════════
-- MARKER_PROJECTORS is the marker-kind data table.
-- The emission-model pass already records marker word positions + consuming-instruction context;
-- this pass only projects those records into the label/branch lookup shape needed by offset computation.
local MARKER_PROJECTORS = {
label = function(state, marker)
state.labels[marker.name] = marker.word_index
end,
offset = function(state, marker)
state.branches[#state.branches + 1] = {
tag = marker.name,
target = marker.target,
branch_word = marker.word_index,
consuming_encoder = marker.consuming_encoder,
consuming_arg_pos = marker.consuming_arg_pos,
}
end,
}
--- Project canonical marker records into the two lookup tables used by the offset renderer.
--- No source text, body text, or body token is inspected.
--- @param markers table[] -- atom.paths.markers
--- @return table<string, integer>, table[]
local function project_markers(markers)
local state = { labels = {}, branches = {} }
for _, marker in ipairs(markers or {}) do
local project = MARKER_PROJECTORS[marker.kind]
if project then project(state, marker) end
end
return state.labels, state.branches
end
-- ════════════════════════════════════════════════════════════════════════════
-- Offset computation + header generation
-- ════════════════════════════════════════════════════════════════════════════
--- Compute branch offsets per consuming instruction.
--- Disposition table:
--- `branch_*` -> relative offset: `target_word - branch_word - 1` (MIPS branch-immediate encoding).
--- `jump` / `call_addr` -> same value as `branch_*` (a relative word offset).
--- The duffle headers' `enc_i` macro truncates the value to the immediate-field width (16 bits for branches, 26 bits for jumps).
--- For tape-atom bodies within a single module, this works for `j`/`jal` because the linker's symbol resolution produces the correct 26-bit absolute target via standard `j` relocations.
--- For cross-module `j`/`jal` (atom body in one module, target in another), the linker emits a `R_MIPS_26` relocation against the lower 26 bits; the upper 4 bits come from the PC of the delay slot following the `j`.
--- The metaprogram doesn't know either at compile time, so the emitted value is the relative word offset that the duffle `enc_i` macro places in the immediate field; the toolchain handles the rest.
--- `jump_reg` / `call_reg` / `jump_link` -> ERROR. Register-form jumps have no offset field; `atom_offset` is invalid.
---
--- Top-level `atom_offset(F, T)` markers (where the marker is the entire token — `consuming_encoder` == nil) default to `branch_*` behavior (relative offset).
--- This preserves backward compatibility for any top-level marker that may exist outside a control-transfer instruction.
--- @param labels table<string, integer>
--- @param branches table[]
--- @return BranchOffset[]
local function compute_offsets(labels, branches)
local results = {}
for _, br in ipairs(branches) do
local target = labels[br.target]
if not target then
error("Branch target '" .. br.target .. "' has no atom_label (at word " .. br.branch_word .. ")")
end
local consuming = br.consuming_encoder
local offset
if consuming == "jump_reg" or consuming == "call_reg" or consuming == "jump_link" then
-- Register-form jumps have no offset field. `atom_offset` cannot be used here.
error("atom_offset cannot be used with " .. consuming
.. " (register-form jumps have no offset field); at word " .. br.branch_word)
end
-- All other consuming instructions (including `branch_*`, `jump`, `call_addr`, and nil for top-level markers) use the same relative offset value.
-- The MIPS encoding differs per opcode but the duffle `enc_i` macro handles the truncation to the immediate-field width.
offset = target - br.branch_word - 1
results[#results + 1] = {
target = br.target,
tag = br.tag,
branch_word = br.branch_word,
offset = offset,
consuming_encoder = br.consuming_encoder,
consuming_arg_pos = br.consuming_arg_pos,
}
end
return results
end
--- Right-pad `s` with spaces to width `w`. If `s` is already `w` or wider, no padding is added.
--- @param s string
--- @param w integer
--- @return string
local function pad_right(s, w)
return s .. string.rep(" ", math.max(0, w - #s))
end
--- (internal) Build a constant-table entry `{macro_name, enum_name, value}` from a BranchOffset.
--- @param bo BranchOffset
--- @return table
local function make_offset_const(bo)
return {
macro_name = OFFSET_MACRO_PREFIX .. bo.tag .. "_" .. bo.target,
enum_name = OFFSET_ENUM_PREFIX .. bo.tag .. "_" .. bo.target,
value = bo.offset,
}
end
--- (internal) Emit one atom's offset constants + enum into the lines buffer.
--- @param add fun(s: string)
--- @param atom AtomData
local function emit_atom_offsets(add, atom)
if #atom.offsets == 0 then return end
add("// --- atom: " .. atom.name .. " (" .. atom.total_words .. " words) ---")
add("")
local consts = {}
for _, r in ipairs(atom.offsets) do
consts[#consts + 1] = make_offset_const(r)
end
for _, c in ipairs(consts) do
add("#define " .. pad_right(c.macro_name, OFFSET_MACRO_COL) .. " " .. c.value)
end
add("")
add("enum {")
for _, c in ipairs(consts) do
add(" " .. c.enum_name .. " = " .. c.macro_name .. ",")
end
add("};")
add("")
end
--- Generate the per-directory .offsets.h header.
--- @param dir string -- the absolute source directory
--- @param sources table[] -- sources contributing to this directory (for the header comment)
--- @param atoms_data AtomData[]
--- @return string
local function generate_header(dir, sources, atoms_data)
local dir_basename = duffle.basename_no_ext(dir)
local lines = {}
local function add(s) lines[#lines + 1] = s end
add("// Auto-generated by ps1_meta.lua (passes/offsets.lua) — DO NOT EDIT")
add("// Directory: " .. dir:gsub("/", "\\") .. "\\")
for _, src in ipairs(sources) do
add("// source: " .. src.path:gsub("/", "\\"))
end
add("#pragma once")
add("")
add("#pragma region " .. dir_basename)
add("")
add("")
for _, atom in ipairs(atoms_data) do
emit_atom_offsets(add, atom)
end
add("#pragma endregion " .. dir_basename)
add("")
return table.concat(lines, "\n") .. "\n"
end
local M = {}
--- (internal) Aggregate atoms from every source in one directory, render the per-directory `offsets.h`.
--- Returns the offsets_h path if a header was written, or nil.
--- @param ctx PassCtx
--- @param dir string -- the absolute source directory
--- @param sources SourceFile[] -- sources in this directory
--- @return string|nil -- the offsets_h path
local function process_directory(ctx, dir, sources)
local atoms_data = {}
local function append_atom(atom)
local paths = atom and atom.paths
if not paths then return end
local labels, branches = project_markers(paths.markers)
atoms_data[#atoms_data + 1] = {
name = atom.raw_name or atom.name,
total_words = #(paths.word_events or {}),
offsets = compute_offsets(labels, branches),
}
end
for _, src in ipairs(sources) do
local scan = src.scan or {}
for _, atom in ipairs(scan.atoms or {}) do append_atom(atom) end
for _, atom in ipairs(scan.raw_atoms or {}) do append_atom(atom) end
end
if #atoms_data == 0 then return nil end
local out_path = dir .. "/gen/offsets.h"
duffle.ensure_dir(duffle.dirname(out_path))
duffle.write_file(out_path, generate_header(dir, sources, atoms_data))
return out_path
end
--- Run the offsets pass.
--- For each canonical source-directory, emits a per-directory `gen/offsets.h`
--- containing constants for every marker recorded in atom.paths across every source in that directory.
--- @param ctx PassCtx
--- @return PassResult
function M.run(ctx)
local outputs = {}
local errors = {}
local warnings = {}
local corpus = ctx.shared and ctx.shared.corpus
if type(corpus) ~= "table" then
error("offsets.run requires ctx.shared.corpus", 0)
end
if type(corpus.source_order) ~= "table" then
error("offsets.run requires ctx.shared.corpus.source_order.", 0)
end
-- Per-directory aggregation: every source in the same directory contributes to one `gen/offsets.h`.
local sources_by_dir = corpus.sources_by_dir or duffle.group_sources_by_dir(corpus.source_order)
for dir, sources in pairs(sources_by_dir) do
local out_path = process_directory(ctx, dir, sources)
if out_path then
outputs[#outputs + 1] = { offsets_h = out_path }
end
end
return { outputs = outputs, errors = errors, warnings = warnings }
end
return M
+599
View File
@@ -0,0 +1,599 @@
--- passes/report.lua — Per-MODULE annotation report renderer + project-wide summary writer.
---
--- Two output files per build:
--- - `build/gen/<dir_basename>.annotations.txt` — one per source-directory containing atoms; aggregates across all sources in the directory.
--- - `build/gen/annotation_validation.txt` — the project summary.
---
--- The annotation pass emits `errors.h` files per module and the canonical `corpus.sources_by_dir` projection groups sources by directory.
--- This pass iterates the dir projection directly and re-validates each source via `annotation.validate()` to get the detailed per-source results.
-- ════════════════════════════════════════════════════════════════════════════
-- Module-scope requires + package.path setup
-- ════════════════════════════════════════════════════════════════════════════
-- Resolve `arg[0]` to an absolute-ish script directory so that `require("duffle")` resolves against `scripts/` regardless of CWD.
-- Bootstrap: See `ps1_meta.lua` for the rationale.
-- Bootstrap: Load `scripts/duffle_paths.lua` (sets package.path + package.cpath).
-- Uses `debug.getinfo` to find this file's own directory, so it works both standalone and when require'd from the orchestrator.
-- Bootstrap: Load `duffle_paths.lua` via `debug.getinfo(1, "S").source` (works both standalone + when require'd).
-- duffle_paths.lua sets package.path then returns `require("duffle")` at the bottom, so the dofile value IS the duffle module.
local _bootstrap_dir = debug.getinfo(1, "S").source:match("^@?(.*[/\\])") or "./"
local duffle = dofile(_bootstrap_dir .. "../duffle_paths.lua")
-- Load the annotation pass so we can re-validate each source against the canonical corpus projection.
-- The annotation pass exposes `M.validate`, which returns the per-source AnnotationResult (atoms / annots / macros / binds / errors / warnings)
-- that the report pass renders into the per-module `<dir_basename>.annotations.txt` output.
local annotation = dofile(_bootstrap_dir .. "annotation.lua")
-- Load atoms_source_map for the `render_source_map` / `render_provenance` module functions (used by `render_module_atoms_md` to produce `<module>.atoms.md` without re-walking source tokens).
-- The pass itself emits no per-source files anymore; we only consume the two pure renderers here.
-- Defined BEFORE the renderer functions below so their upvalues resolve to this local (not the global `atoms_source_map`, which is nil).
local atoms_source_map = dofile(_bootstrap_dir .. "atoms_source_map.lua")
-- ════════════════════════════════════════════════════════════════════════════
-- Constants
-- ════════════════════════════════════════════════════════════════════════════
-- Section separators used in the rendered text reports.
-- The thin rules are hand-tuned to align with the per-section content width; do not change without also checking the section renderers below.
local RULE_THICK = "========================================================"
local SECTION_HEADER_ATOMS = "── Atoms ────────────────────────────────────────────────"
local SECTION_HEADER_ANNOTS = "── Annotations ──────────────────────────────────────────"
local SECTION_HEADER_BINDS = "── Binds_* structs ──────────────────────────────────────"
local SECTION_HEADER_MACROS = "── Macro word-count declarations ─────────────────────────"
local SECTION_HEADER_ERRORS = "── Errors ──────────────────────────────────────────────"
local SECTION_HEADER_WARNINGS = "── Warnings ────────────────────────────────────────────"
-- Lua pattern that captures the basename (last path segment) of a forward- or back-slash separated path.
local BASENAME_PATTERN = "([^/\\]+)$"
-- Debug flag name — set to truthy in `_G` to enable verbose logging.
local DEBUG_FLAG = "_DEBUG_REPORT"
-- Pass identifier for log messages.
local PASS_NAME = "report"
-- ════════════════════════════════════════════════════════════════════════════
-- Type declarations
-- ════════════════════════════════════════════════════════════════════════════
--- @class SourceFile
--- @field path string -- Absolute path to the source file
--- @field text string -- Full source text
--- @field dir string -- Directory containing the source
--- @field basename string -- Filename without extension
--- @class PassCtx
--- @field sources SourceFile[] -- All source files in the build
--- @field metadata_path string -- Path to word_count.metadata.h
--- @field shared table -- Cross-pass shared state
--- @field out_root string -- Output root (e.g. "build/gen")
--- @field project_root string -- Project root (e.g. "code/")
--- @field upstream table<string, table> -- Per-pass upstream outputs
--- @field flags table -- CLI flags + per-pass stash
--- @field verbose boolean -- If true, log diagnostic info
--- @class PassResult
--- @field outputs table[] -- {kind=, path=} entries describing emit files
--- @field errors table[] -- {line=, msg=} entries; build-stops
--- @field warnings table[] -- {line=, msg=} entries; build-succeeds
-- Shapes produced by `passes/annotation.lua`'s `M.validate()`.
--- @class AtomEntry
--- @field name string -- Atom name (e.g. "cube_g4_face")
--- @field line integer -- Source line of the atom declaration
--- @class AnnotEntry
--- @field line integer -- Source line
--- @field macro string -- Macro name (e.g. "atom_reads")
--- @field name string -- Atom name (if a `name(...)` was given)
--- @field kind string -- "atom_info" | "atom_bind" | ...
--- @field binds string|nil -- Binds_X name if any
--- @field reads string[] -- R_* names (read targets)
--- @field writes string[] -- R_* names (write targets)
--- @field error string|nil -- Error message if annotation was malformed
--- @class BindsField
--- @field name string -- Field name
--- @field offset integer -- Byte offset within the Binds_X struct
--- @class BindsStruct
--- @field name string -- Struct name (e.g. "Binds_Floor")
--- @field line integer -- Source line of the typedef
--- @field bytes integer -- Total byte size
--- @field fields BindsField[] -- The field list
--- @class MacroEntry
--- @field name string -- Macro name (e.g. "WORD_COUNT(my_macro, 4)")
--- @field line integer -- Source line
--- @field words integer -- Declared word count
--- @class Finding
--- @field line integer -- Source line
--- @field msg string -- Finding message
--- @class AnnotationResult
--- @field source string -- Set by this pass; original source path
--- @field atoms AtomEntry[] -- Atom declarations in this source
--- @field annots AnnotEntry[] -- Annotation entries
--- @field macros MacroEntry[] -- Macro word-count declarations
--- @field binds BindsStruct[] -- Binds_* struct declarations
--- @field errors Finding[] -- Errors from validation
--- @field warnings Finding[] -- Warnings from validation
--- @field info table -- Info summary (not rendered here)
--- @class ModuleEntry
--- @field dir string -- Absolute directory path
--- @field dir_basename string -- Basename (e.g. "duffle", "gte_hello")
--- @field atoms_count integer -- Pre-counted atoms for filtering
--- @class ModuleReport
--- @field dir string -- Module directory
--- @field sources SourceFile[] -- Sources in this module
--- @field results AnnotationResult[] -- Per-source validate() results
--- @class ProjectReport
--- @field results AnnotationResult[] -- All per-source results
-- ════════════════════════════════════════════════════════════════════════════
-- Per-MODULE annotation report (aggregated across all sources in a dir)
-- ════════════════════════════════════════════════════════════════════════════
--- Extract the basename (last path segment) of a forward- or back-slash separated path. Returns the input unchanged if no separator is found.
--- @param path string
--- @return string
local function source_basename(path)
return path:match(BASENAME_PATTERN) or path
end
-- ════════════════════════════════════════════════════════════════════════════
-- Markdown renderers (consolidated-report-files refactor, 2026-07-26)
-- ════════════════════════════════════════════════════════════════════════════
--- Render the thin project-wide summary (`build/atom_meta_report.summary.md`).
--- @param all_results {
--- module:string,
--- atoms:integer,
--- annots:integer,
--- binds:integer,
--- macros:integer,
--- findings:integer,
--- errors:integer,
--- warnings:integer,
--- info:integer }[]
--- @return string
local function render_project_summary(all_results)
local lines = {
"# Project summary",
"> Auto-generated by ps1_meta.lua (passes/report.lua).",
"",
"| module | atoms | annots | binds | macros | findings | errors | warnings | info |",
"|--------|-------|--------|-------|--------|----------|--------|----------|------|",
}
local totals = { atoms = 0, annots = 0, binds = 0, macros = 0, findings = 0, errors = 0, warnings = 0, info = 0 }
for _, e in ipairs(all_results) do
lines[#lines + 1] = string.format("| %s | %d | %d | %d | %d | %d | %d | %d | %d |"
, e.module, e.atoms, e.annots, e.binds, e.macros, e.findings, e.errors, e.warnings, e.info)
totals.atoms = totals.atoms + e.atoms
totals.annots = totals.annots + e.annots
totals.binds = totals.binds + e.binds
totals.macros = totals.macros + e.macros
totals.findings = totals.findings + e.findings
totals.errors = totals.errors + e.errors
totals.warnings = totals.warnings + e.warnings
totals.info = totals.info + e.info
end
lines[#lines + 1] = string.format("| **TOTAL** | %d | %d | %d | %d | %d | %d | %d | %d |"
, totals.atoms, totals.annots, totals.binds, totals.macros, totals.findings, totals.errors, totals.warnings, totals.info)
return table.concat(lines, "\n") .. "\n"
end
--- Render the per-module verbose source-map markdown (`build/<module>.atoms.md`).
--- Per-source sub-section, per-atom stanza with sourcemap + provenance rows.
--- Pulls sourcemap + provenance from `atoms_source_map` (no second source walk).
--- @param dir string
--- @param dir_sources SourceFile[]
--- @param wc table<string, integer>
--- @return string
local function render_module_atoms_md(dir, dir_sources, wc)
local dir_basename = source_basename(dir)
local lines = {
"# " .. dir_basename .. " — atoms (verbose source map)",
"> Per-word call-site + provenance. Auto-generated.",
"",
}
for _, src in ipairs(dir_sources) do
local src_name = source_basename(src.path)
lines[#lines + 1] = "## " .. src_name
lines[#lines + 1] = ""
-- For each atom with a projection, render its sourcemap + provenance.
local atoms_list = {}
for _, atom in ipairs((src.scan or {}).atoms or {}) do
if atom.paths then atoms_list[#atoms_list + 1] = atom end
end
for _, atom in ipairs((src.scan or {}).raw_atoms or {}) do
if atom.paths then atoms_list[#atoms_list + 1] = atom end
end
if #atoms_list == 0 then
lines[#lines + 1] = "_(no atom projections)_"
lines[#lines + 1] = ""
else
-- Per-source forward-slash path (same one `emit_atom_stanza` / `emit_provenance_stanza` would derive;
-- computed once per `## <source>` heading and reused by each atom's `WORD N CALL ...` field).
local rel_path = src.path:gsub("\\\\", "/")
for _, atom in ipairs(atoms_list) do
lines[#lines + 1] = string.format(
"### atom: %s (line %d, %d words)",
atom.name, atom.line or 0, #(atom.paths.items or {}))
lines[#lines + 1] = ""
lines[#lines + 1] = "**Sourcemap** — per-word call site:"
lines[#lines + 1] = "```"
-- Per-atom invariant: call the per-atom renderers, NOT the per-source ones.
-- The per-source renderers enumerate every atom in `src`;
-- calling them in a per-atom loop would repeat the whole source under every `### atom:` heading.
lines[#lines + 1] = atoms_source_map.render_atom_source_map(atom):gsub("\n+$", "")
lines[#lines + 1] = "```"
lines[#lines + 1] = ""
lines[#lines + 1] = "**Provenance** — per-word definition + body:"
lines[#lines + 1] = "```"
lines[#lines + 1] = atoms_source_map.render_atom_provenance(atom, wc, rel_path):gsub("\n+$", "")
lines[#lines + 1] = "```"
lines[#lines + 1] = ""
end
end
end
return table.concat(lines, "\n") .. "\n"
end
--- Render the consolidated per-module markdown (`build/<module>.atom_meta_report.md`).
--- Aggregates annotation + static-analysis content across all sources in `dir`.
--- Annotations come from re-running `annotation.validate()` per source (the existing pattern);
--- static-analysis comes from `corpus.static_analysis_results[dir_basename]` (populated by `static_analysis.lua` — no second corpus_pipe_ctx build).
--- @param dir string
--- @param dir_sources SourceFile[]
--- @param annot_results AnnotationResult[]
--- @param sa_results table -- corpus.static_analysis_results[dir_basename]
--- @return string
local function render_module_meta_report(dir, dir_sources, annot_results, sa_results)
local dir_basename = source_basename(dir)
local lines = {
"# " .. dir_basename .. " — atom meta report",
"> Auto-generated by ps1_meta.lua (passes/report.lua). Do not edit.",
"",
}
local function add(s) lines[#lines + 1] = s end
-- Module summary table.
local n_atoms = 0
local n_annot = 0
local n_binds = 0
local n_macros = 0
local n_bare, n_proc = 0, 0
for _, r in ipairs(annot_results) do
n_atoms = n_atoms + #r.atoms
n_annot = n_annot + #r.annots
n_binds = n_binds + #r.binds
n_macros = n_macros + #r.macros
end
for _, a in ipairs(sa_results.atoms or {}) do
if a.kind == "comp_bare" then n_bare = n_bare + 1
elseif a.kind == "comp_proc" then n_proc = n_proc + 1
end
end
add("## Module summary"); add("")
add("| metric | value |"); add("|--------|-------|")
add(string.format("| sources | %d |", #dir_sources))
add(string.format("| atoms | %d (atoms: %d, comp_bare: %d, comp_proc: %d) |",
#(sa_results.atoms or {}),
#(sa_results.atoms or {}) - n_bare - n_proc, n_bare, n_proc))
add(string.format("| annotations | %d |", n_annot))
add(string.format("| binds structs | %d |", n_binds))
add(string.format("| macro decls | %d |", n_macros))
add(string.format("| findings | %d (errors: %d, warnings: %d, info: %d) |",
#(sa_results.findings or {}),
#(sa_results.errors or {}),
#(sa_results.warnings or {}),
#(sa_results.info or {})))
add("")
-- Sources
add("## Sources"); add("")
for _, s in ipairs(dir_sources) do add("- `" .. s.path .. "`") end
add("")
-- Atoms (annotation)
add("## Atoms"); add("")
add("| kind | name | source | line |"); add("|------|------|--------|------|")
for _, r in ipairs(annot_results) do
local src_name = source_basename(r.source)
for _, a in ipairs(r.atoms) do
add(string.format("| atom | %s | %s | %d |", a.name, src_name, a.line))
end
end
add("")
-- Annotations
add("## Annotations"); add("")
if #annot_results == 0 then
add("_(none)_")
else
add("| source | line | name | binds | reads | writes |")
add("|--------|------|------|-------|-------|--------|")
for _, r in ipairs(annot_results) do
local src_name = source_basename(r.source)
for _, a in ipairs(r.annots) do
local binds = a.binds or ""
local reads = (#a.reads > 0 and table.concat(a.reads, ",")) or ""
local writes = (#a.writes > 0 and table.concat(a.writes, ",")) or ""
add(string.format("| %s | %d | %s | %s | %s | %s |"
, src_name, a.line, a.name, binds, reads, writes))
end
end
end
add("")
-- Binds_* structs
add("## Binds_* structs"); add("")
if #annot_results == 0 then
add("_(none)_")
else
for _, r in ipairs(annot_results) do
local src_name = source_basename(r.source)
for _, b in ipairs(r.binds) do
add(string.format("### %s (%s:%d, %d bytes)",
b.name, src_name, b.line, b.bytes))
for _, f in ipairs(b.fields) do
add(string.format("- `+%d %s`", f.offset, f.name))
end
add("")
end
end
end
-- Macro decls
add("## Macro word-count declarations"); add("")
if #annot_results == 0 then
add("_(none)_")
else
add("| source | line | macro declaration |")
add("|--------|------|-------------------|")
for _, r in ipairs(annot_results) do
local src_name = source_basename(r.source)
for _, m in ipairs(r.macros) do
add(string.format("| %s | %d | %s |",
src_name, m.line, m.name))
end
end
end
add("")
-- Findings by atom (static-analysis)
add("## Static analysis — findings by atom"); add("")
local by_atom = {}
for _, f in ipairs(sa_results.findings or {}) do
by_atom[f.atom] = by_atom[f.atom] or {}
by_atom[f.atom][#by_atom[f.atom] + 1] = f
end
if next(by_atom) == nil then
add("_(no findings)_")
else
for _, a in ipairs(sa_results.atoms or {}) do
local fs = by_atom[a.name]
if fs then
add(string.format("### %s", a.name))
for _, f in ipairs(fs) do
add(string.format("- `[%s] %s`", f.check, f.msg))
end
add("")
end
end
end
-- Errors / Warnings / Info
local function add_findings(label, entries)
add(string.format("## %s", label))
if #entries == 0 then
add("_(none)_")
else
for _, e in ipairs(entries) do
add(string.format("- line %d %s", e.line, e.msg))
end
end
add("")
end
add_findings("Errors", sa_results.errors or {})
add_findings("Warnings", sa_results.warnings or {})
add_findings("Info", sa_results.info or {})
-- Per-atom cycle counts (path-aware)
add("## Per-atom cycle counts (path-aware, best case, no stalls)"); add("")
add("| atom | source | min | max | branches | paths | notes |")
add("|------|--------|-----|-----|----------|-------|-------|")
local sorted = {}
for _, a in ipairs(sa_results.atoms or {}) do sorted[#sorted + 1] = a end
table.sort(sorted, function(x, y)
return ((x.paths or {}).cycles_max or 0) > ((y.paths or {}).cycles_max or 0)
end)
for _, a in ipairs(sorted) do
local p = a.paths or {}
local src_name = a.source_path and source_basename(a.source_path) or ""
local notes = ""
if p.has_loops then notes = notes .. " [loop!]" end
if p.unknown_macros and #p.unknown_macros > 0 then
notes = notes .. " [unknown: " .. table.concat(p.unknown_macros, ", ") .. "]"
end
add(string.format("| %s | %s | %d | %d | %d | %d | %s |",
a.name, src_name,
p.cycles_min or 0, p.cycles_max or 0,
p.branches or 0, p.paths or 0, notes))
end
add("")
-- Per-source scan summary
add("## Per-source scan summary"); add("")
for _, src in ipairs(dir_sources) do
local src_atoms = {}
for _, a in ipairs(sa_results.atoms or {}) do
if a.source_path == src.path then src_atoms[#src_atoms + 1] = a end
end
if #src_atoms > 0 then
local mn, mx = math.huge, -1
for _, a in ipairs(src_atoms) do
local p = a.paths or {}
if (p.cycles_min or 0) < mn then mn = p.cycles_min or 0 end
if (p.cycles_max or 0) > mx then mx = p.cycles_max or 0 end
end
local path_str
if mx > 0 then
path_str = string.format(" cycles=%d..%d", mn, mx)
else
path_str = string.format(" %d cycles", mn)
end
add(string.format("- `%s` — %d atom%s%s",
src.basename, #src_atoms,
#src_atoms == 1 and "" or "s", path_str))
end
end
add("")
return table.concat(lines, "\n") .. "\n"
end
-- ════════════════════════════════════════════════════════════════════════════
-- REPORT_RENDERERS — data-driven report dispatch (one row per file kind)
-- ════════════════════════════════════════════════════════════════════════════
-- `once = true` means render once at the project level (not per-module).
-- `basename(dir_basename)` yields the file's basename for that kind.
-- `gather(ctx, dir, dir_sources [, all_modules])` returns the rendered string.
local REPORT_RENDERERS = {
{
name = "atom_meta_report",
ext = "md",
basename = function(dir_basename) return dir_basename .. ".atom_meta_report" end,
once = false,
gather = function(ctx, dir, dir_sources)
-- Annotations: re-run `annotation.validate()` per source (the existing pattern).
local annot_results = {}
for _, src in ipairs(dir_sources) do
if src.scan then
local r = annotation.validate(ctx, src, nil)
r.source = src.path
annot_results[#annot_results + 1] = r
end
end
-- Static-analysis: read stashed projection (no re-validate).
local dir_basename = dir:match("([^/\\]+)$") or dir
local sa_results = (ctx.shared.corpus.static_analysis_results or {})[dir_basename] or {}
return render_module_meta_report(dir, dir_sources, annot_results, sa_results)
end,
},
{
name = "atoms",
ext = "md",
basename = function(dir_basename) return dir_basename .. ".atoms" end,
once = false,
gather = function(ctx, dir, dir_sources)
return render_module_atoms_md(dir, dir_sources,
ctx.shared.corpus.word_counts or {})
end,
},
{
name = "summary",
ext = "md",
basename = function(_dir_basename) return "atom_meta_report.summary" end,
once = true,
gather = function(_ctx, _dir, _dir_sources, all_modules)
return render_project_summary(all_modules)
end,
},
}
-- ════════════════════════════════════════════════════════════════════════════
-- M — public pass surface
-- ════════════════════════════════════════════════════════════════════════════
local M = {}
--- Run the report pass. Emits 1 `atom_meta_report.summary.md` per build + 2 `atom_meta_report.md` + 2 `atoms.md` files per module (duffle + gte_hello).
--- Reads `corpus.static_analysis_results` (added in Phase 1) to populate per-module findings without re-running validate().
--- @param ctx PassCtx
--- @return PassResult
function M.run(ctx)
local outputs = {}
local corpus = ctx.shared and ctx.shared.corpus
local by_dir = (corpus and corpus.sources_by_dir) or {}
-- `out_path_root`: when the conventional `out_root` is `build/gen` (any spelling — relative, absolute, separator variants).
-- Write the md files to `build/` (parent of `gen/`) instead of nested under `gen/`.
-- Mirrors the `gdb_tape_atoms_runtime.gdb` relocation.
local function ends_with_gen(p)
return type(p) == "string" and (p:match("[/\\]gen[/\\]?$") ~= nil
or p == "build/gen" or p == "build\\gen")
end
local out_root_effective = ends_with_gen(ctx.out_root)
and ctx.out_root:gsub("[/\\]gen[/\\]?$", "")
or ctx.out_root
duffle.ensure_dir(out_root_effective)
-- Aggregator for the project-wide `once = true` summary renderer.
local all_modules = {}
for dir, dir_sources in pairs(by_dir) do
local dir_basename = dir:match("([^/\\]+)$") or dir
-- Per-renderer dispatch for the per-module renderers (once = false).
for _, renderer in ipairs(REPORT_RENDERERS) do
if not renderer.once then
local body = renderer.gather(ctx, dir, dir_sources)
local out_path = out_root_effective .. "/" .. renderer.basename(dir_basename) .. "." .. renderer.ext
duffle.write_file(out_path, body)
outputs[#outputs + 1] = { kind = renderer.name, path = out_path }
end
end
-- For the summary, compute per-module totals once (re-validating annotations per source — same pattern as the meta_report renderer).
local annot_results = {}
for _, src in ipairs(dir_sources) do
if src.scan then
local r = annotation.validate(ctx, src, nil)
r.source = src.path
annot_results[#annot_results + 1] = r
end
end
local n_annot, n_binds, n_macros = 0, 0, 0
for _, r in ipairs(annot_results) do
n_annot = n_annot + #r.annots
n_binds = n_binds + #r.binds
n_macros = n_macros + #r.macros
end
local sa_results = (corpus.static_analysis_results or {})[dir_basename] or {}
all_modules[#all_modules + 1] = {
module = dir_basename,
atoms = #(sa_results.atoms or {}),
annots = n_annot,
binds = n_binds,
macros = n_macros,
findings = #(sa_results.findings or {}),
errors = #(sa_results.errors or {}),
warnings = #(sa_results.warnings or {}),
info = #(sa_results.info or {}),
}
end
-- Project-wide renderer (once = true): write the summary file.
for _, renderer in ipairs(REPORT_RENDERERS) do
if renderer.once then
local body = renderer.gather(ctx, nil, nil, all_modules)
local out_path = out_root_effective .. "/" .. renderer.basename("") .. "." .. renderer.ext
duffle.write_file(out_path, body)
outputs[#outputs + 1] = { kind = renderer.name, path = out_path }
end
end
return { outputs = outputs, errors = {}, warnings = {} }
end
return M
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+128
View File
@@ -0,0 +1,128 @@
--- word_count_eval.lua — Word-counting logic for the tape-atom metaprogram pipeline.
---
--- Two responsibilities:
--- 1. **Public utility** `M.count_token_words(token, wc)`: Used by `passes/offsets.lua`, `passes/annotation.lua`, and other passes.
--- 2. **Pass entry** `M.run(ctx)`: Loads the authored `word_count.metadata.h` into `ctx.shared.corpus.word_counts` for downstream passes.
--- The generated `.macs.h` files are OUTPUT artifacts and are NOT inputs to this pass;
--- Current component counts are owned by `passes/components.lua` (which populates `corpus.word_counts` and `corpus.component_body_index`
--- AFTER computing each current count from the just-built body + `corpus.word_counts`).
---
--- **Canonical contract**:
--- * `ctx.shared.corpus.word_counts` is the count table.
--- * `corpus.word_counts` is the sole count table. Consumers read `corpus.word_counts` directly.
--- * `ctx.shared.components` and `ctx.shared.component_body_index` are NOT created by this pass (projections only).
--- * No `.macs.h` recursive discovery (no `scan_dir`, no scan cache, no `_invalidate_scan_cache`).
---
--- **Conventions**: tabs (1/level), EmmyLua annotations, no regex,
--- Lua 5.3 compatible.
-- ════════════════════════════════════════════════════════════════════════════
-- Module-scope requires + package.path setup
-- ════════════════════════════════════════════════════════════════════════════
-- Bootstrap: load `scripts/duffle_paths.lua` (sets package.path + package.cpath).
-- Uses `debug.getinfo` to find this file's own directory, so it works both standalone and when require'd from the orchestrator.
-- duffle_paths.lua sets package.path then returns `require("duffle")` at the bottom, so the dofile value IS the duffle module.
local _bootstrap_dir = debug.getinfo(1, "S").source:match("^@?(.*[/\\])") or "./"
local duffle = dofile(_bootstrap_dir .. "../duffle_paths.lua")
-- ════════════════════════════════════════════════════════════════════════════
-- Type declarations
-- ════════════════════════════════════════════════════════════════════════════
--- @class WordCounts
--- @field [string] integer -- macro name -> word count
--- @class SourceFile
--- @field path string -- absolute path to the source file
--- @field text string -- the full source text
--- @field dir string -- the directory containing the source
--- @field basename string -- filename without extension
--- @class PassCtx
--- @field sources SourceFile[] -- all source files in the build
--- @field metadata_path string -- path to word_count.metadata.h
--- @field shared table -- cross-pass shared state
--- @field shared.corpus table -- canonical corpus (required)
--- @field shared.corpus.word_counts WordCounts -- canonical count table (populated by this pass)
--- @field out_root string -- output root (e.g. "build/gen")
--- @field project_root string -- project root (e.g. "code/")
--- @field upstream table<string, table> -- per-pass upstream outputs
--- @field flags table -- CLI flags
--- @field verbose boolean -- if true, log diagnostic info
--- @class PassResult
--- @field outputs table[] -- {kind=, path=} entries describing emit files
--- @field errors table[] -- {line=, msg=} entries; build-stops
--- @field warnings table[] -- {line=, msg=} entries; build-succeeds
-- ════════════════════════════════════════════════════════════════════════════
-- Module exports
-- ════════════════════════════════════════════════════════════════════════════
local M = {}
-- ┌────────────────────────────────────────────────────────────────────┐
-- │ Shared utility: count_token_words │
-- └────────────────────────────────────────────────────────────────────┘
--- Count words emitted by a single comma-separated token inside an atom body.
--- For most tokens (regular MIPS instructions) this returns 1.
--- For `mac_X(...)` calls, this returns the resolved word count from `wc` (recursively if needed). For `nop2` etc., returns wc[name].
--- For unknown macros, returns 1 and (optionally) warns.
--- @param token string -- a single token from split_top_level_commas
--- @param wc WordCounts -- the shared word-count table
--- @return integer
function M.count_token_words(token, wc)
local s = duffle.trim(token)
if s == "" then return 0 end
local name, after = duffle.read_ident(s, 1)
if not name then return 1 end
if wc[name] then return wc[name] end
local paren_pos = duffle.skip_ws_and_cmt(s, after)
if s:sub(paren_pos, paren_pos) == "(" then
io.stderr:write(" warning: unknown macro '" .. name .. "', assuming 1 word\n")
end
return 1
end
-- ┌────────────────────────────────────────────────────────────────────┐
-- │ Pass entry: M.run(ctx) — "word-counts" pass │
-- └────────────────────────────────────────────────────────────────────┘
--- Load the authored `word_count.metadata.h` into `ctx.shared.corpus.word_counts`.
--- Generated `.macs.h` files are OUTPUT artifacts and are NOT scanned as inputs.
--- Current component counts are computed and inserted by `passes/components.lua`
--- after the components pass iterates `corpus.source_order` and writes each source-directory's `gen/macs.h` file.
---
--- Contract:
--- * `ctx.shared.corpus` MUST exist (canonical corpus ownership).
--- * `ctx.metadata_path` MUST be a readable file path to the authored `word_count.metadata.h`.
--- * The pass assigns exactly one table to `corpus.word_counts`.
--- Consumers read the corpus-owned table directly.
--- Consumers must read `corpus.word_counts` directly.
--- @param ctx PassCtx
--- @return PassResult
function M.run(ctx)
-- 1. Canonical-corpus ownership gate.
local corpus = ctx.shared and ctx.shared.corpus
if type(corpus) ~= "table" then
error("word_count_eval.run requires ctx.shared.corpus (canonical corpus). The fixture must install the corpus before running this pass.", 0)
end
-- 2. metadata_path gate.
if type(ctx.metadata_path) ~= "string" or ctx.metadata_path == "" then
error("word_count_eval.run requires ctx.metadata_path (path to the authored word_count.metadata.h).", 0)
end
-- 3. Load authored metadata. Generated .macs.h files are NOT scanned
-- (the pass computes their counts from the just-built bodies after disk emission; see passes/components.lua).
local wc = duffle.load_word_counts(ctx.metadata_path)
-- 4. Assign the count table. ONE assignment, no copy. The assignment creates no secondary alias.
corpus.word_counts = wc
return { outputs = {}, errors = {}, warnings = {} }
end
return M
+75
View File
@@ -0,0 +1,75 @@
-- autoexec.lua - pcsx_debug_helper plugin entry point.
-- Packaged in scripts/pcsx_debug_helper.zip. Loaded by pcsx-redux via the -archive CLI flag (see scripts/launch_pcsx_debug.ps1).
--
-- Registers two web handlers for external CLI tools:
-- /api/v1/lua/gte - full GTE state (32 data + 32 control regs + PC)
-- /api/v1/lua/gp - GP state summary (screenshot endpoint + VRAM endpoint refs)
--
-- The GTE handler reads COP2 regs via PCSX.getRegisters().CP2D/CP2C.
-- The pcsx-redux gdb stub doesn't expose COP2, so this is the only way for external tools to see GTE state.
--
-- The GP handler is a thin pointer:
-- pcsx-redux's Lua API exposes only PCSX.GPU.takeScreenShot() (no GPUSTAT, no GP0/GP1 command log, no display state). For richer GP state, the existing web endpoints are the practical path:
-- /api/v1/state/still - PNG screenshot
-- /api/v1/gpu/vram/raw - VRAM raw bytes (1MB)
--
-- Companion: scripts/gdb/gdb_tape_atoms.gdb (covers GPRs + atom-aware stepping).
local function register_handlers()
if not PCSX.WebServer then PCSX.WebServer = {} end
if not PCSX.WebServer.Handlers then PCSX.WebServer.Handlers = {} end
-- ── GTE state ──
PCSX.WebServer.Handlers.gte = function(req)
local r = PCSX.getRegisters()
local out = { "pc=0x" .. string.format("%x", r.pc) }
for i = 0, 31 do
out[#out + 1] = string.format("D[%d]=0x%08x C[%d]=0x%08x",
i, r.CP2D.r[i], i, r.CP2C.r[i])
end
return table.concat(out, "\n")
end
-- ── GP state (pointer to existing endpoints) ──
-- pcsx-redux's Lua GPU API exposes only takeScreenShot(); no GPUSTAT / GP0 / GP1 command log / display state.
-- We point to the existing web endpoints that DO expose those (when the emulator is actually rendering. Paused-at-BP frames won't have a fresh frame).
PCSX.WebServer.Handlers.gp = function(req)
local out = {
"gpu_screenshot_png=http://localhost:8080/api/v1/state/still",
"vram_raw=http://localhost:8080/api/v1/gpu/vram/raw (1MB VRAM)",
"gpustat=NOT_AVAILABLE_VIA_LUA",
"gp_command_log=NOT_AVAILABLE_VIA_LUA (use pcsx-redux Debug > GPU Logger)",
"hint_run_emulator_unpaused_for_screenshot",
}
return table.concat(out, "\n")
end
end
local ok, err = pcall(register_handlers)
if ok then print("[pcsx_debug_helper] handlers registered: gte, gp")
else print("[pcsx_debug_helper] registration failed: " .. tostring(err))
end
-- ── reload handler (Task 6) ──
-- After gte and gp register successfully, load reload.lua through Support.extra.dofile and call its install(pcsx, support).
-- The whole sequence runs inside pcall so a missing zip, missing module table,
-- or throwing install never disturbs the gte and gp handlers already registered above (handler isolation).
--
-- The failure messages are intentionally single-line so the helper's boot log stays scannable.
if type(Support) == "table"
and type(Support.extra) == "table"
and type(Support.extra.dofile) == "function" then
local load_ok, reload_mod = pcall(Support.extra.dofile, "reload.lua")
if load_ok and type(reload_mod) == "table" and type(reload_mod.install) == "function" then
local install_ok, install_err = pcall(reload_mod.install, PCSX, Support)
if install_ok then
print("[pcsx_debug_helper] reload handler registered")
else
print("[pcsx_debug_helper] reload registration failed: " .. tostring(install_err))
end
else
print("[pcsx_debug_helper] reload load failed: " .. tostring(reload_mod))
end
else
print("[pcsx_debug_helper] reload load failed: Support.extra.dofile unavailable")
end
+902
View File
@@ -0,0 +1,902 @@
-- reload.lua - Side-effect-free hot-reload helper for the
-- pcsx_redux_hot_reload track (Task 2). This file owns the HTTP request
-- surface that the launch / reload client targets:
--
-- POST /api/v1/lua/reload?mode=prime&target=hello_camera&path=<encoded-elf>
-- POST /api/v1/lua/reload?mode=elf&target=hello_camera&path=<encoded-elf>
-- POST /api/v1/lua/reload?mode=patch&target=hello_camera&addr=...&hex=...
--
-- This module exposes the public surface used by the contract harness
-- (tests/reload_helper_contract.lua) and the runtime installed by
-- scripts/pcsx_debug_helper/autoexec.lua. The module must not reference
-- the global PCSX table at load time; the host is passed in explicitly
-- through M.new(host) and M.install(pcsx, support).
--
-- Public surface:
-- M.parse_query(query) -> table, nil OR nil, err_string
-- M.json_response(fields) -> string (sorted keys)
-- M.parse_manifest(...) -> Task 3 (real impl uses elf32.lua)
-- M.new(host) -> runtime object (Task 4; stub here)
-- M.install(pcsx, support) -> registers web handler (Task 6; stub here)
--
-- Companion: scripts/pcsx_debug_helper/autoexec.lua.
-- ---------------------------------------------------------------------------
-- Load the shared ELF32 helpers.
--
-- **The bane of this refactor:** the helper VM (PCSX-Redux) does not expose
-- `require` for paths outside the helper zip. The production loader is
-- `Support.extra.dofile("elf32.lua")` — Support.extra.dofile resolves the
-- name against the helper zip's contents (the zip is generated by the
-- build script and includes both `reload.lua` and `elf32.lua` after Task 6).
--
-- The test harness at `tests/reload_helper_contract.lua` loads `reload.lua`
-- via standard Lua `dofile` with an absolute path; it does not install a
-- `Support` object. We detect the runtime context: if `Support.extra.dofile`
-- exists, use it (production path); otherwise fall back to standard `dofile`
-- with an absolute path (test harness path).
-- ---------------------------------------------------------------------------
local function load_elf32()
if type(Support) == "table"
and type(Support.extra) == "table"
and type(Support.extra.dofile) == "function" then
return Support.extra.dofile("elf32.lua")
end
-- Test harness + any other context that supplies standard Lua dofile.
return dofile("C:/projects/Pikuma/ps1/scripts/elf32.lua")
end
local E = load_elf32()
local M = {}
-- ---------------------------------------------------------------------------
-- parse_query(query)
--
-- Parses an application/x-www-form-urlencoded query string into a table.
--
-- Rules (per spec §8 + plan.md Task 2 Step 3):
-- * Each pair is split on the first '='; the key is to the left, the value
-- to the right. A pair without '=' is a malformed_pair.
-- * Percent escapes '%HH' (HH = two hex digits) decode to the corresponding
-- byte. A '%' not followed by two hex digits is a malformed_escape.
-- * '+' decodes to a literal space (applied after percent decode).
-- * A key appearing more than once is a duplicate_key error.
--
-- Returns the parsed table on success. On failure returns nil and a stable
-- error string suitable for the JSON error envelope. An empty / nil query
-- returns an empty table (not an error).
-- ---------------------------------------------------------------------------
local function percent_decode(s)
-- Walk the string once, byte by byte. A '%' must be followed by exactly
-- two hex digits; '+' decodes to ' '; everything else is passed through.
local out = {}
local i = 1
local len = #s
while i <= len do
local c = s:sub(i, i)
if c == "%" then
if i + 2 > len then
return nil -- truncated escape (e.g., '%' at end or '%X')
end
local hex = s:sub(i + 1, i + 2)
local hd1, hd2 = hex:sub(1, 1), hex:sub(2, 2)
-- Validate both characters are hex digits.
if not (hd1:match("[0-9A-Fa-f]") and hd2:match("[0-9A-Fa-f]")) then
return nil -- malformed escape
end
out[#out + 1] = string.char(tonumber(hex, 16))
i = i + 3
else
out[#out + 1] = c
i = i + 1
end
end
return table.concat(out)
end
local function plus_to_space(s)
-- Standalone helper so callers can decode '+' after percent decoding.
return (s:gsub("+", " "))
end
function M.parse_query(query)
if query == nil or query == "" then
return {}, nil
end
local result = {}
local seen = {}
for pair in query:gmatch("[^&]+") do
-- Split on the first '=' only.
local eq = pair:find("=", 1, true)
if not eq then
return nil, "malformed_pair"
end
local raw_key = pair:sub(1, eq - 1)
local raw_value = pair:sub(eq + 1)
-- Percent-decode first, then convert '+' to space. The order matters:
-- a '%2B' should decode to '+' (literal plus), not be re-converted to a
-- space. Per RFC 1866 §8.2.1, '+' is a literal plus in the encoded form
-- only when it represents a space.
local key = percent_decode(raw_key)
if key == nil then
return nil, "malformed_escape"
end
key = plus_to_space(key)
local val = percent_decode(raw_value)
if val == nil then
return nil, "malformed_escape"
end
val = plus_to_space(val)
if seen[key] then
return nil, "duplicate_key"
end
seen[key] = true
result[key] = val
end
return result, nil
end
-- ---------------------------------------------------------------------------
-- json_response(fields)
--
-- Deterministic JSON object encoder. Returns a string. Keys are sorted
-- alphabetically before emission so byte-for-byte equality is testable
-- across runs and across PS1 captures.
--
-- Supported value types: string, number, boolean, nil (encoded as null).
-- Strings escape '\', '"', and the C0 control range (0x00..0x1F). The
-- named escapes use the conventional single-char forms: \\, \", \b, \f,
-- \n, \r, \t. Everything else in 0x00..0x1F is \uXXXX.
-- ---------------------------------------------------------------------------
local function json_escape_string(s)
-- Two passes: first the named escapes, then the catch-all C0 range
-- (%c covers 0x00..0x1F in Lua patterns). Using plain string.gsub
-- with a literal replacement table covers the named escapes; a
-- second gsub handles the rest.
s = s:gsub('[\\"]', {
["\\"] = "\\\\",
['"'] = '\\"',
})
s = s:gsub("\b", "\\b")
s = s:gsub("\f", "\\f")
s = s:gsub("\n", "\\n")
s = s:gsub("\r", "\\r")
s = s:gsub("\t", "\\t")
-- Remaining C0 control characters (0x00..0x1F) become \uXXXX. We
-- intentionally keep the named escapes above (which are already
-- single backslashes in the output) from being re-escaped: gsub on
-- the literal control char bytes doesn't match the backslashes we
-- already inserted.
s = s:gsub("([%c])", function(c)
return string.format("\\u%04x", string.byte(c))
end)
return s
end
function M.json_response(fields)
if type(fields) ~= "table" then
error("json_response: expected table, got " .. type(fields))
end
-- Sort keys for deterministic output. Lua's table.sort is byte-wise
-- and stable for strings; JSON object key order is not significant
-- but tests rely on a fixed order to compare against fixtures.
local keys = {}
for k in pairs(fields) do
keys[#keys + 1] = k
end
table.sort(keys)
local parts = {}
parts[#parts + 1] = "{"
for i = 1, #keys do
local k = keys[i]
if i > 1 then
parts[#parts + 1] = ","
end
parts[#parts + 1] = '"'
parts[#parts + 1] = json_escape_string(k)
parts[#parts + 1] = '":'
local v = fields[k]
local tv = type(v)
if tv == "string" then
parts[#parts + 1] = '"'
parts[#parts + 1] = json_escape_string(v)
parts[#parts + 1] = '"'
elseif tv == "number" then
parts[#parts + 1] = tostring(v)
elseif tv == "boolean" then
parts[#parts + 1] = v and "true" or "false"
elseif v == nil then
parts[#parts + 1] = "null"
else
error("json_response: unsupported value type " .. tv .. " for key " .. tostring(k))
end
end
parts[#parts + 1] = "}"
return table.concat(parts)
end
-- ---------------------------------------------------------------------------
-- ELF32 manifest parser (Task 3).
--
-- Parses a little-endian ELF32 file exposed through a file_adapter that
-- provides read_u8_at/read_u16_at/read_u32_at/read_size. The parser validates the
-- magic, class, data encoding, and machine before reading anything else.
-- It resolves section names through the .shstrtab table and symbols
-- through every SHT_SYMTAB section (and its linked string table).
--
-- The output manifest contains the state ABI the reload gate must
-- preserve plus the addresses the helper writes to the CPU on a reload.
-- Loaded sections (SHF_ALLOC, non-SHT_NOBITS) are recorded so the runtime
-- can reject any ELF whose loaded range overlaps the preserved smem.
--
-- **Refactor:** the format-constant tables + the byte-level walker live in
-- scripts/elf32.lua (loaded above via `load_elf32()`). This module retains
-- only the manifest-specific validation: required symbols, smem size, stack
-- alignment, loaded-section overlap. The net effect is ~80 lines shorter.
--
-- Stable error codes (returned as the second value):
-- bad_magic, unsupported_elf_class, unsupported_elf_data,
-- non_mips_machine, truncated_header, truncated_section_headers,
-- missing_shstrtab, missing_symtab_strtab, missing_smem,
-- missing_data_start, missing_data_end, missing_bss_start,
-- missing_bss_end, missing_stack_top, missing_hot_reload_entry,
-- zero_smem_size, stack_misaligned, stack_out_of_main_ram,
-- section_overlaps_smem, bad_file_adapter
-- ---------------------------------------------------------------------------
-- Convert a KSEG0/KSEG1/physical address to its physical main-RAM offset.
local function to_physical(addr)
if addr >= 0x80000000 and addr < 0x80200000 then
return addr - 0x80000000
elseif addr >= 0xa0000000 and addr < 0xa0200000 then
return addr - 0xa0000000
end
return addr
end
-- Strip KSEG0 / KSEG1 alias from an address and return the physical main-RAM
-- offset. Used by M.elf_reload and M.patch_handler. Returns nil when the
-- address falls outside physical main RAM (0..0x1fffff), KSEG0 main RAM
-- (0x80000000..0x801fffff), or KSEG1 main RAM (0xa0000000..0xa01fffff).
-- Per spec §7 the patch path MUST reject scratchpad (0x1F800000+), BIOS
-- (0x1FC00000+), MMIO, and expansion aliases; this helper centralizes the
-- strip + range check so callers cannot forget the upper bound.
local function strip_kseg(addr)
if type(addr) ~= "number" then return nil end
if addr >= 0x80000000 and addr < 0x80200000 then
return addr - 0x80000000
elseif addr >= 0xa0000000 and addr < 0xa0200000 then
return addr - 0xa0000000
elseif addr >= 0 and addr < 0x200000 then
return addr
end
return nil
end
-- Parse a hex string ("0xHHHH..." or "HHHH...") into a 32-bit unsigned
-- integer. Returns nil + stable error on absent / non-hex / out-of-range.
-- Used for both the patch path's addr/hex query parameters and any other
-- 32-bit hex field the API may add. Accepts up to 8 hex digits.
local function parse_hex_u32(s, missing_err, badhex_err)
if type(s) ~= "string" or #s == 0 then
return nil, missing_err or "missing_hex"
end
local clean = s:match("^0[xX]([0-9A-Fa-f]+)$")
or s:match("^([0-9A-Fa-f]+)$")
if not clean then return nil, badhex_err or "non_hex" end
if #clean > 8 then return nil, badhex_err or "non_hex" end
return tonumber(clean, 16), nil
end
-- Trap on a missing E.* — keeps the existing one-line-error pattern when
-- the helper zip is stale or absent.
local function stack()
io.stderr:write("[reload.parse_manifest] FATAL: scripts/elf32.lua not loaded; aborting\n")
error("elf32 module not loaded")
end
local function parse_manifest_impl(file_adapter, target, path, require_entry)
-- Wrap the body in a pcall so any thrown exception (e.g. a bad
-- adapter method or a malformed section header) surfaces as a
-- parse_error with the message and traceback instead of being lost
-- into the with_busy_guard xpcall as a generic internal_error.
local inner_ok, inner_result, inner_err = pcall(function()
-- Validate the adapter surface. E.validate_adapter returns the same
-- "bad_file_adapter" error code the prior implementation used.
local ok, err = E.validate_adapter(file_adapter)
if not ok then return nil, err end
-- Magic, class, data encoding. E.parse_elf32_headers reads fields at
-- the wire offsets specified in E.ELF32_HEADER.
local hdr, hdr_err = E.parse_elf32_headers(file_adapter)
if not hdr then return nil, hdr_err end
-- Machine check (e.g. EM_MIPS = 8). e_machine is at offset 0x12 (18).
-- The reload helper rejects non-MIPS ELFs before any symbol work.
-- Explicit pass style: E.read_u16(adapter, off). The helper wraps the
-- Support.File adapter once to strip its implicit `self` so the
-- parser shape stays flat-function, not colon-dispatch.
local machine = E.read_u16(file_adapter, 0x12)
if not machine then return nil, "truncated_header" end
if machine ~= E.EM_MIPS then
return nil, "non_mips_machine"
end
-- Walk sections. E.walk_sections also resolves .shstrtab names.
local sections, walk_err = E.walk_sections(file_adapter, hdr)
if not sections then return nil, walk_err end
-- Walk symbols. E.collect_symbols includes both STB_LOCAL and STB_GLOBAL
-- (the live ELF stores smem as a local symbol).
local symbols, sym_err = E.collect_symbols(file_adapter, sections)
if not symbols then return nil, sym_err end
-- Required symbols.
local smem = symbols["smem"]
local data_start = symbols["__data_start"]
local data_end = symbols["__data_end"]
local bss_start = symbols["__bss_start"]
local bss_end = symbols["__bss_end"]
local stack_top_s = symbols["__sp"]
local entry_s = symbols["hot_reload_entry"]
if not smem then return nil, "missing_smem" end
if not data_start then return nil, "missing_data_start" end
if not data_end then return nil, "missing_data_end" end
if not bss_start then return nil, "missing_bss_start" end
if not bss_end then return nil, "missing_bss_end" end
if not stack_top_s then return nil, "missing_stack_top" end
if require_entry and not entry_s then
return nil, "missing_hot_reload_entry"
end
-- Validate smem size.
if smem.size == 0 then
return nil, "zero_smem_size"
end
-- Validate stack alignment and range.
local stack_top = stack_top_s.value
if stack_top % 8 ~= 0 then
return nil, "stack_misaligned"
end
local p = to_physical(stack_top)
if p < 0 or p > 0x1fffff then
return nil, "stack_out_of_main_ram"
end
-- Collect loaded (SHF_ALLOC, non-SHT_NOBITS) sections and check overlap.
local loaded = {}
local smem_lo = smem.value
local smem_hi = smem.value + smem.size
for _, s in ipairs(sections) do
-- bit 1 (SHF_ALLOC = 0x2) of sh_flags. The modulo-4 trick matches
-- the prior implementation; canonicalising on E.SHF_ALLOC would
-- gain readability but lose the exact prior behavior.
local is_alloc = (s.sh_flags % 4) >= 2
if is_alloc and s.sh_type ~= E.SHT_NOBITS and s.sh_size > 0 then
loaded[#loaded + 1] = { name = s.name, addr = s.sh_addr, size = s.sh_size }
local lo = s.sh_addr
local hi = s.sh_addr + s.sh_size
if lo < smem_hi and hi > smem_lo then
return nil, "section_overlaps_smem"
end
end
end
return {
target = target,
elf_path = path,
elf_entry = hdr.e_entry,
smem_addr = smem.value,
smem_size = smem.size,
bss_start = bss_start.value,
bss_end = bss_end.value,
data_start = data_start.value,
data_end = data_end.value,
hot_reload_entry = entry_s and entry_s.value or nil,
stack_top = stack_top,
loaded_sections = loaded,
}
end)
if inner_ok then
return inner_result, inner_err
end
-- pcall captured a thrown error; surface as parse_error with the
-- message + traceback so the caller can render it.
local tb = debug.traceback(inner_result, 2)
local err = {
parse_error = true,
detail = tostring(inner_result),
tb = tb,
}
return nil, err
end
function M.parse_manifest(file_adapter, target, path, require_entry)
if type(E) ~= "table" or type(E.parse_elf32_headers) ~= "function" then
stack()
end
return parse_manifest_impl(file_adapter, target, path, require_entry)
end
-- ---------------------------------------------------------------------------
-- Runtime + dispatch (Task 4)
--
-- M.new(host) returns a runtime object that owns:
-- active -- the most recently primed manifest, or nil
-- busy -- boolean guard; only one request runs at a time
-- host -- the bound host surface (pause / memory_file / open_file
-- / binary_load / invalidate_cache / get_registers)
--
-- runtime:handle(req) parses the query through M.parse_query, validates
-- the mode against a dispatch table, then acquires the busy guard through
-- xpcall so any error inside the handler releases the guard. The response
-- is always a JSON string built by M.json_response.
--
-- M.prime_active and M.elf_reload are the two handler bodies Task 4 ships.
-- prime_active always parses with require_entry=false (Phase 0 binary
-- compatibility). elf_reload always parses with require_entry=true (the
-- new binary must expose hot_reload_entry). Both validate the parsed
-- manifest; elf_reload runs the five-field ABI gate before declaring
-- success. Full host.pause / memory_file / binary_load / invalidate_cache
-- / get_registers sequencing is Task 5.
-- ---------------------------------------------------------------------------
-- Convert a manifest into the JSON-serializable field subset. loaded_sections
-- is excluded because json_response only supports scalars + nil.
local function manifest_to_response(m)
local fields = {
ok = true,
target = m.target,
elf_path = m.elf_path,
elf_entry = m.elf_entry,
smem_addr = m.smem_addr,
smem_size = m.smem_size,
bss_start = m.bss_start,
bss_end = m.bss_end,
data_start = m.data_start,
data_end = m.data_end,
stack_top = m.stack_top,
}
if m.hot_reload_entry then
fields.hot_reload_entry = m.hot_reload_entry
end
return fields
end
-- Open the new ELF through the host and parse its manifest.
-- Returns manifest on success; nil + stable error on failure.
local function parse_manifest_via_host(host, target, path, require_entry)
local adapter = host.open_file(path)
if not adapter then
return nil, "open_file_failed"
end
return M.parse_manifest(adapter, target, path, require_entry)
end
-- prime_active: parse with require_entry=false. Accepts Phase 0 binaries
-- that lack hot_reload_entry. Stores the manifest in runtime.active.
function M.prime_active(runtime, parsed)
local manifest, err = parse_manifest_via_host(
runtime.host, parsed.target, parsed.path, false)
if not manifest then
return M.json_response({ ok = false, error = err, restart_required = true })
end
runtime.active = manifest
return M.json_response(manifest_to_response(manifest))
end
-- elf_reload: full host-driven reload sequence.
--
-- Per conductor/tracks/ps1_pcsx_redux_hot_reload_20260802/spec.md §5 +
-- plan.md Task 5 Step 4. The canonical 11-entry success log is:
--
-- pause, memory_file, state_read, open_new_elf, binary_load,
-- state_restore, invalidate_cache, get_registers, write_sp,
-- write_ra, write_pc
--
-- Sequencing:
--
-- 1. Validate the request (target == active.target, path present).
-- 2. Compute the physical address of `active.smem_addr` via
-- strip_kseg; reject if outside physical main RAM.
-- 3. PARSE PHASE (before pause):
-- a. elf_handle = host.open_file(parsed.path)
-- b. manifest = M.parse_manifest(elf_handle, ..., require_entry=true)
-- c. Run the five-field ABI gate against runtime.active.
-- d. On any rejection here, return BEFORE pause — the runtime
-- has invoked host.open_file once (logging "open_file") and
-- no other host methods.
-- 4. Pause + snapshot:
-- host.pause()
-- mem = host.memory_file()
-- saved = mem:readAtToSlice(active.smem_size, smem_phys)
-- 5. RELOAD PHASE:
-- elf_handle = host.open_new_elf(parsed.path) -- second open
-- loaded = host.binary_load(elf_handle, mem)
-- if loaded == nil then return binary_load_failed
-- 6. Restore state: mem:writeAtMoveSlice(saved, smem_phys)
-- 7. host.invalidate_cache()
-- 8. Rewrite SP / RA / PC through the FFI register pointer.
-- 9. Replace runtime.active last.
-- 10. Return the JSON envelope.
--
-- The two opens are an intentional test-discoverability choice. The
-- PARSE phase uses host.open_file (it is an existing Task 4 surface
-- also used by prime); the RELOAD phase uses host.open_new_elf (a
-- dedicated Task 5 method). In production both methods bind to
-- Support.File.open so the runtime cost is identical to a single open
-- — the distinction lives in the test log for ordering verification.
local function abi_mismatch_response(field, expected, actual)
return M.json_response({
ok = false, error = "state_abi_mismatch", field = field,
expected = expected, actual = actual,
restart_required = true,
})
end
function M.elf_reload(runtime, parsed)
-- 1. Pre-pause request validation. Pure-Lua, no host calls.
if not runtime.active then
return M.json_response({
ok = false, error = "not_primed", restart_required = false })
end
if parsed.target ~= runtime.active.target then
return M.json_response({
ok = false, error = "target_mismatch",
expected = runtime.active.target, actual = parsed.target,
restart_required = true })
end
if type(parsed.path) ~= "string" or parsed.path == "" then
return M.json_response({
ok = false, error = "missing_path",
restart_required = false })
end
-- 2. SMEM range check on `active` (the new ELF has not been
-- parsed yet; the ABI gate below enforces it cannot relocate).
local smem_phys = strip_kseg(runtime.active.smem_addr)
if smem_phys == nil or smem_phys < 0 or smem_phys > 0x1fffff then
return M.json_response({
ok = false, error = "smem_out_of_main_ram",
restart_required = true })
end
-- 3. PARSE PHASE — open + parse + ABI gate. On any rejection here,
-- only host.open_file has been called. Pause and downstream
-- mutations do NOT occur.
local elf_handle_for_parse = runtime.host.open_file(parsed.path)
if not elf_handle_for_parse then
return M.json_response({
ok = false, error = "open_file_failed",
restart_required = true })
end
local manifest, parse_err = M.parse_manifest(
elf_handle_for_parse, parsed.target, parsed.path, true)
if not manifest then
return M.json_response({
ok = false, error = parse_err,
restart_required = true })
end
local active = runtime.active
if manifest.smem_addr ~= active.smem_addr then
return abi_mismatch_response(
"smem_addr", active.smem_addr, manifest.smem_addr)
end
if manifest.smem_size ~= active.smem_size then
return abi_mismatch_response(
"smem_size", active.smem_size, manifest.smem_size)
end
if manifest.bss_start ~= active.bss_start then
return abi_mismatch_response(
"bss_start", active.bss_start, manifest.bss_start)
end
if manifest.bss_end ~= active.bss_end then
return abi_mismatch_response(
"bss_end", active.bss_end, manifest.bss_end)
end
-- 4. Pause + snapshot smem bytes.
runtime.host.pause()
local mem = runtime.host.memory_file()
local saved = mem:readAtToSlice(active.smem_size, smem_phys)
-- 5. RELOAD PHASE — second open for binary_load.
local elf_handle = runtime.host.open_new_elf(parsed.path)
if not elf_handle then
return M.json_response({
ok = false, error = "open_file_failed",
restart_required = true })
end
local loaded = runtime.host.binary_load(elf_handle, mem)
if loaded == nil then
-- Do NOT restore state; PCSX.Binary.load may have partially
-- written RAM. Keep ACTIVE untouched and tell the caller to
-- restart the emulator.
return M.json_response({
ok = false, error = "binary_load_failed",
restart_required = true })
end
-- 6. Restore the smem snapshot over the freshly-loaded code.
mem:writeAtMoveSlice(saved, smem_phys)
-- 7. Flush the CPU instruction cache (.text/.rodata changed).
runtime.host.invalidate_cache()
-- 8. Rewrite SP / RA / PC through the FFI register pointer. The
-- PC write must happen last; the CPU starts consuming
-- instructions at the new PC the moment the emulator resumes.
local regs = runtime.host.get_registers()
regs.GPR.n.sp = manifest.stack_top
regs.GPR.n.ra = 0
regs.pc = manifest.hot_reload_entry
-- 9. Replace ACTIVE last so a failed reload cannot poison the
-- next request's gate.
runtime.active = manifest
-- 10. Return the JSON envelope.
return M.json_response({
ok = true,
target = manifest.target,
elf_path = manifest.elf_path,
elf_entry = manifest.elf_entry,
smem_addr = manifest.smem_addr,
smem_size = manifest.smem_size,
bss_start = manifest.bss_start,
bss_end = manifest.bss_end,
data_start = manifest.data_start,
data_end = manifest.data_end,
hot_reload_entry = manifest.hot_reload_entry,
stack_top = manifest.stack_top,
})
end
-- patch_handler: one-word RAM patch through MemoryAsFile.
--
-- Per spec §7 + plan.md Task 5 Step 5, the order is:
-- 1. Parse addr and hex query parameters
-- 2. Reject non-hex / missing inputs
-- 3. Reject unaligned addresses (addr & 3)
-- 4. Normalize through strip_kseg; reject out-of-main-RAM
-- (scratchpad 0x1F800000+, BIOS 0x1FC00000+, MMIO, expansion)
-- 5. host.pause()
-- 6. mem = host.memory_file()
-- 7. mem:writeU32At(value, physical_offset)
-- 8. host.invalidate_cache()
-- 9. Return JSON envelope ok=true with the requested addr and value.
local function patch_error(err, restart)
return M.json_response({
ok = false, error = err,
restart_required = restart or false,
})
end
function M.patch_handler(runtime, parsed)
local addr_str = parsed.addr
local hex_str = parsed.hex
-- 1. Presence checks.
if type(addr_str) ~= "string" or addr_str == "" then
return patch_error("missing_addr", false)
end
if type(hex_str) ~= "string" or hex_str == "" then
return patch_error("missing_value", false)
end
-- 2. Hex parse.
local addr = parse_hex_u32(addr_str, "missing_addr", "non_hex_addr")
if not addr then
return patch_error(
addr == false and "missing_addr" or "non_hex_addr", false)
end
local value = parse_hex_u32(hex_str, "missing_value", "non_hex_value")
if not value then
return patch_error(
value == false and "missing_value" or "non_hex_value", false)
end
-- 3. Alignment (checked on the canonical KSEG/physical addr).
if addr % 4 ~= 0 then
return patch_error("addr_unaligned", false)
end
-- 4. Range check via strip_kseg (rejects KSEG0 > 0x801fffff, KSEG1 >
-- 0xa01fffff, scratchpad, BIOS, MMIO, expansion, etc.).
local phys = strip_kseg(addr)
if phys == nil then
return patch_error("addr_out_of_main_ram", false)
end
-- 5-8. Pause / write / cache invalidate.
runtime.host.pause()
local mem = runtime.host.memory_file()
mem:writeU32At(value, phys)
runtime.host.invalidate_cache()
-- 9. Return the JSON envelope. Echo the requested address and the
-- value in normalized hex so log captures stay stable across runs.
return M.json_response({
ok = true,
addr = addr_str,
value = "0x" .. string.format("%x", value),
})
end
-- Mode dispatch table. Each handler is invoked with (runtime, parsed).
-- Tasks 5 adds patch (M.patch_handler); the previous placeholder removed.
local DISPATCH = {
prime = M.prime_active,
elf = M.elf_reload,
patch = M.patch_handler,
}
-- Wrap a handler call with the busy guard. The guard is acquired only
-- after the mode is validated, so unknown-mode requests do not deadlock
-- the runtime. xpcall guarantees the guard is released even if the
-- handler throws.
local function with_busy_guard(runtime, fn)
if runtime.busy then
return M.json_response({
ok = false, error = "reload_busy", restart_required = false })
end
runtime.busy = true
-- Capture both the error text and a full Lua traceback so the user
-- can see the actual failing call site instead of a generic
-- "internal_error". debug.traceback("", 2) skips this xpcall frame
-- and the json_response frame so the trace starts at the handler.
local ok, result = xpcall(fn, function(e)
return { msg = tostring(e), tb = debug.traceback("", 2) }
end)
runtime.busy = false
if not ok then
return M.json_response({
ok = false, error = "internal_error",
detail = result.msg, tb = result.tb,
restart_required = true })
end
return result
end
function M.new(host)
if type(host) ~= "table" then
error("M.new: host must be a table, got " .. type(host))
end
local runtime = {
active = nil,
busy = false,
host = host,
}
function runtime:handle(req)
-- 1. Parse the query (M.parse_query returns nil, err on failure).
local query = req and req.urlData and req.urlData.query or ""
local parsed, parse_err = M.parse_query(query)
if not parsed then
return M.json_response({
ok = false, error = parse_err, restart_required = false })
end
-- 2. Validate the mode against the dispatch table.
local mode = parsed.mode
local handler = DISPATCH[mode]
if not handler then
return M.json_response({
ok = false, error = "unknown_mode", restart_required = false })
end
-- 3. Acquire busy and dispatch via xpcall. Mode validation
-- happens BEFORE busy is acquired so unknown-mode requests
-- cannot deadlock the runtime.
return with_busy_guard(self, function()
return handler(self, parsed)
end)
end
return runtime
end
-- Install the reload handler on a PCSX-Redux instance.
--
-- Per plan.md Task 5 Step 5 the adapter binds the canonical host method
-- names to the PCSX-Lua FFI surface:
--
-- pause -> PCSX.pauseEmulator
-- memory_file -> PCSX.getMemoryAsFile
-- open_file -> Support.File.open(path, "READ")
-- binary_load -> PCSX.Binary.load
-- invalidate_cache -> PCSX.invalidateCache
-- get_registers -> PCSX.getRegisters
--
-- The returned closure dispatches each request through M.new(host)'s
-- runtime:handle so the same prime/elf/patch dispatch machinery is used
-- (including the busy guard from Task 4).
--
-- Missing `PCSX.WebServer.Handlers` is created on demand so callers do
-- not have to wire that themselves; if `PCSX` or `Support` is absent a
-- single line is printed and the function returns without registering
-- a handler.
function M.install(pcsx, support)
if type(pcsx) ~= "table" then
print("[reload] install failed: PCSX is not a table")
return
end
if type(support) ~= "table"
or type(support.File) ~= "table"
or type(support.File.open) ~= "function" then
print("[reload] install failed: Support.File.open unavailable")
return
end
if type(pcsx.pauseEmulator) ~= "function" then print("[reload] install failed: PCSX.pauseEmulator missing"); return end
if type(pcsx.getMemoryAsFile) ~= "function" then print("[reload] install failed: PCSX.getMemoryAsFile missing"); return end
if type(pcsx.Binary) ~= "table"
or type(pcsx.Binary.load) ~= "function" then print("[reload] install failed: PCSX.Binary.load missing"); return end
if type(pcsx.invalidateCache) ~= "function" then print("[reload] install failed: PCSX.invalidateCache missing"); return end
if type(pcsx.getRegisters) ~= "function" then print("[reload] install failed: PCSX.getRegisters missing"); return end
-- ---------------------------------------------------------------------------
-- File adapter wrap.
--
-- The production pcsx-redux Support.File wrapper (see
-- toolchain/pcsx-redux/src/lua/fileffi.lua:225-232 + size() around line 203)
-- exposes byte-read methods as colon-syntax closures with camelCase names:
-- readU8At = function(self, pos) ... end
-- readU16At = function(self, pos) ... end
-- readU32At = function(self, pos) ... end
-- size = function(self) ... end
--
-- The ELF32 parser (scripts/elf32.lua) uses an explicit-pass shape with
-- snake_case names:
-- adapter.read_u8_at(off) / adapter.read_u16_at(off) /
-- adapter.read_u32_at(off) / adapter.read_size()
--
-- The install boundary wraps the Support.File return value in a thin
-- adapter whose methods forward to the production closures, stripping
-- the implicit `self` and re-exporting the names the parser validates.
-- Without this wrap, E.validate_adapter returns "bad_file_adapter"
-- because adapter.read_u8_at / read_u16_at / read_u32_at / read_size
-- are not present on the raw Support.File return.
local function wrap_file(f)
return {
read_u8_at = function(off) return f:readU8At(off) end,
read_u16_at = function(off) return f:readU16At(off) end,
read_u32_at = function(off) return f:readU32At(off) end,
read_size = function() return f:size() end,
}
end
local host = {
pause = function() pcsx.pauseEmulator() end,
memory_file = function() return pcsx.getMemoryAsFile() end,
open_file = function(path) return wrap_file(support.File.open(path, "READ")) end,
-- open_new_elf returns the raw Support.File object because the
-- RELOAD phase passes it directly to PCSX.Binary.load which
-- expects a real File (with readAt / size), NOT the elf32
-- parser adapter (read_u8_at / read_u16_at / read_u32_at /
-- read_size). Wrapping it in the adapter here triggers the
-- binffi.lua "Expected a File object as first argument" error.
open_new_elf = function(path) return support.File.open(path, "READ") end,
binary_load = function(elf, mem) return pcsx.Binary.load(elf, mem) end,
invalidate_cache = function() pcsx.invalidateCache() end,
get_registers = function() return pcsx.getRegisters() end,
}
local runtime = M.new(host)
if type(pcsx.WebServer) ~= "table" then pcsx.WebServer = {} end
if type(pcsx.WebServer.Handlers) ~= "table" then pcsx.WebServer.Handlers = {} end
pcsx.WebServer.Handlers.reload = function(req)
return runtime:handle(req)
end
print("[reload] handler installed: reload")
end
return M
+758
View File
@@ -0,0 +1,758 @@
--- ps1_meta.lua — Orchestrator entry point for the tape-atom metaprogram.
---
--- Dispatches to pass modules under `scripts/passes/`, resolving dependencies topologically (Kahn's algorithm + cycle detection).
---
--- Architecture:
--- - PASSES table: Declarative dep graph (data, not code).
--- - FLAG_HANDLERS table: Maps CLI flags to handlers.
--- - parse_args → build_ctx (resolves unity/direct includes or exact sources) → topo_sort → dispatch_passes.
--- - The first pass in the dep graph is `scan-source` (see `passes/scan_source.lua`).
--- It calls `duffle.scan_source` once per source to produce the fat `SourceScan` payload, which is attached to each `src.scan`.
--- Every other pass that reads source structure depends on `scan-source` and consumes `src.scan` as a read-only.
---
-- ════════════════════════════════════════════════════════════════════════════
-- Module-scope requires + package.path setup
-- ════════════════════════════════════════════════════════════════════════════
-- Bootstrap: load `duffle_paths.lua` via this script's own path.
-- Use `arg[0]` when this file is the entry script (`arg[0]` ends in "ps1_meta.lua");
-- fall back to `debug.getinfo(1, "S").source` when this file is being dofile()'d or require()'d (in which case `arg[0]` is the *caller's* path).
-- That single statement: (a) sets `package.path` + `package.cpath`, (b) at the bottom returns `require("duffle")`.
-- So the dofile's return value is the duffle module.
local _is_entry_script = arg and arg[0] and arg[0]:match("ps1_meta%.lua$") ~= nil
local _bootstrap_src
if _is_entry_script then
_bootstrap_src = arg[0]
else
-- debug.getinfo(1, "S").source returns "@<path>" for the current chunk;
-- strip the leading "@" so the directory match works in both cases.
_bootstrap_src = debug.getinfo(1, "S").source:sub(2)
end
local duffle = dofile((_bootstrap_src:match("(.*[/\\])") or "./") .. "duffle_paths.lua")
-- ════════════════════════════════════════════════════════════════════════════
-- Constants
-- ════════════════════════════════════════════════════════════════════════════
-- Exit codes (per the --help text and the post-build summary convention).
local EXIT_OK = 0
local EXIT_VALIDATION_ERRORS = 1
local EXIT_INTERNAL_ERROR = 2
-- Default --out-root value if not provided.
local DEFAULT_OUT_ROOT = "build/gen"
-- Sentinel for "all passes" in `PASS_FLAG_TO_NAME`. Distinguishes `--all` from the per-pass flags (which map to individual pass names).
local ALL_PASSES_SENTINEL = "__all__"
-- Sentinel key for the pass-flag dispatcher in `FLAG_HANDLERS`.
-- The actual pass names are looked up via `PASS_FLAG_TO_NAME`, not direct dispatch, so this key never matches a real flag.
local PASS_FLAG_DISPATCH_KEY = "__pass__"
-- ════════════════════════════════════════════════════════════════════════════
-- Type declarations
-- ════════════════════════════════════════════════════════════════════════════
--- @class PassDescriptor
--- @field module string -- Module name passed to require()
--- @field kind string -- "shared" | "header-output" | "validation" | "diagnostic" | "report"
--- -- Report severity is independent from process exit policy (see PASS_KIND_STOP_ON_ERROR).
--- @field deps string[] -- Names of upstream passes
--- @field groups string[]? -- OPTIONAL build-phase groups this pass is a root of (e.g. { "pre-link" }, { "post-link" }); absent ⇒ dependency-only
--- @class SourceFile
--- @field path string -- Absolute path to the source file
--- @field text string -- Full source text
--- @field dir string -- Directory containing the source
--- @field basename string -- Filename without extension
--- @class PassCtx
--- @field metadata_path string -- Path to word_count.metadata.h
--- @field shared table -- Cross-pass shared state
--- @field shared.corpus table -- Authored-source/project projection
--- @field out_root string -- Output root (e.g. "build/gen")
--- @field project_root string -- PS1 repository root
--- @field flags table -- CLI flags + per-pass stash
--- @field verbose boolean -- If true, log diagnostic info
--- @class Finding
--- @field line integer -- Source line (or 0 for pass-level)
--- @field msg string -- Finding message
--- @class PassResult
--- @field outputs PassOutputEntry[] -- Emitted file paths
--- @field errors Finding[] -- Build-stops (per-pass kind policy)
--- @field warnings Finding[] -- Informational
--- @class ParsedArgs
--- @field requested_set string[] -- Pass names to run (explicit --all expanded)
--- @field sources string[] -- Exact --source values, retained in CLI order
--- @field unity_root string|nil -- --unity-root value; mutually exclusive with sources
--- @field metadata string -- --metadata value
--- @field out_root string -- --out-root value (default "build/gen")
--- @field project_root string -- PS1 repository root (derived from metadata by default)
--- @field verbose boolean -- If true, log diagnostic info
-- ════════════════════════════════════════════════════════════════════════════
-- PASSES Table
-- ════════════════════════════════════════════════════════════════════════════
-- Build-phase groups: Each PASSES row may declare membership in one or more named groups via `groups = { ... }`.
-- The CLI flags --pre-link and --post-link request the *roots* of their group; topo_sort then closes transitive dependencies from those roots,
-- and dispatch_passes runs every pass in the resulting closure without phase-filtering.
--
-- A row without a `groups` entry is dependency-only: it runs only when a transitive dep requests it,
-- but it remains directly requestable through its explicit CLI flag (e.g. --atoms-source-map, --scan-source).
local PASSES = {
["scan-source"] = {
module = "passes.scan_source",
kind = "shared", deps = {},
},
["word-counts"] = {
module = "passes.word_count_eval",
kind = "shared", deps = {},
},
components = {
module = "passes.components",
kind = "header-output",
deps = {"scan-source", "word-counts"},
},
["emission-model"] = {
module = "passes.emission_model",
kind = "validation",
deps = {"components"},
},
annotation = {
module = "passes.annotation",
kind = "validation",
deps = {"scan-source", "word-counts"},
},
offsets = {
module = "passes.offsets",
kind = "header-output",
deps = {"scan-source", "word-counts", "components", "emission-model"},
groups = { "pre-link" },
},
["static-analysis"] = {
module = "passes.static_analysis",
-- "diagnostic" — every `error`/`warning` finding is written to the report file;
-- The orchestrator does NOT exit non-zero on these findings (see PASS_KIND_STOP_ON_ERROR).
-- Report severity is independent from process exit policy.
kind = "diagnostic",
deps = {"scan-source", "word-counts", "components", "emission-model"},
},
["atoms-source-map"] = {
module = "passes.atoms_source_map",
kind = "header-output",
deps = {"word-counts", "components", "emission-model"},
},
["dwarf-injection"] = {
module = "passes.dwarf_injection",
kind = "shared",
deps = {"scan-source", "atoms-source-map"},
groups = { "post-link" },
},
report = {
module = "passes.report",
kind = "report",
deps = {"annotation", "static-analysis", "atoms-source-map"}, -- +atoms-source-map (consolidated-report-files refactor, 2026-07-26)
groups = { "pre-link" },
},
}
-- ────────────────────────────────────────────────────────────────────────────
-- Phase-root selection: Derive the sorted set of roots belonging to a named build-phase group, then append them to `args.requested_set`.
-- topo_sort closes the transitive deps from there; dispatch_passes runs every resolved pass without phase-filtering.
-- ────────────────────────────────────────────────────────────────────────────
--- @param group_name string -- Build-phase group ("pre-link" | "post-link")
--- @return string[] -- Sorted root pass names belonging to that group
local function roots_for_group(group_name)
local names = {}
for name, pass in pairs(PASSES) do
if pass.groups then
for _, g in ipairs(pass.groups) do
if g == group_name then
names[#names + 1] = name
break
end
end
end
end
table.sort(names)
return names
end
--- Append every root belonging to `group_name` to `args.requested_set`.
--- Errors loudly if no PASSES row declares the group, so a typo'd or future-removed group name
--- cannot silently fall through to pre-link (or any other default) and dispatch nothing.
--- @param args ParsedArgs
--- @param group_name string
local function request_roots_for_group(args, group_name)
local roots = roots_for_group(group_name)
if #roots == 0 then
error(string.format("ps1_meta: build-phase group %q has zero roots in PASSES; check PASSES rows for a `groups = { %q }` field"
, group_name, group_name))
end
for _, name in ipairs(roots) do
args.requested_set[#args.requested_set + 1] = name
end
end
-- Pass-kind taxonomy: Which kinds stop the build on errors?
--
-- Report severity is independent from process exit policy.
-- A "diagnostic" pass still writes every `error`/`warning` finding into its report file,
-- but `report_validation_errors` returns early for non-stopping kinds, so nothing is printed to stderr and the orchestrator does not exit non-zero.
-- Adding a new pass kind requires listing it here explicitly; An unknown kind must not silently fall back to "true".
local PASS_KIND_STOP_ON_ERROR = {
["shared"] = false,
["header-output"] = true,
["validation"] = true,
["diagnostic"] = false,
["report"] = false,
}
-- Closed set of CLI flags -> pass names.
-- Per-pass flags (e.g. --word-counts); phase flags (--pre-link, --post-link, --all) are within FLAG_HANDLERS because they own side effects or invoke group-derivation logic.
-- dwarf-injection is *also* a per-pass opt-in flag, but its selection + opt-in state are both owned by the explicit FLAG_HANDLERS entry below
-- (it sets args.flags.dwarf_injection and appends "dwarf-injection" to requested_set), so it is intentionally absent from this table.
local PASS_FLAG_TO_NAME = {
["--word-counts"] = "word-counts",
["--components"] = "components",
["--validate"] = "annotation",
["--offsets"] = "offsets",
["--static-analysis"] = "static-analysis",
["--atoms-source-map"] = "atoms-source-map",
["--report"] = "report",
["--scan-source"] = "scan-source",
["--all"] = ALL_PASSES_SENTINEL,
}
--- Append every pass name to args.requested_set.
--- Names are derived from PASSES (no parallel name list); used by --all and by any caller that wants the full closure.
--- @param args ParsedArgs
local function request_all_passes(args)
local names = {}
for name in pairs(PASSES) do names[#names + 1] = name end
table.sort(names)
for _, n in ipairs(names) do
args.requested_set[#args.requested_set + 1] = n
end
end
-- Per-flag handlers. Each handler takes (args, argv, arg_idx) and returns the new arg_idx (so multi-arg flags like --source FILE advance it).
-- Returning nil + os.exit() handles termination flags (--help).
local FLAG_HANDLERS = {}
-- ════════════════════════════════════════════════════════════════════════════
-- CLI parsing
-- ════════════════════════════════════════════════════════════════════════════
--- Print the CLI usage to stdout and exit 0.
local function print_help()
io.write([[
ps1_meta.lua - Tape-atom metaprogram orchestrator
USAGE:
ps1_meta.lua [PASS_FLAGS] [COMMON_FLAGS]
PASS_FLAGS:
Pick a phase or one-or-more individual passes:
--pre-link [phase; default] Run the pre-link group + transitive deps.
The root set is data-driven from each PASSES row's groups` field; no parallel name list is maintained.
--post-link [phase] Run the post-link group + transitive deps.
Requires --elf. Sets --gdb-runtime and --dwarf-injection opt-in flags as well.
--all Select every row of the PASSES table. Pass-local opt-in guards remain active, so --dwarf-injection still requires
--elf and --gdb-runtime still requires a runtime emission.
Or pick any subset:
--scan-source Scan sources into the fat SourceScan payload
--word-counts Load metadata.h + scan for existing .macs.h
--components Generate <srcdir>/gen/macs.h (per-directory aggregation)
--validate Run atom annotation DSL validation
--offsets Generate <srcdir>/gen/offsets.h (per-directory aggregation)
--atoms-source-map Generate <basename>.atoms.sourcemap.txt per source
--dwarf-injection [opt-in] Select the post-link dwarf-injection pass + set the opt-in flag. Requires --elf.
--static-analysis Static analysis: GTE pipeline-fill, mac_yield, ABI handoff, cycle budget
--report Render per-project summary
COMMON_FLAGS:
--unity-root FILE Unity source root: load root + direct quoted authored includes only. Mutually exclusive with --source.
--source FILE Exact source file to process (repeatable, never expands includes). Mutually exclusive with --unity-root.
--metadata PATH Path to metadata.h (required)
--out-root DIR Output root for reports (default: build/gen)
--project-root DIR PS1 repository root (default: derived from <repo>/code/duffle/word_count.metadata.h)
--gdb-runtime Also emit <out_root>/gdb_tape_atoms_runtime.gdb (post-link, requires --elf)
--elf PATH Path to linked .elf (for --gdb-runtime / --dwarf-injection)
--verbose Print per-pass debug output
--help Show this help and exit
EXIT CODES:
0 All requested passes succeeded
1 Validation errors found
2 Metaprogram internal error
EXAMPLES:
ps1_meta.lua --pre-link --metadata code/duffle/word_count.metadata.h --unity-root code/gte_hello/hello_gte.c
ps1_meta.lua --post-link --metadata code/duffle/word_count.metadata.h --unity-root code/gte_hello/hello_gte.c --elf build/hello_gte.elf
ps1_meta.lua --all --metadata metadata.h --source code/foo.c --source code/bar.c
]])
end
local FLAG_VALUE_NAMES = {
["--source"] = "FILE",
["--unity-root"] = "FILE",
["--metadata"] = "PATH",
["--out-root"] = "DIR",
["--project-root"] = "DIR",
["--elf"] = "PATH",
}
local function require_flag_value(argv, arg_idx, flag)
local value = argv[arg_idx + 1]
local next_known = type(value) == "string"
and (FLAG_HANDLERS[value] ~= nil or PASS_FLAG_TO_NAME[value] ~= nil)
if value == nil or next_known then
io.stderr:write("ps1_meta: " .. flag .. " requires " .. FLAG_VALUE_NAMES[flag] .. "\n")
os.exit(EXIT_INTERNAL_ERROR)
end
return value, arg_idx + 1
end
-- Per-flag handlers. Each takes (args, argv, arg_idx) and returns the new arg_idx (so multi-arg flags like --source FILE advance it).
-- Termination flags like --help call os.exit() instead.
-- Populated AFTER print_help so the --help handler can reference it as an upvalue (Lua resolves locals at closure-call time,
-- but if the closure is defined before the local, it falls back to _G).
FLAG_HANDLERS["--help"] = function(args) print_help(); os.exit(0) end
FLAG_HANDLERS["--verbose"] = function(args) args.verbose = true end
FLAG_HANDLERS["--source"] = function(args, argv, arg_idx)
local value, value_idx = require_flag_value(argv, arg_idx, "--source")
args.sources[#args.sources + 1] = value
return value_idx
end
FLAG_HANDLERS["--unity-root"] = function(args, argv, arg_idx)
local value, value_idx = require_flag_value(argv, arg_idx, "--unity-root")
args.unity_root = value
return value_idx
end
FLAG_HANDLERS["--metadata"] = function(args, argv, arg_idx)
local value, value_idx = require_flag_value(argv, arg_idx, "--metadata")
args.metadata = value
return value_idx
end
FLAG_HANDLERS["--out-root"] = function(args, argv, arg_idx)
local value, value_idx = require_flag_value(argv, arg_idx, "--out-root")
args.out_root = value
return value_idx
end
FLAG_HANDLERS["--project-root"] = function(args, argv, arg_idx)
local value, value_idx = require_flag_value(argv, arg_idx, "--project-root")
args.project_root = value
return value_idx
end
-- Per-pass stash flags. Read by `passes/atoms_source_map.lua` to opt into the post-link gdb-runtime emission.
-- Same shape as the existing per-flag handlers. mutates `args.flags` (which propagates into `ctx.flags`).
FLAG_HANDLERS["--gdb-runtime"] = function(args)
args.flags = args.flags or {}
args.flags.gdb_runtime = true
end
FLAG_HANDLERS["--elf"] = function(args, argv, arg_idx)
local value, value_idx = require_flag_value(argv, arg_idx, "--elf")
args.flags = args.flags or {}
args.flags.elf_path = value
return value_idx
end
-- Enable DWARF injection (default OFF). Opts in to the post-link pass and sets the flag in one shot.
-- The explicit handler below owns both selection and opt-in state, so --dwarf-injection is intentionally absent from PASS_FLAG_TO_NAME.
FLAG_HANDLERS["--dwarf-injection"] = function(args)
args.flags = args.flags or {}
args.flags.dwarf_injection = true
args.requested_set[#args.requested_set + 1] = "dwarf-injection"
end
-- Build-phase flags: --pre-link and --post-link request the roots of their declared groups (see roots_for_group).
-- topo_sort closes transitive deps from those roots; dispatch_passes runs every pass in the resolved closure without phase-filtering.
FLAG_HANDLERS["--pre-link"] = function(args)
request_roots_for_group(args, "pre-link")
end
-- Batch post-link phase: gdb-runtime + dwarf-injection in one luajit cold start.
-- Sets the same opt-in flags as --gdb-runtime + --dwarf-injection and selects the post-link build-phase group.
-- elf is required; parse_args enforces it after all flags are parsed.
FLAG_HANDLERS["--post-link"] = function(args)
args.flags = args.flags or {}
args.flags.gdb_runtime = true
args.flags.dwarf_injection = true
request_roots_for_group(args, "post-link")
end
-- `--dwarf-injection` also emits atom-local debug data.
-- Pass-flag handler. Reads the closed-set table, expands --all, appends to requested_set.
FLAG_HANDLERS[PASS_FLAG_DISPATCH_KEY] = function(args, a)
local name = PASS_FLAG_TO_NAME[a]
if name == ALL_PASSES_SENTINEL then
request_all_passes(args)
return
end
args.requested_set[#args.requested_set + 1] = name
end
--- Parse argv into a structured table. Validates against a closed enum.
--- @param argv string[]
--- @return ParsedArgs
local function parse_args(argv)
local args = {
requested_set = {},
sources = {},
unity_root = nil,
metadata = nil,
out_root = DEFAULT_OUT_ROOT,
project_root = nil,
verbose = false,
}
local pos = 1
while pos <= #argv do
local a = argv[pos]
local handler = FLAG_HANDLERS[a]
if handler then
pos = handler(args, argv, pos) or pos
elseif PASS_FLAG_TO_NAME[a] then
FLAG_HANDLERS[PASS_FLAG_DISPATCH_KEY](args, a)
else
io.stderr:write("ps1_meta: unknown flag '" .. a .. "'\n")
io.stderr:write("Run with --help for usage.\n")
os.exit(EXIT_INTERNAL_ERROR)
end
pos = pos + 1
end
-- Default: --pre-link if no explicit pass flags were given.
-- The first invocation of a build is always pre-link, so this avoids silently also invoking post-link work in builds without an ELF artifact.
if #args.requested_set == 0 then request_roots_for_group(args, "pre-link") end
if not args.metadata then
io.stderr:write("ps1_meta: --metadata PATH is required\n")
os.exit(EXIT_INTERNAL_ERROR)
end
-- `<repo>/code/duffle/word_count.metadata.h` is the canonical metadata location.
-- `project_root` names `<repo>`; the resolver derives `<project_root>/code` separately.
if not args.project_root then
local metadata_dir = duffle.dirname(duffle.normalize_path(args.metadata))
local code_root = duffle.dirname(metadata_dir)
args.project_root = duffle.dirname(code_root)
else
args.project_root = duffle.normalize_path(args.project_root)
end
local has_unity = type(args.unity_root) == "string" and args.unity_root ~= ""
if has_unity and #args.sources > 0 then
io.stderr:write("ps1_meta: --unity-root FILE and --source FILE are mutually exclusive\n")
os.exit(EXIT_INTERNAL_ERROR)
end
if not has_unity and #args.sources == 0 then
io.stderr:write("ps1_meta: either --unity-root FILE or at least one --source FILE is required\n")
os.exit(EXIT_INTERNAL_ERROR)
end
-- Post-link opt-ins (--gdb-runtime, --dwarf-injection) write output that depends on the linked ELF.
-- Without --elf the metaprogram can't satisfy those requests, so refuse loud and early.
-- This covers the explicit --post-link batch, --dwarf-injection by itself, and --gdb-runtime by itself.
local flags = args.flags or {}
local elf_path = flags.elf_path
local has_elf = type(elf_path) == "string" and #elf_path > 0
local post_links = flags.gdb_runtime or flags.dwarf_injection
if post_links and not has_elf then
io.stderr:write("ps1_meta: --elf PATH is required for post-link output\n")
os.exit(EXIT_INTERNAL_ERROR)
end
return args
end
-- ════════════════════════════════════════════════════════════════════════════
-- Build ctx from parsed args
-- ════════════════════════════════════════════════════════════════════════════
--- Build the PassCtx from parsed args. Exact mode opens only the repeated `--source` inputs;
--- unity mode delegates direct-include resolution to duffle.resolve_source_corpus`.
--- Scanning remains pass-owned (`src.scan`).
--- @param args ParsedArgs
--- @return PassCtx
local function build_ctx(args)
local normalized_project_root = duffle.normalize_path(args.project_root)
local project_root = normalized_project_root
local project_root_is_absolute = normalized_project_root:match("^%a:/")
or normalized_project_root:sub(1, 2) == "//"
or normalized_project_root:sub(1, 1) == "/"
if not project_root_is_absolute then
-- canonical_path_key validates ordinary relative paths and rejects drive-relative paths before the absolute-path rewrite is performed.
duffle.canonical_path_key(normalized_project_root)
project_root = duffle.normalize_path(duffle.to_absolute_path(normalized_project_root))
else
-- Do not route POSIX/UNC/drive-absolute paths through to_absolute_path.
duffle.canonical_path_key(project_root)
end
local resolution
if args.unity_root then
local ok_resolve, resolved = pcall(duffle.resolve_source_corpus, {
unity_root = args.unity_root,
project_root = project_root,
})
if not ok_resolve then
io.stderr:write("ps1_meta: cannot resolve --unity-root " .. tostring(args.unity_root) .. ": " .. tostring(resolved) .. "\n")
os.exit(EXIT_INTERNAL_ERROR)
end
resolution = resolved
else
local source_order = {}
local sources_by_path = {}
local resolver = {
resolved = {},
skipped = {},
shadowed = {},
}
for _, input_path in ipairs(args.sources) do
local path = duffle.normalize_path(input_path)
local key_ok, key_or_error = pcall(duffle.canonical_path_key, path)
if not key_ok then
error("ps1_meta: invalid --source " .. input_path .. ": " .. tostring(key_or_error), 0)
end
local file = io.open(path, "r")
if not file then
io.stderr:write("ps1_meta: cannot open --source " .. input_path .. "\n")
os.exit(EXIT_INTERNAL_ERROR)
end
local text = file:read("*a")
file:close()
local source = {
path = path,
text = text,
dir = duffle.dirname(path),
basename = duffle.basename_no_ext(path),
}
source_order[#source_order + 1] = source
local key = key_or_error
if not sources_by_path[key] then sources_by_path[key] = source end
resolver.resolved[#resolver.resolved + 1] = {
include_path = path,
include_text = nil,
root_source = nil,
root_line = nil,
candidate_a = path,
candidate_b = nil,
selected_path = path,
disposition = "exact",
}
end
resolution = {
unity_root = nil,
project_root = project_root,
code_root = duffle.normalize_path(project_root .. "/code"),
source_order = source_order,
sources_by_path = sources_by_path,
sources_by_dir = duffle.group_sources_by_dir(source_order),
resolver = resolver,
}
end
local corpus = {
unity_root = resolution.unity_root,
project_root = resolution.project_root,
code_root = resolution.code_root,
source_order = resolution.source_order,
sources_by_path = resolution.sources_by_path,
sources_by_dir = resolution.sources_by_dir,
atoms_by_name = {},
binds_by_name = {},
atom_infos = {},
register_alias_registry = {},
type_name_registry = {},
atom_views = {},
atom_ctxs = {},
atom_phases = {},
word_counts = {},
components = {},
component_body_index = {},
collisions = {},
resolver = resolution.resolver,
}
local ctx = {
metadata_path = args.metadata,
shared = { corpus = corpus },
out_root = args.out_root,
project_root = corpus.project_root,
flags = args.flags or {},
verbose = args.verbose,
}
-- Source records and directory buckets are owned by the corpus.
-- Consumers read `corpus.source_order` and `corpus.sources_by_dir` directly.
-- The corpus is the sole source of truth for source records and module grouping; `ctx` only holds per-pass execution state.
return ctx
end
-- ════════════════════════════════════════════════════════════════════════════
-- Topological sort (Kahn's algorithm + cycle detection)
-- ════════════════════════════════════════════════════════════════════════════
--- Topologically sort the requested pass set, augmented with all transitive deps.
--- Detects cycles and errors out with details.
--- @param passes table<string, PassDescriptor>
--- @param requested_set string[]
--- @return string[] -- execution order
---
--- Dependency closure, in-degree calculation, queue seeding, and sorting are local blocks.
--- Keeping these blocks local makes the topological sort self-contained.
local function topo_sort(passes, requested_set)
-- Dependency closure: include every pass transitively required by `requested_set`.
local needed = {}
for _, name in ipairs(requested_set) do needed[name] = true end
local changed = true
while changed do
changed = false
for name, _ in pairs(needed) do
local pass = passes[name]
if not pass then error("unknown pass '" .. name .. "' requested") end
for _, dep in ipairs(pass.deps) do
if not needed[dep] then
needed[dep] = true
changed = true
end
end
end
end
-- In-degree calculation: count each needed pass's needed dependencies.
local in_degree = {}
for name, _ in pairs(needed) do in_degree[name] = 0 end
for name, _ in pairs(needed) do
for _, dep in ipairs(passes[name].deps) do
if needed[dep] then
in_degree[name] = in_degree[name] + 1
end
end
end
-- Ready-queue seeding: add zero-in-degree passes in deterministic order.
local ready = {}
for name, deg in pairs(in_degree) do
if deg == 0 then ready[#ready + 1] = name end
end
table.sort(ready)
-- Ready-queue drain: decrement dependents when each pass is emitted.
-- Newly-zero-degree passes are inserted back into the ready queue (kept sorted).
local order = {}
while #ready > 0 do
local just_finished = table.remove(ready, 1)
order[#order + 1] = just_finished
for name, _ in pairs(needed) do
if name ~= just_finished then
for _, dep in ipairs(passes[name].deps) do
if dep == just_finished then
in_degree[name] = in_degree[name] - 1
if in_degree[name] == 0 then
ready[#ready + 1] = name
table.sort(ready)
end
end
end
end
end
end
-- Cycle detection: if `order` doesn't include all needed passes, some are stuck with in_degree > 0
-- (the cycle closed on itself before Kahn could process them).
-- Without this check, a fully-closed cycle (e.g. A -> B -> A) would silently return an empty order list, leaving the orchestrator to dispatch nothing.
local needed_count = 0
for _ in pairs(needed) do needed_count = needed_count + 1 end -- count hash entries; Lua's #t doesn't work
if #order ~= needed_count then
for name, deg in pairs(in_degree) do
if deg > 0 then
error("dependency cycle detected involving pass '" .. name .. "'")
end
end
end
return order
end
-- ════════════════════════════════════════════════════════════════════════════
-- Main Orchestrator
-- ════════════════════════════════════════════════════════════════════════════
--- (internal) If the pass's kind is in PASS_KIND_STOP_ON_ERROR and it reported errors, write each error to stderr.
--- Returns true if any validation errors were reported.
--- @param pass_name string
--- @param pass PassDescriptor
--- @param result PassResult
--- @return boolean
local function report_validation_errors(pass_name, pass, result)
local has_errors = result.errors and #result.errors > 0
if not (has_errors and PASS_KIND_STOP_ON_ERROR[pass.kind]) then return false end
for _, e in ipairs(result.errors) do
io.stderr:write(string.format("[%s] line %d: %s\n", pass_name, e.line or 0, e.msg or ""))
end
return true
end
--- (internal) Run each pass in `order` in topological sequence.
--- @param ctx PassCtx
--- @param order string[]
--- @return boolean -- true if any validation errors were reported
local function dispatch_passes(ctx, order)
local had_errors = false
for _, pass_name in ipairs(order) do
local pass = PASSES[pass_name]
local mod = require(pass.module)
local result = mod.run(ctx)
if report_validation_errors(pass_name, pass, result) then
had_errors = true
end
end
return had_errors
end
--- Main entry point. Runs the requested passes in dep-topological order.
--- @param argv string[]
local function main(argv)
local ok, err = pcall(function()
local args = parse_args(argv)
local ctx = build_ctx(args)
local requested = args.requested_set
local closed = topo_sort(PASSES, requested)
local had_errors = dispatch_passes(ctx, closed)
if had_errors then os.exit(EXIT_VALIDATION_ERRORS) end
end)
if not ok then
io.stderr:write("[ps1_meta] internal error: " .. tostring(err) .. "\n")
os.exit(EXIT_INTERNAL_ERROR)
end
os.exit(EXIT_OK)
end
-- Module export for in-process consumers (tests that dofile this script).
-- The conditional `main(...)` call below only fires when this file is invoked as the entry script (arg[0] ends in "ps1_meta.lua");
-- in dofile() mode (test's arg[0] does not match), main() is skipped and the chunk returns `_M` to the caller.
local _M = {
PASSES = PASSES,
PASS_KIND_STOP_ON_ERROR = PASS_KIND_STOP_ON_ERROR,
parse_args = parse_args,
build_ctx = build_ctx,
}
if arg and arg[0] and arg[0]:match("ps1_meta%.lua$") then
main({...})
end
return _M
+82
View File
@@ -0,0 +1,82 @@
# scripts/reload.ps1
#
# PCSX-Redux Lua helper reload client.
#
# Modes:
# elf - Request a full ELF reload. Requires -ElfPath.
# patch - Request a single-word RAM patch. Requires -Address and -Word.
#
# -RequestOnly prints the URI and exits before any network I/O.
# -Quiet suppresses the compact-JSON printout on the real path.
[CmdletBinding()]
param(
[ValidateSet('elf', 'patch')][string]$Mode = 'elf',
[string]$Target = 'hello_camera',
[string]$ElfPath = '',
[string]$Address = '',
[string]$Word = '',
[int]$Port = 8080,
[switch]$RequestOnly,
[switch]$Quiet
)
# mode-specific argument guards
switch ($Mode) {
'patch' {
if ([string]::IsNullOrEmpty($Address) -or [string]::IsNullOrEmpty($Word)) {
Write-Error "patch mode requires both -Address and -Word"
exit 1
}
}
'elf' {
if ([string]::IsNullOrEmpty($ElfPath)) {
Write-Error "elf mode requires -ElfPath"
exit 1
}
}
}
# Build the URL-encoded query string.
$queryParts = New-Object System.Collections.Generic.List[string]
[void]$queryParts.Add("mode=$([uri]::EscapeDataString($Mode))")
[void]$queryParts.Add("target=$([uri]::EscapeDataString($Target))")
switch ($Mode) {
'elf' {
[void]$queryParts.Add("path=$([uri]::EscapeDataString($ElfPath))")
}
'patch' {
[void]$queryParts.Add("addr=$([uri]::EscapeDataString($Address))")
[void]$queryParts.Add("hex=$([uri]::EscapeDataString($Word))")
}
}
$uri = "http://localhost:$Port/api/v1/lua/reload?$($queryParts -join '&')"
# RequestOnly path: emit URI and return before any network I/O.
if ($RequestOnly) {
Write-Output $uri
return
}
# Real request path: POST, decode body if it is a byte array, parse JSON.
$response = Invoke-WebRequest -Method Post -Uri $uri
if ($response.Content -is [byte[]]) {
$text = [System.Text.Encoding]::UTF8.GetString([byte[]]$response.Content)
}
else {
$text = [string]$response.Content
}
$obj = $text | ConvertFrom-Json
if (-not $Quiet) {
$obj | ConvertTo-Json -Compress | Write-Output
}
if (-not $obj.ok) {
$errCode = if ($obj.error) { [string]$obj.error } else { 'unknown' }
throw "Reload failed: $errCode"
}
-587
View File
@@ -1,587 +0,0 @@
#!/usr/bin/env lua
-- tape_atom_offset_gen.lua
--
-- Finds every `MipsAtom_(name) { ... }` declaration in the given sources,
-- counts the words in each body using the WORD_COUNT manifest, computes
-- branch offsets for atom_label(name) / atom_offset(tag, name) markers,
-- and writes one header per source into <source_dir>/gen/<basename>.offsets.h
--
-- Generated header layout (per source):
-- #pragma region <basename>
-- #undef atom_offset
-- #define atom_offset(tag, name) atom_offset_##tag##_##name
-- // --- atom: <name> (<n> words) ---
-- #define atom_offset_<tag>_<target> (N) // preprocessor form
-- #undef atom_offset_<tag>_<target> // (so enum can reuse)
-- enum {
-- atom_offset_<tag>_<target> = N, // C enum form
-- };
-- #define atom_offset_<tag>_<target> (N) // re-define for preprocessor
-- #pragma endregion <basename>
--
-- Usage:
-- lua gen_atom_offsets.lua <metadata.h> <source1> [source2 ...]
-- ============================================================
-- Character classification
-- ============================================================
local function is_space(c) return c == " " or c == "\t" or c == "\n" or c == "\r" or c == "\v" or c == "\f" end
local function is_alpha(c)
if not c or #c == 0 then return false end
if c >= "a" and c <= "z" then return true end
if c >= "A" and c <= "Z" then return true end
return c == "_"
end
local function is_digit(c) return c and c >= "0" and c <= "9" end
local function is_alnum(c) return is_alpha(c) or is_digit(c) end
-- ============================================================
-- I/O
-- ============================================================
local function read_file(path)
local f = io.open(path, "r")
if not f then error("Cannot open " .. path) end
local content = f:read("*a")
f:close()
return content
end
local function write_file(path, content)
local f = io.open(path, "w")
if not f then error("Cannot write " .. path) end
f:write(content)
f:close()
end
-- PowerShell aliases `mkdir` to New-Item, which treats `-p` as a path, so guard the call.
local function ensure_dir(path)
local is_win = package.config:sub(1, 1) == "\\"
os.execute(is_win and ('if not exist "' .. path .. '" mkdir "' .. path .. '"') or ('mkdir -p "' .. path .. '" 2>/dev/null'))
end
-- ============================================================
-- String primitives
-- ============================================================
local function trim(s)
local a = 1; while a <= #s and is_space(s:sub(a, a)) do a = a + 1 end
local b = #s; while b >= a and is_space(s:sub(b, b)) do b = b - 1 end
return s:sub(a, b)
end
local function starts_with(s, prefix)
if #s < #prefix then return false end
for i = 1, #prefix do
if s:sub(i, i) ~= prefix:sub(i, i) then return false end
end
return true
end
local function ends_with(s, suffix)
if #s < #suffix then return false end
local off = #s - #suffix
for i = 1, #suffix do
if s:sub(off + i, off + i) ~= suffix:sub(i, i) then return false end
end
return true
end
local function find_byte(haystack, target, start)
for i = start or 1, #haystack do
if haystack:sub(i, i) == target then return i end
end
return nil
end
local function dirname(path)
local last_sep = 0
for i = 1, #path do
local c = path:sub(i, i)
if c == "/" or c == "\\" then last_sep = i end
end
if last_sep == 0 then return "." end
return path:sub(1, last_sep - 1)
end
local function basename_no_ext(path)
local last_sep = 0
for i = 1, #path do
local c = path:sub(i, i)
if c == "/" or c == "\\" then last_sep = i end
end
local a = last_sep + 1
local last_dot = #path + 1
for i = #path, a, -1 do
if path:sub(i, i) == "." then last_dot = i; break end
end
return path:sub(a, last_dot - 1)
end
local function to_upper(s) return s:upper() end
local function to_alnum_underscore(s)
local out = ""
for i = 1, #s do
local c = s:sub(i, i)
if is_alnum(c) then out = out .. c
else out = out .. "_" end
end
return out
end
local function pad_right(s, w) return s .. string.rep(" ", w - #s) end
-- ============================================================
-- Lexer helpers
-- ============================================================
-- If position i starts a C string literal ("..."), char literal ('.'),
-- // line comment, or /* block comment, advance past it and return the
-- position just after the construct (or #s+1 if unterminated).
-- Otherwise return i unchanged.
local function skip_str_or_cmt(s, i)
local c = s:sub(i, i)
if c == '"' or c == "'" then
i = i + 1
while i <= #s do
if s:sub(i, i) == "\\" then i = i + 2
elseif s:sub(i, i) == c then return i + 1
else i = i + 1 end
end
return #s + 1
elseif c == "/" then
local nx = s:sub(i+1, i+1)
if nx == "/" then
while i <= #s and s:sub(i, i) ~= "\n" do i = i + 1 end
return i
elseif nx == "*" then
i = i + 2
while i <= #s - 1 do
if s:sub(i, i) == "*" and s:sub(i+1, i+1) == "/" then
return i + 2
end
i = i + 1
end
return #s + 1
end
end
return i
end
local function skip_ws_and_cmt(s, i)
while i <= #s do
if is_space(s:sub(i, i)) then i = i + 1
else
local nx = skip_str_or_cmt(s, i)
if nx > i then i = nx else break end
end
end
return i
end
local function read_ident(source, i)
if not is_alpha(source:sub(i, i)) then return nil, i end
local a = i
i = i + 1
while i <= #source and is_alnum(source:sub(i, i)) do i = i + 1 end
return source:sub(a, i - 1), i
end
local function read_balanced(s, open_char, close_char, i)
if s:sub(i, i) ~= open_char then return nil, i end
i = i + 1
local len = #s
local depth = 1
local a = i
while i <= len and depth > 0 do
local c = s:sub(i, i)
if c == open_char then
depth = depth + 1
i = i + 1
elseif c == close_char then
depth = depth - 1
if depth == 0 then break end
i = i + 1
else
local nx = skip_str_or_cmt(s, i)
if nx > i then i = nx else i = i + 1 end
end
end
return s:sub(a, i - 1), i + 1
end
local read_parens = function(s, i) return read_balanced(s, "(", ")", i) end
local read_braces = function(s, i) return read_balanced(s, "{", "}", i) end
local read_brackets = function(s, i) return read_balanced(s, "[", "]", i) end
local function scan_to_char(s, target, start)
local i = start
while i <= #s do
local c = s:sub(i, i)
if c == target then return i end
if c == "(" then local _, a = read_balanced(s, "(", ")", i); i = a
elseif c == "{" then local _, a = read_balanced(s, "{", "}", i); i = a
elseif c == "[" then local _, a = read_balanced(s, "[", "]", i); i = a
else
local nx = skip_str_or_cmt(s, i)
if nx > i then i = nx else i = i + 1 end
end
end
end
-- ============================================================
-- Extract comma-separated identifier args from a parenthesized group
-- after a function-like macro call.
-- ============================================================
local function extract_ident_args(token, after_ident)
local arg_start = skip_ws_and_cmt(token, after_ident)
if token:sub(arg_start, arg_start) ~= "(" then return {}, nil end
local inner, after_paren = read_parens(token, arg_start)
local args = {}
local n = 1
local len = #inner
while n <= len do
n = skip_ws_and_cmt(inner, n)
if n > len then break end
local ident, after = read_ident(inner, n)
if ident and ident ~= "" then
table.insert(args, ident)
n = after
else
n = n + 1
end
n = skip_ws_and_cmt(inner, n)
if n <= len and inner:sub(n, n) == "," then n = n + 1 end
end
return args, after_paren
end
-- ============================================================
-- Load WORD_COUNT manifest
-- ============================================================
local function load_word_counts(metadata_path)
local counts = {}
local content = read_file(metadata_path)
local len = #content
local i = 1
local prefix = "WORD_COUNT("
while i <= len do
local nl = find_byte(content, "\n", i)
local line_end = nl or (len + 1)
local line = content:sub(i, line_end - 1)
local trimmed = trim(line)
if starts_with(trimmed, prefix) and ends_with(trimmed, ")") then
local inner = trimmed:sub(#prefix + 1, #trimmed - 1)
local comma = find_byte(inner, ",", 1)
if comma then
counts[trim(inner:sub(1, comma - 1))] = tonumber(trim(inner:sub(comma + 1)))
end
end
i = line_end + 1
end
return counts
end
-- ============================================================
-- Count words for a single comma-separated token
-- ============================================================
local function word_count_of_token(token, wc)
local s = trim(token)
if s == "" then return 0 end
local name, after = read_ident(s, 1)
if not name then return 1 end
if wc[name] then return wc[name] end
local j = skip_ws_and_cmt(s, after)
if s:sub(j, j) == "(" then
io.stderr:write(" warning: unknown macro '" .. name .. "', assuming 1 word\n")
end
return 1
end
-- ============================================================
-- Split brace-body into top-level comma-separated tokens
-- ============================================================
local function split_top_level_commas(body)
local tokens = {}
local i = 1
local token_start = 1
while i <= #body do
local c = body:sub(i, i)
if c == "(" then local _, a = read_parens(body, i); i = a
elseif c == "{" then local _, a = read_braces(body, i); i = a
elseif c == "[" then local _, a = read_brackets(body, i); i = a
elseif c == "," then
table.insert(tokens, body:sub(token_start, i - 1))
i = i + 1
token_start = i
else
local nx = skip_str_or_cmt(body, i)
if nx > i then i = nx else i = i + 1 end
end
end
local last = body:sub(token_start)
if trim(last) ~= "" then table.insert(tokens, last) end
return tokens
end
-- ============================================================
-- Scan token for atom_label/atom_offset markers, walking through
-- balanced groups transparently (so nested calls are found)
-- ============================================================
local function scan_for_atom_markers(token, at_pos, labels, branches)
local i = 1
local len = #token
while i <= len do
i = skip_ws_and_cmt(token, i)
if i > len then break end
local c = token:sub(i, i)
if is_alpha(c) then
local ident, after = read_ident(token, i)
if ident == "atom_label" then
local args, after_paren = extract_ident_args(token, after)
if #args >= 1 then labels[args[1]] = at_pos end
if after_paren then i = after_paren else i = after end
elseif ident == "atom_offset" then
local args, after_paren = extract_ident_args(token, after)
if #args >= 2 then table.insert(branches, {pos = at_pos, target = args[2], tag = args[1]}) end
if after_paren then i = after_paren else i = after end
else
i = after
end
else
local nx = skip_str_or_cmt(token, i)
if nx > i then i = nx else i = i + 1 end
end
end
end
-- ============================================================
-- Scan atom body, count words, find markers
-- ============================================================
local function scan_atom_body(body, word_counts)
local pos = 0
local labels = {}
local branches = {}
for _, tok in ipairs(split_top_level_commas(body)) do
local k = 1
local tlen = #tok
while k <= tlen and is_space(tok:sub(k, k)) do k = k + 1 end
local leading_ident = read_ident(tok, k)
if leading_ident == "atom_label" or leading_ident == "atom_offset" then
scan_for_atom_markers(tok, pos, labels, branches)
else
local words = word_count_of_token(tok, word_counts)
scan_for_atom_markers(tok, pos, labels, branches)
pos = pos + words
end
end
return labels, branches, pos
end
-- ============================================================
-- Find every MipsAtom_(name) { ... } in a source
-- ============================================================
local function skip_qualifiers(source, i)
local keywords = {
["static"] = true, ["const"] = true, ["volatile"] = true,
["extern"] = true, ["register"] = true, ["auto"] = true,
["inline"] = true, ["typedef"] = true,
["internal"]= true, ["LP_"] = true, ["global"] = true, ["gkknown"] = true
}
while true do
i = skip_ws_and_cmt(source, i)
local ident, after = read_ident(source, i)
if not ident then return i end
if keywords[ident] then i = after else return i end
end
end
local function find_atoms(source_text)
local atoms = {}
local len = #source_text
local i = 1
local function try_wrapped(after_pos)
local paren_pos = skip_ws_and_cmt(source_text, after_pos)
if source_text:sub(paren_pos, paren_pos) ~= "(" then return nil end
local inner, after_paren = read_parens(source_text, paren_pos)
local n = 1
while n <= #inner and is_space(inner:sub(n, n)) do n = n + 1 end
local ns = n
while n <= #inner and is_alnum(inner:sub(n, n)) do n = n + 1 end
local name = inner:sub(ns, n - 1)
if name == "" then return nil end
local brace_pos = scan_to_char(source_text, "{", after_paren)
if not brace_pos then return nil end
local body, after_brace = read_braces(source_text, brace_pos)
return {name = name, body = body, after_brace = after_brace}
end
local function try_raw(after_pos)
local next_pos = skip_ws_and_cmt(source_text, after_pos)
local next_ident, next_after = read_ident(source_text, next_pos)
if not next_ident then return nil end
if not starts_with(next_ident, "code_") then return nil end
if #next_ident <= 5 then return nil end
local atom_name = next_ident:sub(6)
local brace_pos = scan_to_char(source_text, "{", next_after)
if not brace_pos then return nil end
local body, after_brace = read_braces(source_text, brace_pos)
return {name = atom_name, body = body, after_brace = after_brace}
end
while i <= len do
i = skip_ws_and_cmt(source_text, i); if i > len then break end
i = skip_qualifiers(source_text, i); if i > len then break end
local ident, after = read_ident(source_text, i)
if not ident then
i = i + 1
elseif ident == "MipsAtom_" then
local atom = try_wrapped(after)
if atom then
table.insert(atoms, {name = atom.name, body = atom.body})
i = atom.after_brace
else
i = i + 1
end
elseif ident == "MipsCode" then
local atom = try_raw(after)
if atom then
table.insert(atoms, {name = atom.name, body = atom.body})
i = atom.after_brace
else
i = after
end
else
i = after
end
end
return atoms
end
-- ============================================================
-- Compute branch offsets (target - branch - 1)
-- ============================================================
local function compute_offsets(labels, branches)
local results = {}
for _, br in ipairs(branches) do
local target = labels[br.target]
if not target then
error("Branch target '" .. br.target .. "' has no atom_label (at word " .. br.pos .. ")")
end
table.insert(results, {target = br.target, tag = br.tag, offset = target - br.pos - 1 })
end
return results
end
-- ============================================================
-- Generate header for one source
-- ============================================================
local function generate_header(source_path, atoms_data)
local basename = basename_no_ext(source_path)
local guard = to_alnum_underscore(to_upper(basename)) .. "_OFFSETS_H"
local lines = {}
local function add(s) table.insert(lines, s) end
add("// Auto-generated by tape_atom_offset_gen.meta.lua — DO NOT EDIT")
add("// Source: " .. source_path)
add("#pragma once")
add("")
add("#pragma region " .. basename)
add("")
-- add("// Dispatch macro: token-pastes <tag>_<target> to the enum name")
-- add("#undef atom_offset")
-- add("#define atom_offset(tag, name) atom_offset_##tag##_##name")
add("")
for _, atom in ipairs(atoms_data) do
if #atom.offsets > 0 then
add("// --- atom: " .. atom.name .. " (" .. atom.total_words .. " words) ---")
add("")
local consts = {}
for _, r in ipairs(atom.offsets) do
table.insert(consts, {
macro_name = "_atom_offset_" .. r.tag .. "_" .. r.target,
enum_name = "atom_offset_" .. r.tag .. "_" .. r.target,
value = r.offset
})
end
for _, c in ipairs(consts) do add("#define " .. pad_right(c.macro_name, 44) .. " " .. c.value .. "") end
add("")
add("enum {")
for _, c in ipairs(consts) do add(" " .. c.enum_name .. " = " .. c.macro_name .. ",") end
add("};")
add("")
end
end
add("#pragma endregion " .. basename)
add("")
return table.concat(lines, "\n") .. "\n"
end
-- ============================================================
-- Process one source
-- ============================================================
local function process_source(source_path, word_counts)
local source = read_file(source_path)
local atoms_raw = find_atoms(source)
if #atoms_raw == 0 then
-- io.stderr:write(" note: no MipsAtom_ declarations in " .. source_path .. "\n")
return
end
local atoms_data = {}
for _, atom in ipairs(atoms_raw) do
local labels, branches, total = scan_atom_body(atom.body, word_counts)
local offsets = compute_offsets(labels, branches)
table.insert(atoms_data, {
name = atom.name,
total_words = total,
offsets = offsets
})
end
local basename = basename_no_ext(source_path)
local out_dir = dirname(source_path) .. "/gen"
ensure_dir(out_dir)
local out_path = out_dir .. "/" .. basename .. ".offsets.h"
write_file(out_path, generate_header(source_path, atoms_data))
local total_branches = 0
for _, a in ipairs(atoms_data) do total_branches = total_branches + #a.offsets end
print(" " .. basename .. ": " .. #atoms_data .. " atom(s), " .. total_branches .. " branch(es)")
for _, a in ipairs(atoms_data) do
for _, r in ipairs(a.offsets) do
print(" " .. a.name .. " -> " .. r.tag .. ":" .. r.target .. " : " .. r.offset)
end
end
end
-- ============================================================
-- Main
-- ============================================================
local function main(args)
if #args < 2 then
print("Usage: gen_atom_offsets.lua <metadata.h> <source1> [source2 ...]")
os.exit(1)
end
local word_counts = load_word_counts(args[1])
for i = 2, #args do process_source(args[i], word_counts) end
end
main({...})
File diff suppressed because it is too large Load Diff
+86 -7
View File
@@ -4,25 +4,24 @@ $path_code = join-path $path_root 'code'
$path_scripts = join-path $path_root 'scripts'
$path_toolchain = join-path $path_root 'toolchain'
# Halt on any error (instead of PowerShell's default `Continue`).
$ErrorActionPreference = 'Stop'
$misc = join-path $PSScriptRoot 'helpers/misc.ps1'
. $misc
# TODO(Ed): Review usage of these deps
# I orgiinally cloned them when starting to get to the C runtime usage of the course
# However, based on the heavy reliance of the PSX.Dev extension I might fallback; also
# The gdb server doesn't need the full repo and were only using the src/mips
# which has a standalone repo (nuggets)
# armips may not be used at all but I'm not sure...
$url_armips = 'https://github.com/Kingcom/armips.git'
$url_pcsx_redux = 'https://github.com/grumpycoders/pcsx-redux.git'
$url_psyq_iwyu = 'https://github.com/johnbaumann/psyq_include_what_you_use.git'
$url_lpeg = 'https://github.com/roberto-ieru/LPeg.git'
$path_armips = join-path $path_toolchain 'armips'
$path_pcsx_redux = join-path $path_toolchain 'pcsx-redux'
$path_psyq_iwyu = join-path $path_toolchain 'psyq_iwyu'
$path_lpeg = join-path $path_toolchain 'lpeg'
clone-gitrepo $path_armips $url_armips
clone-gitrepo $path_lpeg $url_lpeg
clone-gitrepo $path_pcsx_redux $url_pcsx_redux
clone-gitrepo $path_psyq_iwyu $url_psyq_iwyu
@@ -37,3 +36,83 @@ pop-location
# $path_pcsx_redux_binaries = join-path $path_pcsx_redux_vsprojects 'x64/Release'
# $psyq_obj_parser = join-path $path_pcsx_redux_binaries 'psyq-obj-parser.exe'
# ════════════════════════════════════════════════════════════════════════════
# PCSX-Redux — built via MSBuild (VS2022)
# Requires: Visual Studio 2022 with the C++ desktop workload.
# Output: toolchain\pcsx-redux\vsprojects\x64\Debug\pcsx-redux.exe
# ════════════════════════════════════════════════════════════════════════════
# Locate MSBuild from the VS2022 install (no hardcoded path — uses vswhere).
$vswhere = "${env:ProgramFiles(x86)}\Microsoft Visual Studio\Installer\vswhere.exe"
if (-not (Test-Path $vswhere)) {
write-error "vswhere not found at '$vswhere'. Install Visual Studio 2022 with the C++ desktop workload."
exit 1
}
$msbuild_exe = & $vswhere -latest -products * -requires Microsoft.Component.MSBuild -find "MSBuild\**\Bin\MSBuild.exe" 2>$null | Select-Object -First 1
if (-not $msbuild_exe) {
write-error "MSBuild not found via vswhere. Install Visual Studio 2022 with the C++ desktop workload."
exit 1
}
$path_pcsx_sln = join-path $path_pcsx_redux 'vsprojects\pcsx-redux.sln'
& $msbuild_exe $path_pcsx_sln /p:Configuration=Release /p:Platform=x64 /p:PlatformToolset=v143 /m /v:minimal
# Locate luajit via scoop. `luajit.exe` is on PATH via scoop's shim;
# we use `scoop prefix` to find the install root for the include dir (needed to compile lpeg against luajit's headers).
# If scoop or luajit is missing, fail fast with an actionable message.
$luajit_prefix = & scoop prefix luajit 2>$null
if (-not $luajit_prefix -or -not (Test-Path (Join-Path $luajit_prefix 'bin/luajit.exe'))) {
write-error "luajit not found via 'scoop prefix luajit'. Install via: scoop install luajit"
exit 1
}
# Discover the luajit include dir by globbing `include/luajit-*`.
# This avoids hardcoding a specific version (e.g. `luajit-2.1`).
$luajit_include_root = Join-Path $luajit_prefix 'include'
$lua_inc_dir = Get-ChildItem -Path $luajit_include_root -Directory -Filter 'luajit-*' -ErrorAction SilentlyContinue |
Select-Object -First 1 -ExpandProperty FullName
if (-not $lua_inc_dir) {
write-error "No 'luajit-*' include dir found under '$luajit_include_root'. The scoop luajit install may be broken."
exit 1
}
# Generate lpeg.dll by compiling the 6 source files directly.
# `gcc` is on PATH (scoop's shim puts it there).
# The source files: lpcap.c lpcode.c lpcset.c lpprint.c lptree.c lpvm.c
# Link against luajit's import library (`libluajit-5.1.a`) for the Lua C API symbols (lua_*, luaL_*).
$luajit_lib_dir = Join-Path $luajit_prefix 'lib'
$lpeg_sources = @('lpcap.c', 'lpcode.c', 'lpcset.c', 'lpprint.c', 'lptree.c', 'lpvm.c')
$lpeg_compile_args = @(
'-O2', '-shared',
"-I$lua_inc_dir",
"-L$luajit_lib_dir",
'-o', 'lpeg.dll'
) + $lpeg_sources + @('-lluajit-5.1')
push-location $path_lpeg
& gcc @lpeg_compile_args
pop-location
# ════════════════════════════════════════════════════════════════════════════
# lfs (LuaFileSystem) — compiled from pcsx-redux's vendored luafilesystem source.
# Source: toolchain/pcsx-redux/third_party/luafilesystem/src/lfs.c
# Output: toolchain/lfs/lfs.dll
# ════════════════════════════════════════════════════════════════════════════
$path_lfs = join-path $path_toolchain 'lfs'
verify-path $path_lfs
$lfs_src = join-path $path_pcsx_redux 'third_party\luafilesystem\src\lfs.c'
$lfs_dll = join-path $path_lfs 'lfs.dll'
$lfs_dll_import = join-path $luajit_lib_dir 'libluajit-5.1.dll.a'
& gcc -O2 -shared "-I$lua_inc_dir" -o $lfs_dll $lfs_src $lfs_dll_import
# ════════════════════════════════════════════════════════════════════════════
# OpenBIOS — built from the PCSX-Redux source tree via make + mipsel-none-elf
# Output: toolchain\pcsx-redux\src\mips\openbios\openbios.bin
# ════════════════════════════════════════════════════════════════════════════
$path_openbios = join-path $path_pcsx_redux 'src\mips\openbios'
push-location $path_openbios
& make clean
& make
pop-location