mirror of
https://github.com/Ed94/pikuma_ps1.git
synced 2026-08-04 14:48:48 +00:00
WIP: Better step debug on atom components, better db_skip annotation, lots of curation passes on lua.
Still don't have this thing in its final state for the curse but its close.
This commit is contained in:
+249
-200
@@ -1,15 +1,16 @@
|
||||
--- duffle.lua — Shared primitives + domain tables for the tape-atom
|
||||
--- metaprograms.
|
||||
--- duffle.lua — shared primitives + domain tables for the tape-atom metaprograms.
|
||||
---
|
||||
--- This module is the source for:
|
||||
--- - **Character classification** (`is_space`, `is_alpha`, `is_alnum`, `is_digit`, plus the byte-fast `_byte` variants).
|
||||
--- - **String/path primitives** (`trim`, `dirname`, `basename_no_ext`, `normalize_path`, `canonical_path_key`, `find_byte`).
|
||||
--- - **I/O primitives** (`read_file`, `write_file`, `ensure_dir`).
|
||||
--- - **Canonical corpus resolution** (`parse_direct_quoted_includes`, `resolve_source_corpus`).
|
||||
--- - **C-language scanner** (`skip_ws_and_cmt`, `skip_str_or_cmt`, `read_ident`, `read_parens`, `read_braces`, `read_brackets`, `read_balanced`, `scan_to_char`, `split_top_level_commas`).
|
||||
--- - **Word-count loader** (`load_word_counts` for `WORD_COUNT(...)` metadata files).
|
||||
--- - **Line lookup** (`LineIndex` returns an O(log N) `line_of(pos)` closure for source-mapping).
|
||||
--- - **Domain tables** (`TAPE_ATOM_MACROS`, `GTE_PIPELINE_LATENCY`, `GP0_CMD_SIZE`, `GP0_CMD_BY_SHAPE`, `GP0_MACRO_CONTRIB`, `INSTRUCTION_LATENCY`).
|
||||
--- One ownership statement, then the rest is signal:
|
||||
--- * **Character classification** (`is_space`, `is_alpha`, `is_alnum`, `is_digit`, plus the byte-fast `_byte` variants).
|
||||
--- * **String / path primitives** (`trim`, `dirname`, `basename_no_ext`, `normalize_path`, `canonical_path_key`, `find_byte`).
|
||||
--- * **I/O primitives** (`read_file`, `write_file`, `ensure_dir`).
|
||||
--- * **Corpus resolution** (`parse_direct_quoted_includes`, `resolve_source_corpus`).
|
||||
--- * **C-language scanner** (`skip_ws_and_cmt`, `skip_str_or_cmt`, `read_ident`, `read_parens`, `read_braces`, `read_brackets`,
|
||||
--- `read_balanced`, `scan_to_char`, `split_top_level_commas`).
|
||||
--- * **Word-count loader** (`load_word_counts` for `WORD_COUNT(...)` metadata files).
|
||||
--- * **Line lookup** (`LineIndex` returns an O(log N) `line_of(pos)` closure for source-mapping).
|
||||
--- * **Domain tables** (`TAPE_ATOM_MACROS`, `GTE_PIPELINE_LATENCY`, `GP0_CMD_SIZE`, `GP0_CMD_BY_SHAPE`,
|
||||
--- `GP0_MACRO_CONTRIB`, `INSTRUCTION_LATENCY`).
|
||||
---
|
||||
--- **Conventions**: tabs (1/level), EmmyLua annotations, no regex.
|
||||
|
||||
@@ -73,27 +74,22 @@ local BYTE_DIGIT_9 = 0x39 -- '9'
|
||||
-- Section -1: Bootstrap (path-setup at module load)
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
--
|
||||
-- Path setup is done by `scripts/duffle_paths.lua`, which derives the repo root from `debug.getinfo(1, "S").source` (NO subprocess, ~0ms) and then calls `require("duffle")`.
|
||||
-- Repository paths come from `scripts/duffle_paths.lua` because:
|
||||
-- 1. Entry and pass scripts load `duffle_paths.lua`.
|
||||
-- The `find_repo_root` / `setup_package_path` defined here was dead code in practice.
|
||||
-- 2. `git rev-parse` costs ~100-180ms per subprocess spawn on Windows.
|
||||
-- `debug.getinfo` is <1ms. There's no reason to keep the slow path even as a "fallback".
|
||||
-- Path setup runs through `scripts/duffle_paths.lua`, which derives the repo root from `debug.getinfo(1, "S").source`
|
||||
-- (no subprocess, ~0ms) and then calls `require("duffle")`.
|
||||
-- Entry and pass scripts load `duffle_paths.lua` first; a `find_repo_root` / `setup_package_path` defined here was dead code in practice.
|
||||
-- `git rev-parse` costs ~100-180ms per subprocess spawn on Windows; `debug.getinfo` is <1ms, so we keep only the fast path.
|
||||
--
|
||||
-- If a future use case ever needs to load `duffle.lua` WITHOUT going through `duffle_paths.lua`, set `package.path` manually before `require`.
|
||||
-- To load `duffle.lua` outside `duffle_paths.lua`, set `package.path` manually before `require`.
|
||||
-- See `docs/guide_metaprogram_ssdl.md` §"I/O primitives" for the pattern.
|
||||
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
-- Section 0: LPeg patterns (compiled once at module load)
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
--
|
||||
-- LPeg is a required dependency (PEG library, no regex).
|
||||
-- It's loaded via `package.cpath` (configured by `duffle_paths.lua` to find `toolchain/lpeg/lpeg.dll`).
|
||||
-- LPeg handles the high-level scanner, while Section 1 handles byte classification.
|
||||
-- only relevant at the high-level scanner stage; the byte-by-byte helpers in Section 1 are sufficient for the classification primitives.
|
||||
-- LPeg is a required dependency (PEG library, no regex). It's loaded via `package.cpath` — `duffle_paths.lua` wires the path to `toolchain/lpeg/lpeg.dll`.
|
||||
-- LPeg handles the high-level scanner; the byte-by-byte helpers in Section 1 handle classification primitives that LPeg's CPython-level cost would dominate.
|
||||
--
|
||||
-- If the require fails, fail loud with an actionable message. The build script (`update_deps.ps1`) builds lpeg.dll into `toolchain/lpeg/`;
|
||||
-- if it's missing, run `update_deps.ps1`.
|
||||
-- If the require fails, fail loud with an actionable message. The build script (`update_deps.ps1`) builds lpeg.dll into `toolchain/lpeg/`; run it when the dll is missing.
|
||||
local lpeg_ok, lpeg = pcall(require, "lpeg")
|
||||
if not lpeg_ok then
|
||||
io.stderr:write("[duffle] require('lpeg') failed: ", lpeg, "\n")
|
||||
@@ -395,7 +391,7 @@ end
|
||||
|
||||
--- Group a list of `SourceFile`-shaped records by their `dir` field.
|
||||
--- Used by the annotation / static-analysis / report passes to partition sources into per-DIRECTORY (per-module) buckets before emitting per-module reports.
|
||||
--- Insertion order preserved within each bucket (matches source order in `ctx.sources`).
|
||||
--- Insertion order is preserved within each bucket (matches source order in `corpus.source_order`).
|
||||
--- @param sources table[] -- list of source records (each having a `dir` string field)
|
||||
--- @return table<string, table[]> -- map of `dir` -> sources in that dir
|
||||
function M.group_sources_by_dir(sources)
|
||||
@@ -562,7 +558,7 @@ local function splice_c_lines(source)
|
||||
end
|
||||
|
||||
--- Parse direct quoted preprocessor includes from one source buffer.
|
||||
--- Translation Line splicing occurs ahead of comment, string, and directive processing.
|
||||
--- Line splicing occurs ahead of comment, string, and directive processing.
|
||||
--- Interpreted records retain original physical include text and line numbers.
|
||||
--- Angle includes and include-like text inside comments/strings are ignored.
|
||||
--- @param source_text string
|
||||
@@ -850,13 +846,10 @@ function M.split_top_level_commas(body)
|
||||
if has_real_content(chunk) then
|
||||
tokens[#tokens + 1] = chunk
|
||||
elseif #tokens > 0 then
|
||||
-- Pure comment/string chunk at top level (no preceding instruction content within this chunk).
|
||||
-- APPEND it to the LAST token so emit-context callers (components.lua build_component_lines)
|
||||
-- can convert `// trailing comment` to `/* */` and emit it with the macro body.
|
||||
-- For word counting, count_token_words only inspects the leading ident, so a trailing comment doesn't affect the count.
|
||||
--
|
||||
-- This is the second-half fix to commit 98e27c2: the first fix correctly broke top-level comments off from the NEXT statement (fixing macro-call word counts);
|
||||
-- This fix preserves them on the PREVIOUS statement (restoring the comments in the emitted .macs.h output).
|
||||
-- Pure comment/string chunk at top level.
|
||||
-- Append it to the LAST token so emit-context callers (components.lua build_component_lines) can convert
|
||||
-- `// trailing comment` to `/* */` and emit it with the macro body.
|
||||
-- count_token_words only inspects the leading ident, so a trailing comment does not affect the count.
|
||||
tokens[#tokens] = tokens[#tokens] .. chunk
|
||||
end
|
||||
end
|
||||
@@ -1048,18 +1041,16 @@ M.TAPE_ATOM_MACROS = {
|
||||
|
||||
-- GTE command-alias resolution table.
|
||||
--
|
||||
-- Maps each GTE command macro that may appear in source to its CANONICAL short form.
|
||||
-- Both forms resolve to the same PSX-SPX-documented pipeline semantics;
|
||||
-- The canonical name is the only one that appears in `GTE_COMMAND_INPUTS` and the per-check producer / consumer reports.
|
||||
-- Aliases resolve exactly once; unknown idents (e.g. an MVMVA with a custom `(sf, mx, v, cv, lm)` payload that is not on this list)
|
||||
-- are reported as "command unknown" by the check, not silently treated as 0-cycle.
|
||||
-- Maps every source-side GTE command macro to its canonical short ident.
|
||||
-- Both forms run the same PSX-SPX-documented pipeline semantics.
|
||||
-- Aliases resolve exactly once; an unknown ident (an MVMVA with a custom `(sf, mx, v, cv, lm)` payload that is not on this list) lands as
|
||||
-- "command unknown" from the check rather than being silently treated as 0-cycle.
|
||||
--
|
||||
-- Source conventions (per `code/duffle/gte.h`): The C source ships both short canonical macros
|
||||
-- (`gte_cmdw_rtps`, `gte_cmdw_rtpt`, `gte_cmdw_nclip`, `gte_cmdw_avsz3`, `gte_cmdw_avsz4`, `gte_cmdw_mvmva`, `gte_cmdw_op`)
|
||||
-- and human-readable aliases (`gte_cmdw_rotate_translate_perspective_*`, `gte_cmdw_avg_sort_z3`, etc.).
|
||||
-- Every alias row maps source ident -> canonical short ident.
|
||||
-- Source conventions (per `code/duffle/gte.h`): the C source ships short idents (`gte_cmdw_rtps`, `gte_cmdw_rtpt`, `gte_cmdw_nclip`,
|
||||
-- `gte_cmdw_avsz3`, `gte_cmdw_avsz4`, `gte_cmdw_mvmva`, `gte_cmdw_op`) and human-readable aliases
|
||||
-- (`gte_cmdw_rotate_translate_perspective_*`, `gte_cmdw_avg_sort_z3`, etc.). Each alias row maps the source ident to its short form.
|
||||
M.GTE_COMMAND_ALIASES = {
|
||||
-- Canonical -> canonical (identity).
|
||||
-- Identity rows: short form resolves to itself.
|
||||
["gte_cmdw_rtps"] = "gte_cmdw_rtps",
|
||||
["gte_cmdw_rtpt"] = "gte_cmdw_rtpt",
|
||||
["gte_cmdw_nclip"] = "gte_cmdw_nclip",
|
||||
@@ -1067,7 +1058,7 @@ M.GTE_COMMAND_ALIASES = {
|
||||
["gte_cmdw_op"] = "gte_cmdw_op",
|
||||
["gte_cmdw_avsz3"] = "gte_cmdw_avsz3",
|
||||
["gte_cmdw_avsz4"] = "gte_cmdw_avsz4",
|
||||
-- Aliases -> canonical.
|
||||
-- Long-form aliases resolve to the short form.
|
||||
["gte_cmdw_rotate_translate_perspective_single"] = "gte_cmdw_rtps",
|
||||
["gte_cmdw_rotate_translate_perspective_triple"] = "gte_cmdw_rtpt",
|
||||
["gte_cmdw_avg_sort_z3"] = "gte_cmdw_avsz3",
|
||||
@@ -1082,28 +1073,27 @@ M.GTE_COMMAND_ALIASES = {
|
||||
|
||||
-- GTE command input-set table.
|
||||
--
|
||||
-- For each canonical command, the set of C2 registers whose recent CPU-to-COP2 write
|
||||
-- must retire before the command can issue. Per PSX-SPX `docs/psx-spx/docs/cpuspecifications.md:407-419`:
|
||||
-- For each command, the set of C2 registers whose recent CPU-to-COP2 write must retire before the command can issue.
|
||||
-- Per PSX-SPX `docs/psx-spx/docs/cpuspecifications.md:407-419`:
|
||||
-- * A store to COP2 registers (mtc2/ctc2) has a delay of 2..3 clock cycles.
|
||||
-- * In most cases the delay is 2 cycles; special cases like writes to IRGB
|
||||
-- (which additionally affect IR1/IR2/IR3) take 3 cycles.
|
||||
-- * In most cases the delay is 2 cycles; special cases like writes to IRGB (which additionally affect IR1/IR2/IR3) take 3 cycles.
|
||||
-- * "Store delays are counted in numbers of clock cycles (not in numbers of opcodes).
|
||||
-- For 3 cycle delay, one must usually insert 3 cached opcodes (or one uncached opcode)."
|
||||
--
|
||||
-- Per PSX-SPX `docs/psx-spx/docs/gtepipelinetimings.md`
|
||||
-- (the per-instruction input-latch measurement, which is the SAME phenomenon modeled from the command side), the values are:
|
||||
-- Per PSX-SPX `docs/psx-spx/docs/gtepipelinetimings.md` (the per-instruction input-latch measurement, which is the same
|
||||
-- phenomenon modeled from the command side), the values are:
|
||||
-- rtps: every data register, every control register (RT/TR/OFX/OFY/H/DQA/DQB)
|
||||
-- rtpt: same superset (rtpt reads V0..V2, the RT matrix, the TR vector, OFX/OFY, H, DQA, DQB)
|
||||
-- nclip: SXY0, SXY1, SXY2 (no RT/TR/OFX inputs)
|
||||
-- mvmva: variable (depends on the chosen mx / v / cv selector); treated conservatively as the union of all RT + TR + BK + IR columns (the data inputs the command can read).
|
||||
-- mvmva: variable (depends on the chosen mx / v / cv selector); treated conservatively as the union of all RT + TR + BK + IR columns
|
||||
-- (the data inputs the command can read).
|
||||
-- op: IR1, IR2, IR3 (cross-product output, atomic; consumers treat as fan-out only)
|
||||
-- avsz3/avsz4: SZ0..SZ3 + ZSF3/ZSF4
|
||||
--
|
||||
-- We model the data-register + control-register superset.
|
||||
-- Per PSX-SPX `gtepipelinetimings.md`, every relevant input is in this set;
|
||||
-- the per-input latching values listed there are the SAME number's command-side view
|
||||
-- We model the data-register + control-register superset. Every relevant input is in this set per PSX-SPX `gtepipelinetimings.md`;
|
||||
-- the per-input latching values there describe the same number's command-side view
|
||||
-- (a recent mtc2/ctc2 to that register must retire the same number of cycles before the command issues).
|
||||
-- Anything not in the set is safe to clobber immediately after a prior command.
|
||||
-- Anything outside this set is safe to clobber immediately after a prior command.
|
||||
M.GTE_COMMAND_INPUTS = {
|
||||
-- RTPS / RTPT: every data + every rotation/translation control + screen offset + projection.
|
||||
["gte_cmdw_rtps"] = {
|
||||
@@ -1166,15 +1156,15 @@ M.GTE_COMMAND_INPUTS = {
|
||||
|
||||
-- GTE command output-set + semantic role table.
|
||||
--
|
||||
-- For each canonical command, the SET of C2 data registers the command writes as outputs, paired with the SEMANTIC ROLE of each output.
|
||||
-- The semantic role is the basis for the `_post_<cmd>` contract validation:
|
||||
-- The contract says "after <cmd>, the latest screen-XY is C2_SXY2" (NOT C2_SXY0. The FIFO side effects do NOT make SXY0 the newest result).
|
||||
-- For each command, the set of C2 data registers the command writes as outputs, paired with the SEMANTIC ROLE of each output.
|
||||
-- The semantic role is the basis for the `_post_<cmd>` contract validation.
|
||||
-- The contract says "after <cmd>, the latest screen-XY is C2_SXY2" (C2_SXY0 is wrong; the FIFO side effects leave SXY0 as an older FIFO entry, never the newest).
|
||||
--
|
||||
-- Per PSX-SPX `docs/psx-spx/docs/geometrytransformationenginegte.md`:
|
||||
-- * RTPS: writes VXY/VZ -> MAC results; the SINGLE projected screen coordinate is written to C2_SXY2 (the IRGB -> SXY2 path via the perspective divide).
|
||||
-- C2_SXY0 and C2_SXY1 are NOT written.
|
||||
-- * RTPS: writes VXY/VZ -> MAC results; the single projected screen coordinate is written to C2_SXY2 (the IRGB -> SXY2 path via the perspective divide).
|
||||
-- C2_SXY0 and C2_SXY1 are untouched.
|
||||
-- * RTPT: writes three projected screen coordinates into SXY0, SXY1, SXY2 in pipeline order.
|
||||
-- The LAST projection is in C2_SXY2; a reader that wants "the last RTPT result" must read C2_SXY2.
|
||||
-- The last projection lives in C2_SXY2; a reader that wants "the last RTPT result" reads C2_SXY2.
|
||||
-- * NCLIP: writes a single MAC result into C2_SZ3 (the inner-product sum); no screen XY output.
|
||||
-- * AVSZ3 / AVSZ4: write average Z into C2_OTZ (single output).
|
||||
-- * OP: writes C2_IR1, C2_IR2, C2_IR3 (cross-product result; no projection).
|
||||
@@ -1190,22 +1180,21 @@ M.GTE_COMMAND_INPUTS = {
|
||||
-- * "mac_result" : generic MAC output (nclip, op, mvmva)
|
||||
--
|
||||
-- Consumers:
|
||||
-- * passes/static_analysis.lua::analyze_hardware_relations (the walker consults this table after a GTE command to update `forward_state.post_command_roles` for `gte_result_position`).
|
||||
-- * passes/static_analysis.lua::analyze_hardware_relations (the walker reads this after a GTE command to update
|
||||
-- `forward_state.post_command_roles` for `gte_result_position`).
|
||||
-- * passes/static_analysis.lua::check_gte_result_position (per-atom CHECK_RULES reader; renders role mismatches).
|
||||
-- This table is consumed by the hardware-relation analyzer and result-position check.
|
||||
M.GTE_COMMAND_OUTPUTS = {
|
||||
-- RTPS: writes ONE screen coordinate (the perspective-divide result)
|
||||
-- into C2_SXY2; the FIFO side effects do NOT make SXY0 / SXY1 newest.
|
||||
-- `latest_screen_xy` is C2_SXY2.
|
||||
-- RTPS: writes one screen coordinate (the perspective-divide result) into C2_SXY2.
|
||||
-- The FIFO side effects leave SXY0 / SXY1 untouched, so `latest_screen_xy` is C2_SXY2.
|
||||
["gte_cmdw_rtps"] = {
|
||||
{ register = "C2_SXY2", role = "latest_screen_xy" },
|
||||
{ register = "C2_SZ2", role = "latest_screen_z" },
|
||||
{ register = "C2_OTZ", role = "otz" },
|
||||
{ register = "C2_IR0", role = "latest_color" },
|
||||
},
|
||||
-- RTPT: writes THREE screen coordinates; the LAST projection lands in
|
||||
-- C2_SXY2. `latest_screen_xy` is C2_SXY2; C2_SXY0 / C2_SXY1 are the
|
||||
-- earlier projections of the batched triple.
|
||||
-- RTPT: writes three screen coordinates; the last projection lands in C2_SXY2 (`latest_screen_xy`).
|
||||
-- C2_SXY0 / C2_SXY1 carry the earlier projections of the batched triple.
|
||||
["gte_cmdw_rtpt"] = {
|
||||
{ register = "C2_SXY0", role = "screen_xy[0]" },
|
||||
{ register = "C2_SXY1", role = "screen_xy[1]" },
|
||||
@@ -1213,8 +1202,7 @@ M.GTE_COMMAND_OUTPUTS = {
|
||||
{ register = "C2_SZ3", role = "latest_screen_z" },
|
||||
{ register = "C2_OTZ", role = "otz" },
|
||||
},
|
||||
-- NCLIP: single MAC result; written to C2_SZ3 (the inner-product sum).
|
||||
-- No screen XY output.
|
||||
-- NCLIP: single MAC result; written to C2_SZ3 (the inner-product sum). No screen XY output.
|
||||
["gte_cmdw_nclip"] = {
|
||||
{ register = "C2_SZ3", role = "mac_result" },
|
||||
},
|
||||
@@ -1243,24 +1231,23 @@ M.GTE_COMMAND_OUTPUTS = {
|
||||
-- GTE command/post-command latch-window table.
|
||||
--
|
||||
-- Per PSX-SPX `docs/psx-spx/docs/gtepipelinetimings.md`, a GTE command emits outputs that latch into the pipeline for a measured number of emitted words.
|
||||
-- A subsequent MTC2/CTC2 OVERWRITE of one of those outputs BEFORE the latch window expires is a hazard
|
||||
-- (the latched value in the pipeline is overwritten by the CPU before the pipeline consumes it).
|
||||
-- A subsequent MTC2/CTC2 overwrite of one of those outputs before the latch window expires is a hazard:
|
||||
-- the latched value in the pipeline gets overwritten by the CPU before the pipeline consumes it.
|
||||
--
|
||||
-- This relation is the COMMAND -> REGISTER direction (the command is the producer, the MTC2/CTC2 is the consumer).
|
||||
-- It is NOT the same relation as the preceding MTC2 -> command input propagation
|
||||
-- (which is the REGISTER -> COMMAND direction and is staged by the producer step of `analyze_hardware_relations`).
|
||||
-- This relation is the command -> register direction (the command is the producer; MTC2/CTC2 is the consumer).
|
||||
-- It is the inverse of the MTC2 -> command input propagation (register -> command direction), which is staged by the
|
||||
-- producer step of `analyze_hardware_relations`.
|
||||
--
|
||||
-- The schema mirrors the producer-side relations (`direction`, `evidence`, `violation_kind`);
|
||||
-- `required` is the number of emitted words strictly between the command's last output word and the overwrite.
|
||||
-- `N=0` permits the immediately following overwrite instruction; `N=4` permits an overwrite that occurs after 4 intervening words.
|
||||
-- The schema mirrors the producer-side relations (`direction`, `evidence`, `violation_kind`); `required` counts the
|
||||
-- emitted words strictly between the command's last output word and the overwrite.
|
||||
-- `required = 0` permits the immediately following overwrite; `required = 4` requires four intervening words.
|
||||
--
|
||||
-- Per PSX-SPX `gtepipelinetimings.md`
|
||||
-- The per-command input latching measurements are the SAME number, just inverted:
|
||||
-- They describe when a recent MTC2/CTC2 must retire before the command issues;
|
||||
-- here we describe when a recent command's outputs latch into the pipeline before a later MTC2/CTC2 may overwrite them.
|
||||
-- Per PSX-SPX `gtepipelinetimings.md` the per-command input latching measurements are the same numbers inverted.
|
||||
-- They describe when a recent MTC2/CTC2 must retire before the command issues; this table describes when a recent
|
||||
-- command's outputs latch into the pipeline before a later MTC2/CTC2 overwrites them.
|
||||
--
|
||||
-- Consumers:
|
||||
-- * passes/static_analysis.lua::analyze_hardware_relations (the walker consults this table after a GTE command to stage post-command latch relations in `pending`).
|
||||
-- * passes/static_analysis.lua::analyze_hardware_relations (stages post-command latch relations in `pending` after a GTE command).
|
||||
-- * passes/static_analysis.lua::check_gte_input_latch (per-atom CHECK_RULES reader; renders the over-the-boundary findings).
|
||||
-- This table is consumed by the hardware-relation analyzer and input-latch check.
|
||||
M.GTE_COMMAND_LATCH_WINDOWS = {
|
||||
@@ -1308,23 +1295,24 @@ M.GTE_COMMAND_LATCH_WINDOWS = {
|
||||
|
||||
-- GTE component result contracts (immutable; keyed by bare component name).
|
||||
--
|
||||
-- Register-role claims that cannot be inferred from the `_post_<cmd>` suffix alone live here.
|
||||
-- The bare name (without the `_post_<cmd>` suffix) is the key; the row carries the expected command, the expected role, and the expected C2 register.
|
||||
-- Register-role claims that the `_post_<cmd>` suffix alone cannot infer live here.
|
||||
-- The bare name (the component name stripped of the `_post_<cmd>` suffix) is the key; the row carries the expected
|
||||
-- command, the expected role, and the expected C2 register.
|
||||
--
|
||||
-- Known rows:
|
||||
-- * `gte_store_g4_p3_post_rtps`: post-RTPS polygon-emit slot reads the newest projected screen coordinate from C2_SXY2
|
||||
-- (NOT C2_SXY0; the FIFO side effects do not make SXY0 the newest result).
|
||||
-- * `gte_store_g4_p3_post_rtps`: post-RTPS polygon-emit slot reads the newest projected screen coordinate from C2_SXY2.
|
||||
-- C2_SXY0 is wrong (C2_SXY0 is an older FIFO entry, never the newest post-RTPS result).
|
||||
--
|
||||
-- Unknown `_post_<cmd>` components (a `<name>_post_<cmd>` suffixed component name whose bare `<name>` is not a row key)
|
||||
-- emit ONE `table_gap` info finding so downstream consumers can detect when the canonical contract table is incomplete for an authored atom body.
|
||||
-- Unknown `_post_<cmd>` components (a `<name>_post_<cmd>`-suffixed component whose bare `<name>` is not a row key) emit one
|
||||
-- `table_gap` info finding so downstream consumers can detect when the contract table is incomplete for an authored atom body.
|
||||
--
|
||||
-- Consumers:
|
||||
-- * passes/static_analysis.lua::check_gte_result_position (per-atom CHECK_RULES reader; renders result-position findings).
|
||||
-- * passes/static_analysis.lua::check_gte_result_position (renders result-position findings).
|
||||
-- * passes/static_analysis.lua::emit_table_gap_warning (called once per atom body; surfaces the missing-row diagnostic).
|
||||
-- This table is consumed by the result-position check.
|
||||
M.GTE_COMPONENT_RESULT_CONTRACTS = {
|
||||
-- Post-RTPS g4 p3 store contract: writes the latest screen XY (C2_SXY2) into the primitive's p3 slot.
|
||||
-- Reads from C2_SXY0 would be a semantic mismatch (C2_SXY0 is the OLDEST post-RTPS SXY, not the newest one).
|
||||
-- Reading from C2_SXY0 is a semantic mismatch — C2_SXY0 is the oldest post-RTPS SXY, not the newest one.
|
||||
["gte_store_g4_p3_post_rtps"] = {
|
||||
command = "gte_cmdw_rtps",
|
||||
role = "latest_screen_xy",
|
||||
@@ -1334,15 +1322,14 @@ M.GTE_COMPONENT_RESULT_CONTRACTS = {
|
||||
|
||||
-- Operand-class table for the COP2->GPR load-delay check.
|
||||
--
|
||||
-- Maps each emitting-token ident to the SET of GPR operand positions it READS (not writes).
|
||||
-- Covers the current encoder vocabulary (`code/duffle/mips.h` + `code/duffle/gte.h`);
|
||||
-- expand by adding rows here as new encoders land.
|
||||
-- Maps each emitting-token ident to the set of GPR operand positions it reads.
|
||||
-- Covers the current encoder vocabulary (`code/duffle/mips.h` + `code/duffle/gte.h`); add rows here as new encoders land.
|
||||
--
|
||||
-- Semantics:
|
||||
-- * A "GPR operand position" is the textual slot in the macro's argument list, 1-based; e.g. `load_word(rt, base, off)` has positional operands 1 (rt), 2 (base), 3 (off);
|
||||
-- The table reads operands 1 + 2 + 3 to find what GPRs the macro touches.
|
||||
-- * A "GPR operand position" is the textual slot in the macro's argument list, 1-based; e.g. `load_word(rt, base, off)` has
|
||||
-- positional operands 1 (rt), 2 (base), 3 (off). The table reads operands 1 + 2 + 3 to find what GPRs the macro touches.
|
||||
-- * The check tracks one entry per destination GPR per MFC2/CFC2 event.
|
||||
-- A subsequent event is considered a "use" iff any of its READ operand positions reference that destination GPR's ident (e.g. `R_T0`).
|
||||
-- A subsequent event counts as a "use" iff any of its read operand positions reference that destination GPR's ident (e.g. `R_T0`).
|
||||
-- * Branch delay slots are out of scope (MIPS control-flow; tracked separately).
|
||||
M.OPERAND_READ_POSITIONS = {
|
||||
-- CPU ALU with one or two GPR operands. Reads every GPR operand.
|
||||
@@ -1374,9 +1361,8 @@ M.OPERAND_READ_POSITIONS = {
|
||||
["shift_lright"] = {1, 2},
|
||||
["shift_aright"] = {1, 2},
|
||||
["shift_lleft_self"] = {1},
|
||||
-- Loads: load_word(rt, base, off); the rt operand is the destination (so it's WRITTEN, not read) and base + off are non-GPR operands.
|
||||
-- Treat load_* as NOT reading any GPR operand position (the rt WRITE is not a read for our purposes).
|
||||
-- The single operand in the table for `load_*` is `rt`, but the check treats it as a write, so we leave the read-positions table empty.
|
||||
-- Loads: load_word(rt, base, off); the rt operand is the destination (it's written, not read) and base + off are non-GPR operands.
|
||||
-- The check treats the rt operand as a write, so the read-positions table for `load_*` is empty.
|
||||
["load_word"] = {},
|
||||
["load_half_u"] = {},
|
||||
["load_byte_u"] = {},
|
||||
@@ -1408,9 +1394,9 @@ M.OPERAND_READ_POSITIONS = {
|
||||
["mov_from_low"] = {},
|
||||
["mov_to_high"] = {1},
|
||||
["mov_to_low"] = {1},
|
||||
-- GTE transfers / loads / stores / commands: the relevant table values live in the check itself
|
||||
-- (gte_mv_to_* writes its rt operand, gte_mv_from_* writes its rt operand, and `gte_*` commands are atomic-from-the-CPU-POV once they issue.
|
||||
-- They don't trigger load-delay violations because the CPU holds until the command completes).
|
||||
-- GTE transfers / loads / stores / commands: the relevant table values live in the check itself.
|
||||
-- `gte_mv_to_*` writes its rt operand; `gte_mv_from_*` writes its rt operand; `gte_*` commands are atomic-from-the-CPU-POV
|
||||
-- once they issue (the CPU holds until the command completes, so load-delay violations don't surface here).
|
||||
["gte_mv_from_data_r"] = {},
|
||||
["gte_mv_from_ctrl_r"] = {},
|
||||
["gte_mv_to_data_r"] = {},
|
||||
@@ -1461,9 +1447,10 @@ M.GP0_CMD_BY_SHAPE = {
|
||||
["g4"] = 0x38, ["gt4"] = 0x3C,
|
||||
}
|
||||
|
||||
-- Per-macro prim-buffer contribution
|
||||
-- (NOT .text instruction count this is "how many 32-bit words does this macro write to the primitive being built in main RAM").
|
||||
-- Sum across `mac_format_X_color` + `mac_gte_store_X_post_*` + `mac_insert_ot_tag_X` calls in an atom body must equal GP0_CMD_SIZE[GP0_CMD_BY_SHAPE[shape]].
|
||||
-- Per-macro prim-buffer contribution: how many 32-bit words each macro writes to the primitive being built in main RAM.
|
||||
-- (This counts RAM-side prim-buffer words, not .text instruction words.)
|
||||
-- The sum across `mac_format_X_color` + `mac_gte_store_X_post_*` + `mac_insert_ot_tag_X` calls in an atom body must equal
|
||||
-- `GP0_CMD_SIZE[GP0_CMD_BY_SHAPE[shape]]`.
|
||||
M.GP0_MACRO_CONTRIB = {
|
||||
["mac_format_f3_color"] = 1,
|
||||
["mac_format_g3_color"] = 3,
|
||||
@@ -1477,20 +1464,18 @@ M.GP0_MACRO_CONTRIB = {
|
||||
}
|
||||
|
||||
-- Per-macro cycle cost (best-case, no stalls). Used by the static-analysis pass to emit per-atom cycle budgets.
|
||||
-- The counts cover the EXPANDED instruction sequence the macro emits (NOT just the token it appears as in source).
|
||||
-- For example:
|
||||
-- mac_pack_color_word(off, cmd, r, g, b) emits:
|
||||
-- The counts cover the expanded instruction sequence the macro emits (not just the surface token in source).
|
||||
-- Worked example — `mac_pack_color_word(off, cmd, r, g, b)` expands to:
|
||||
-- load_upper_i(R_AT, (cmd << 8) | b) -- 1 cycle
|
||||
-- or_i_self(R_AT, (g << 8) | r) -- 1 cycle
|
||||
-- store_word(R_AT, R_PrimCursor, off) -- 1 cycle
|
||||
-- = 3 cycles total
|
||||
-- = 3 cycles total
|
||||
--
|
||||
-- mac_yield emits a control-transfer sequence (load_word, add_ui_self, jump_reg, nop)
|
||||
-- which "yields control" the atom body's cycle budget doesn't include the yield's cost (we model it as 0;
|
||||
-- runtime cost becomes part of the NEXT atom's prologue).
|
||||
-- `mac_yield` emits a control-transfer sequence (load_word, add_ui_self, jump_reg, nop). The atom body's cycle budget excludes
|
||||
-- the yield's cost (we model it as 0); the runtime cost lands in the next atom's prologue.
|
||||
--
|
||||
-- GTE command values are the GTE instruction's intrinsic cycles (the latency AFTER any pre-cmd `nop2` has retired).
|
||||
-- When the source emits `nop2, gte_cmdw_X` the nops' cycles are added separately (1+1) plus the gte_cmdw_X value here:
|
||||
-- GTE command values are the GTE instruction's intrinsic cycles — the latency after any pre-cmd `nop2` has retired.
|
||||
-- When the source emits `nop2, gte_cmdw_X`, the nops' cycles are added separately (1+1) plus the gte_cmdw_X value here:
|
||||
-- rtpt = 23 + 2 nops = 25 total cycles (PSX-SPX says 23 cycles for the cmd itself; the nops are pre-fill)
|
||||
-- rtps = 15 + 2 nops = 17 total
|
||||
-- nclip = 8 + 2 nops = 10 total
|
||||
@@ -1499,11 +1484,11 @@ M.GP0_MACRO_CONTRIB = {
|
||||
-- mvmva = 8 + 2 nops = 10 total
|
||||
-- op = 6 (no pre-cmd nops required; atomic)
|
||||
--
|
||||
-- Note: the "total" above is the pre-fill nops + the GTE intrinsic cycles.
|
||||
-- PSX-SPX documents the GTE intrinsic cycles as the total execution time of the command itself (rtpt=23, rtps=15, nclip=8, etc.).
|
||||
-- The pre-fill nops are a codebase convention for retiring preceding C2 writes, not part of the GTE's own execution time.
|
||||
-- See `docs/psx-spx/docs/geometrytransformationenginegte.md` for the canonical per-command cycle counts and `docs/psx-spx/docs/gtepipelinetimings.md`
|
||||
-- for the hardware-verified input-latch boundaries (which show most inputs are safe to clobber after just 0-4 cycles).
|
||||
-- PSX-SPX reports the GTE intrinsic cycles as the total execution time of the command itself (rtpt=23, rtps=15, nclip=8, etc.).
|
||||
-- The pre-fill nops are a codebase convention for retiring preceding C2 writes.
|
||||
-- See `docs/psx-spx/docs/geometrytransformationenginegte.md` for per-command cycle counts and
|
||||
-- `docs/psx-spx/docs/gtepipelinetimings.md` for the hardware-verified input-latch boundaries (most inputs become
|
||||
-- safe to clobber after 0-4 cycles).
|
||||
M.INSTRUCTION_LATENCY = {
|
||||
-- CPU ALU (single-cycle R3000A ops)
|
||||
["nop"] = 1,
|
||||
@@ -1578,7 +1563,7 @@ M.INSTRUCTION_LATENCY = {
|
||||
["gte_cmdw_op"] = 6, -- OP: 6 cycles (PSX-SPX)
|
||||
["gte_cmdw_outer_product"] = 6, -- alias for OP
|
||||
["gte_cmdw_wedge"] = 6, -- alias for OP
|
||||
-- Long-form aliases (same cost as canonical)
|
||||
-- Long-form aliases (same cycle cost as their short form)
|
||||
["gte_cmdw_rotate_translate_perspective_single"] = 15, -- alias for rtps
|
||||
["gte_cmdw_rotate_translate_perspective_triple"] = 23, -- alias for rtpt
|
||||
["gte_cmdw_avg_sort_z4"] = 6, -- alias for avsz4
|
||||
@@ -1628,29 +1613,30 @@ M.UNKNOWN_INSTRUCTION_CYCLES = 1
|
||||
|
||||
-- Hardware-relation policy table.
|
||||
--
|
||||
-- The single forward-analyzer in `passes/static_analysis.lua::analyze_hardware_relations`
|
||||
-- reads every emitted word_event, matches its `encoder` against `row.token`, and:
|
||||
-- The forward walker in `passes/static_analysis.lua::analyze_hardware_relations` reads every emitted word_event,
|
||||
-- matches its `encoder` against `row.token`, and:
|
||||
-- * stages the event as a producer in `atom.paths.forward_state`; or
|
||||
-- * matches it as a consumer against pending producers and records a hazard on `atom.paths.hazards` when the gap is below `visibility.required`.
|
||||
--
|
||||
-- Each row is the contract for one CPU-to-coprocessor transfer semantic
|
||||
-- (the coprocessor-to-CPU path mirrors the same shape). The `reads` / `writes` sub-tables carry the argument positions the analyzer inspects:
|
||||
-- Each row is the contract for one CPU-to-coprocessor transfer semantic (the coprocessor-to-CPU path mirrors the same shape).
|
||||
-- The `reads` / `writes` sub-tables carry the argument positions the analyzer inspects:
|
||||
-- * `writes.arg` is the destination operand (the producer's effect); the analyzer stages this register as a pending producer.
|
||||
-- * `reads` (when present) lists the operand positions the SAME token reads back from hardware; for MTC2 / CTC2 the producer reads the GPR source it is loading from.
|
||||
-- * `reads` (when present) lists the operand positions the same token reads back from hardware; for MTC2 / CTC2 the producer reads the GPR source it is loading from.
|
||||
-- The `fanout_to` field (MTC2-IRGB row only) tells the consumer-match logic which downstream COP2 registers are transitively updated by the write.
|
||||
--
|
||||
-- Visibility semantics:
|
||||
-- * `kind = "post_producer_words"` means the consumer must observe the producer's effect after `required` independent emitted words that are strictly between the producer and the consumer.
|
||||
-- The producer's own emitted slot does NOT retire the relation (per the canonical PSX-SPX rule: "Store delays are counted in numbers of clock cycles (not in numbers of opcodes).
|
||||
-- For 3 cycle delay, one must usuallys insert 3 cached opcodes (or one uncached opcode).").
|
||||
-- * `kind = "post_producer_words"` means the consumer observes the producer's effect after `required` independent emitted words that are
|
||||
-- strictly between the producer and the consumer. The producer's own emitted slot is implicit (it counts as the slot of issue, not toward
|
||||
-- `required`) — per the PSX-SPX rule: "Store delays are counted in numbers of clock cycles (not in numbers of opcodes). For 3 cycle delay,
|
||||
-- one must usually insert 3 cached opcodes (or one uncached opcode)."
|
||||
-- * `required` is the minimum count of intervening emitted words between producer and consumer.
|
||||
-- `required = 0` is permitted (the consumer may sit on the very next slot);
|
||||
-- `required < 0` would mean the consumer may sit on the same slot as the producer and is reserved for future "self-retires" relations.
|
||||
-- `required = 0` permits the consumer on the very next slot; `required < 0` would place the consumer on the same slot as the producer
|
||||
-- and is reserved for future "self-retires" relations.
|
||||
--
|
||||
-- Evidence:
|
||||
-- * `evidence.confidence` is one of `"exact"`, `"conservative"`, `"unknown"`. The severity comes from `violation_kind`;
|
||||
-- a hardware measurement that the vendor caveats may still be `"conservative"` even when the underlying timing is numerically known.
|
||||
-- * `evidence.source` is the canonical upstream reference (file + line range) the row is sourced from. Doc-edits that add new rows must add the source citation here.
|
||||
-- * `evidence.confidence` is one of `"exact"`, `"conservative"`, `"unknown"`. The severity comes from `violation_kind`; a hardware
|
||||
-- measurement that the vendor caveats may still classify as `"conservative"` even when the underlying timing is numerically known.
|
||||
-- * `evidence.source` is the upstream reference (file + line range) the row is sourced from. New rows must carry this citation.
|
||||
--
|
||||
-- Consumers:
|
||||
-- * passes/static_analysis.lua::analyze_hardware_relations (forward walker).
|
||||
@@ -1715,7 +1701,7 @@ M.HARDWARE_RELATIONS = {
|
||||
semantic = "MFC2",
|
||||
token = "gte_mv_from_data_r",
|
||||
direction = "cop2_data_to_gpr",
|
||||
reads = { domain = "cop2.data", arg = 2 },
|
||||
reads = { domain = "cop2.data", arg = 2 },
|
||||
writes = { domain = "gpr", arg = 1 },
|
||||
visibility = { kind = "post_producer_words", required = 1 },
|
||||
evidence = {
|
||||
@@ -1730,7 +1716,7 @@ M.HARDWARE_RELATIONS = {
|
||||
semantic = "CFC2",
|
||||
token = "gte_mv_from_ctrl_r",
|
||||
direction = "cop2_control_to_gpr",
|
||||
reads = { domain = "cop2.ctrl", arg = 2 },
|
||||
reads = { domain = "cop2.ctrl", arg = 2 },
|
||||
writes = { domain = "gpr", arg = 1 },
|
||||
visibility = { kind = "post_producer_words", required = 1 },
|
||||
evidence = {
|
||||
@@ -1748,7 +1734,7 @@ M.HARDWARE_RELATIONS = {
|
||||
semantic = "MFC0",
|
||||
token = "sys_mov_from_cop0",
|
||||
direction = "cop0_control_to_gpr",
|
||||
reads = { domain = "cop0.ctrl", arg = 2 },
|
||||
reads = { domain = "cop0.ctrl", arg = 2 },
|
||||
writes = { domain = "gpr", arg = 1 },
|
||||
visibility = { kind = "post_producer_words", required = 1 },
|
||||
evidence = {
|
||||
@@ -1775,8 +1761,8 @@ M.HARDWARE_RELATIONS = {
|
||||
violation_kind = "info",
|
||||
clear_on_consumer = true,
|
||||
},
|
||||
-- COP2 data register -> memory (SWC2). This is a read of C2 state, not a CPU-to-COP2 write.
|
||||
-- Keep the policy row for direction/provenance, but do not stage it as a later command-input producer.
|
||||
-- COP2 data register -> memory (SWC2). A read of C2 state, not a CPU-to-COP2 write.
|
||||
-- The policy row stays in for direction/provenance; staging it as a later command-input producer is suppressed.
|
||||
{
|
||||
id = "swc2_memory_write",
|
||||
semantic = "SWC2",
|
||||
@@ -1793,13 +1779,13 @@ M.HARDWARE_RELATIONS = {
|
||||
stage = false,
|
||||
},
|
||||
-- MTC0 Status/SR.CU2. The ordinary COP0 store has no general store-delay relation;
|
||||
-- This row is consumed by the dedicated CU2 transition logic in the same forward walk and is therefore not staged in `pending`.
|
||||
-- this row feeds the dedicated CU2 transition logic in the same forward walk and is therefore not staged in `pending`.
|
||||
{
|
||||
id = "mtc0_cu2_visibility",
|
||||
semantic = "MTC0",
|
||||
token = "sys_mov_to_cop0",
|
||||
direction = "gpr_to_cop0_status",
|
||||
reads = { domain = "gpr", arg = 1 },
|
||||
reads = { domain = "gpr", arg = 1 },
|
||||
writes = { domain = "cop0.status", arg = 2 },
|
||||
status_register = 12,
|
||||
visibility = { kind = "post_producer_words", required = 2 },
|
||||
@@ -1832,18 +1818,18 @@ M.CU2_TRANSITION_POLICY = {
|
||||
-- Maps every CPU/GTE encoder used in production atoms and the focused transfer-hazard tests to its actual GPR operand effects.
|
||||
-- The analyzer applies this table to `atom.paths.forward_state.gpr_values`:
|
||||
-- * a write to a GPR invalidates its constant;
|
||||
-- * a constant-producing transform re-establishes a constant when its inputs are constant (the lattice for `gpr_values` is closed:
|
||||
-- `{kind="unknown"}` and `{kind="constant", value=<U4>}`).
|
||||
-- * a constant-producing transform re-establishes a constant when its inputs are constant
|
||||
-- (the `gpr_values` lattice is closed: `{kind="unknown"}` and `{kind="constant", value=<U4>}`).
|
||||
--
|
||||
-- The schema is:
|
||||
-- reads = {pos1, pos2, ...} -- 1-based argument positions that are GPR reads.
|
||||
-- Schema:
|
||||
-- reads = {pos1, pos2, ...} -- 1-based argument positions that are GPR reads.
|
||||
-- writes = {pos1, pos2, ...} -- 1-based argument positions that are GPR writes.
|
||||
-- The argument positions refer to `word_event.args` (the top-level comma-split args of the emitting token, parsed by `tokenize_body`).
|
||||
-- Operands that are numeric literals, `0x` hex literals, or `U4`/`S4` type keywords are not GPR operand positions and are not listed.
|
||||
-- Numeric literals, `0x` hex literals, and `U4`/`S4` type keywords are not GPR operand positions.
|
||||
--
|
||||
-- Encoders not listed here are treated as "unknown writers" for any GPR they touch;
|
||||
-- Wknown writers invalidate `forward_state.gpr_values` for every GPR operand they touch (the analyzer cannot assume the result is a constant).
|
||||
-- This is deliberately conservative: a row missing for a writer means "we do not know what value the GPR now holds" rather than "the GPR keeps its previous constant".
|
||||
-- Encoders absent from this table are treated as "unknown writers" for every GPR they touch. Unknown writers invalidate
|
||||
-- `forward_state.gpr_values` for those operands — the analyzer cannot assume the result is a constant.
|
||||
-- The shape is deliberately conservative: a row missing for a writer means "we do not know what value the GPR now holds".
|
||||
--
|
||||
-- Consumers:
|
||||
-- * passes/static_analysis.lua::analyze_hardware_relations (forward walker).
|
||||
@@ -1994,10 +1980,9 @@ M.GPR_VALUE_RULES = {
|
||||
-- Control-transfer (branch/jump/call) delay-slot policy table.
|
||||
--
|
||||
-- Used by the emitted-word delay-slot check to identify which emitted machine-word idents are control transfers whose next emitted word is the hardware delay slot.
|
||||
-- One table row per emitted encoder; the `family` field is informational (informational only;
|
||||
-- The check matches by `event.ident` against the row keys).
|
||||
-- `suppress_arg1` (when present) lists first-arg values that should NOT emit a finding even when the next emitted word is
|
||||
-- `nop` or absent — e.g. the fixed `mac_yield()` handshake uses `jump_reg(R_AtomJmp), nop` and is intentionally suppressed.
|
||||
-- One table row per emitted encoder; the `family` field is informational. The check matches by `event.ident` against the row keys.
|
||||
-- `suppress_arg1` (when present) lists first-arg values that suppress the finding even when the next emitted word is `nop` or absent
|
||||
-- — for example, the fixed `mac_yield()` handshake uses `jump_reg(R_AtomJmp), nop` and is suppressed so the check stays signal-only.
|
||||
--
|
||||
-- Consumers:
|
||||
-- * passes/static_analysis.lua::check_control_transfer_delay_slot_use
|
||||
@@ -2025,8 +2010,8 @@ M.CONTROL_TRANSFER_DELAY_SLOT_POLICIES = {
|
||||
-- Section 8: Cross-source component-body index + word-event expansion
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
--
|
||||
-- Two pure helpers that supersede the per-pass local component-body builders (`atoms_source_map.build_cross_source_component_body_index`)
|
||||
-- and provide the shared, memoized "semantic emitted-word event stream" every downstream pass can read from without re-walking the pre-tokenized bodies.
|
||||
-- Shared, memoized helpers: a single emitted-word event stream that every downstream pass reads from,
|
||||
-- built once from the pre-tokenized bodies.
|
||||
|
||||
--- @class ComponentBodyEntry
|
||||
--- @field body_tokens table -- pre-tokenized {{tok=string, rel=integer}, ...}
|
||||
@@ -2036,10 +2021,8 @@ M.CONTROL_TRANSFER_DELAY_SLOT_POLICIES = {
|
||||
--- @field declaration integer -- 1-based line number of the MipsAtomComp_(ac_X) declaration
|
||||
--- @field kind string -- "comp_bare" | "comp_proc"
|
||||
|
||||
-- The cross-source component-body index is owned by the canonical corpus
|
||||
-- (`corpus.component_body_index`, populated by `passes/components.lua`).
|
||||
-- Consumers (`passes/static_analysis.lua`, `passes/emission_model.lua`) read it directly;
|
||||
-- No per-pass memoization helper is needed.
|
||||
-- The cross-source component-body index is owned by the corpus (`corpus.component_body_index`, populated by `passes/components.lua`).
|
||||
-- Consumers (`passes/static_analysis.lua`, `passes/emission_model.lua`) read it directly; per-pass memoization helpers stay out of scope.
|
||||
|
||||
-- ASCII byte constants used by split_top_level_args (kept local to keep Section 8 self-contained).
|
||||
local E_BYTE_OPEN_PAREN = 0x28
|
||||
@@ -2110,45 +2093,66 @@ local E_MAC_PREFIX_LEN = 4
|
||||
---
|
||||
--- Semantics (one event per emitted machine word):
|
||||
--- * **Direct one-word encoders** (`load_word`, `add_ui`, `nop`, `gte_lw`, ...): one event with `ident` = leading ident, `args` = parsed top-level args.
|
||||
--- * **`nop2`** (2-word pseudo-instruction): two events, BOTH with `ident = "nop"` so the canonical "this slot is a no-op" semantic is visible to downstream analyses.
|
||||
--- * **Any other N-word token** in `word_counts` (e.g. `mask_upper` = 2, `load_imm_2w` = 2): N events sharing the same `ident` + `args` so useful CPU words retire slots in the cycle budget.
|
||||
--- * **`nop2`** (2-word pseudo-instruction): two events, both with `ident = "nop"` so the recognized "this slot is a no-op" semantic is visible to downstream analyses.
|
||||
--- * **Any other N-word token** in `word_counts`: N events sharing the same `ident` + `args` so useful CPU words retire slots in the cycle budget.
|
||||
--- * **Known `mac_X(...)` calls**: recursively expand the indexed component body, including nested components. Every event from the expansion carries:
|
||||
--- - `source` / `line` = the COMPONENT'S source path + the line of the token within the component body (i.e. "definition site").
|
||||
--- - `call_source` / `call_line` = the ROOT atom's source path + call-site line, PRESERVED across recursion (nested-nested events still point at the original root, not at an intermediate component).
|
||||
--- * **Unknown `mac_X`** (not in `component_index`): fall back to `word_counts[ident]` if present; otherwise emit exactly one opaque event so the cycle budget still accounts for the word.
|
||||
--- - `call_source` / `call_line` = the ROOT atom's source path + call-site line, PRESERVED across recursion so nested events still point at the original root.
|
||||
--- * **Unknown `mac_X`** (not in `component_index`): fall back to `word_counts[ident]` if present; otherwise emit one opaque event so the cycle budget accounts for the word.
|
||||
--- * **Marker tokens** (`atom_label(...)` / `atom_offset(...)`): zero events (they are pure metaprogram hints, not emitted machine words).
|
||||
---
|
||||
--- Cycle protection: a per-expansion `visiting` set tracks components currently on the expansion stack; a re-entry produces a deterministic `{kind = "cycle", ...}` error and aborts that branch (does NOT hang, does NOT recurse).
|
||||
---
|
||||
--- Pure: does NOT mutate `body_entry`, `component_index`, or `word_counts`. Memoization is the caller's responsibility (callers that want it precomputed for many atoms should memoize `word_events` / `word_event_errors` per atom).
|
||||
--- Pure: reads `body_entry` / `component_index` / `word_counts`. Memoization is the caller's responsibility.
|
||||
--- Callers wanting `word_events` / `word_event_errors` precomputed for many atoms should memoize them per atom.
|
||||
--- @param body_entry table -- `{body_tokens, body_off, line_of, source, declaration}` (declaration = root atom's atom.line)
|
||||
--- @param component_index table -- the bare-name → ComponentBodyEntry map from M.get_component_body_index
|
||||
--- @param word_counts table -- macro name → emitted-word count (from `ctx.shared.word_counts`)
|
||||
--- @return WordEvent[], WordEventError[]
|
||||
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
-- Section 11: project_emission (canonical per-atom emission projection)
|
||||
-- Section 11: project_emission (per-atom emission projection)
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
--
|
||||
-- Canonical per-atom emission projection is owned by `passes/emission_model.lua`.
|
||||
-- The projection is built from the root atom body only (no nested component expansion at this stage);
|
||||
-- Invocation ancestry recursively expands nested components.
|
||||
-- The items stream is the single ordered source of truth; `word_events` and `markers` are dense views over it (never a separate walk).
|
||||
-- Per-atom emission projection is owned by `passes/emission_model.lua`.
|
||||
-- The projection is built from the root atom body only; invocation ancestry recursively expands nested components.
|
||||
-- The items stream is the single ordered source of truth; `word_events` and `markers` are dense views over it.
|
||||
--
|
||||
-- The helper below operates on a body string (not a body_entry) so the canonical pass can call it without depending on the older SourceScan / body_off conventions.
|
||||
-- The helper below operates on a body string (not a body_entry) so the pass can call it without depending on the older SourceScan / body_off conventions.
|
||||
-- component_index argument is reserved for recursive component expansion.
|
||||
-- word_counts table is the canonical authored-metadata + current-component count table.
|
||||
-- word_counts table is authored-metadata + current-component count table.
|
||||
|
||||
--- @class EmissionProjection
|
||||
--- @field items table[] -- ordered stream of word|label|offset|invoke_begin|invoke_end
|
||||
--- @field word_events table[] -- dense view of items where kind == "word"
|
||||
--- @field markers table[] -- dense view of items where kind == "label"|"offset"
|
||||
--- @field invocations table[] -- dense view of items where kind == "invoke_begin"|"invoke_end"
|
||||
--- @field invocations InvocationRecord[] -- dense view of items where kind == "invoke_begin"|"invoke_end"
|
||||
--- @field errors table[] -- token-resolution failures surfaced without fail-loud
|
||||
--- @field warnings table[] -- opaque warnings (e.g. unknown uncounted macro)
|
||||
|
||||
-- Internal recursive walker. Single source of truth for the items stream;
|
||||
-- `word_events`, `markers`, `invocations`, `errors`, `warnings` are dense views / side outputs derived while appending `items`.
|
||||
--- @class InvocationRecord
|
||||
--- Lives at `atom.paths.invocations[*]`. Constructed once at the single invocation-construction site
|
||||
--- (`emit_invoke_begin` inside `_project_emission_inner`); `invoke_begin` / `invoke_end` markers in the items stream share the same `id`.
|
||||
--- @field id integer -- 1-based, monotonic per-atom invocation id (0 is reserved for "no open invocation")
|
||||
--- @field parent_id integer -- 0 for the outermost (root) call; otherwise the id of the immediately enclosing invocation
|
||||
--- @field kind string -- "comp_bare" | "comp_proc" (component form that triggered the expansion)
|
||||
--- @field component_name string -- the bare component name without the `mac_` prefix
|
||||
--- @field call_text string -- the immediate `mac_X(...)` token text (or root call text for the outermost entry)
|
||||
--- @field root_call_text string -- the IMMUTABLE outermost `mac_X(...)` token text for every word emitted in this call's expansion
|
||||
--- @field call_path string -- source path of the call site (root atom source for direct calls, component source for nested expansions)
|
||||
--- @field call_line integer -- source line of the call site
|
||||
--- @field def_path string -- source path of the component definition
|
||||
--- @field def_line integer -- source line of the component declaration
|
||||
--- @field start_pos integer -- 0-based emitted-word position of the FIRST word inside this invocation (the value of `word_idx` AT `emit_invoke_begin` time, BEFORE the first word is emitted). Words emitted inside this invocation occupy `start_pos..start_pos+#body_lines-1` (inclusive, 0-based). Downstream DWARF/provenance consumers MUST read this; do NOT reconstruct it from `start_word` (which is the 1-based items index including `invoke_begin`/`invoke_end` markers).
|
||||
--- @field end_pos integer -- 0-based position of the LAST word inside this invocation (set by `emit_invoke_end` to `word_idx - 1` AFTER all body words are emitted).
|
||||
--- @field start_word integer -- 1-based items index of the `invoke_begin` item
|
||||
--- @field end_word integer -- 1-based items index of the `invoke_end` item (set by `emit_invoke_end`)
|
||||
--- @field word_count integer -- number of `word` items emitted between `start_word` and `end_word` (inclusive)
|
||||
--- @field debug_skip boolean -- `debug_skip` stamp; true iff `corpus.components[name].debug_skip` is true at construction. Always boolean (never `nil`).
|
||||
--- @field errors table[] -- per-invocation construction errors (cycle / count_mismatch); does not include pass-level errors
|
||||
|
||||
-- Internal recursive walker. The items stream holds every emitted event in order; `word_events`, `markers`,
|
||||
-- `invocations`, `errors`, `warnings` are dense views / side outputs appended alongside.
|
||||
--
|
||||
-- Output rules:
|
||||
-- * `word` items record: `invocation_ids` (innermost last) and `outermost_invocation_id` (0 if no invocation is open).
|
||||
@@ -2161,7 +2165,7 @@ local E_MAC_PREFIX_LEN = 4
|
||||
-- * Unknown uncounted macros emit one opaque word + one warning. Unknown metadata-backed macros (entry in `word_counts`) emit the declared word count, no warning.
|
||||
-- * Cycle detection uses an active DFS stack (`visiting`); a cycle appends a construction error to BOTH the projection errors and the cycle invocation's own errors,
|
||||
-- then breaks out without recursing (the cycle entry still receives an invocation ID + paired `invoke_begin` / `invoke_end` items, so the boundary invariant is preserved).
|
||||
-- * Component declared-count mismatch (declared vs. measured) is a construction error (kind = "count_mismatch"); it is recorded on the invocation record and the pass-level errors list.
|
||||
-- * Component declared-count mismatch (declared vs. measured) is a construction error (kind = "count_mismatch"); recorded on the invocation record and pass-level errors list.
|
||||
-- * Final boundary check: if any invocation is still open at end of walk, surface a "unbalanced" construction error.
|
||||
local function _project_emission_inner(root_body_entry, ctx_table)
|
||||
local items = {}
|
||||
@@ -2224,8 +2228,7 @@ local function _project_emission_inner(root_body_entry, ctx_table)
|
||||
immediate_call_text, root_call_text_w)
|
||||
local inv_ids = open_invocation_ids_snapshot()
|
||||
local outermost = inv_ids[1] or 0
|
||||
-- Markers carry the open invocation stack snapshot but do NOT record `call_text` / `root_call_text` —
|
||||
-- markers are zero-width and never participate in the per-word call-site attribution.
|
||||
-- Markers carry the open invocation stack snapshot. `call_text` / `root_call_text` belong to words, not markers — markers are zero-width and skip per-word call-site attribution.
|
||||
local it = {
|
||||
kind = kind,
|
||||
name = name,
|
||||
@@ -2284,6 +2287,25 @@ local function _project_emission_inner(root_body_entry, ctx_table)
|
||||
local function emit_invoke_begin(inv_kind, component_name, call_text,
|
||||
root_call_text, call_path, call_line)
|
||||
next_inv_id = next_inv_id + 1
|
||||
-- Invocation-level debug_skip stamp: Emission pass owns `atom.paths.invocations[*].debug_skip`.
|
||||
-- The stamp is resolved from the `corpus.components[name]` registry (passed in via `ctx_table.components` by `emission_model.run`),
|
||||
-- NOT from a parallel skip map, source-text re-parse, or second pass over `invocations`.
|
||||
-- Unmarked components stamp `false` (not `nil`) so consumers can dispatch on the boolean without nil checks.
|
||||
--
|
||||
-- The walker has already found the component body in `ctx_table.component_index[component_name]`, so the matching entry MUST exist in `ctx_table.components[component_name]`
|
||||
-- (both registries are populated from the same source by the components pass).
|
||||
-- A missing entry is a corpus-plumbing bug; we fail loudly here rather than silently stamp `false` and mask the regression.
|
||||
local components = ctx_table.components
|
||||
local component_def = components and components[component_name] or nil
|
||||
if not component_def then
|
||||
error("duffle.emit_invoke_begin: component " .. string.format("%q", component_name)
|
||||
.. " is present in `component_index` (the walker matched a `mac_" .. component_name .. "()` call) but absent from `components` (the canonical corpus.components registry). "
|
||||
.. "This is a corpus-plumbing bug — the components pass must populate corpus.components[name] for every component it puts in corpus.component_body_index[name]. "
|
||||
.. "The emission pass refuses to silently stamp `debug_skip = false` for a missing registry entry."
|
||||
, 0
|
||||
)
|
||||
end
|
||||
local debug_skip_stamp = component_def.debug_skip == true
|
||||
local inv = {
|
||||
id = next_inv_id,
|
||||
parent_id = 0, -- patched below by caller
|
||||
@@ -2295,29 +2317,39 @@ local function _project_emission_inner(root_body_entry, ctx_table)
|
||||
call_line = call_line,
|
||||
def_path = nil, -- patched below after component lookup
|
||||
def_line = nil,
|
||||
-- 0-based emitted-word position. `word_idx` is the monotonic 0-based counter of `word` items emitted so far in this walk —
|
||||
-- BEFORE this invocation's first word is emitted, it equals the position of the first word inside the invocation.
|
||||
-- `start_word` (1-based items index of `invoke_begin`) is kept for items-walking consumers (Annotation pass bounds checks),
|
||||
-- but DWARF / provenance rows MUST read `start_pos` because those rows are 1-based over the dense `word_events` stream (which has no `invoke_begin` items).
|
||||
start_pos = word_idx,
|
||||
start_word = #items + 1, -- 1-based items index of invoke_begin
|
||||
end_word = nil, -- patched by emit_invoke_end
|
||||
end_pos = nil, -- patched by emit_invoke_end
|
||||
end_word = nil, -- patched by emit_invoke_end
|
||||
word_count = 0,
|
||||
debug_skip = debug_skip_stamp,
|
||||
errors = {},
|
||||
}
|
||||
invocations[#invocations + 1] = inv
|
||||
items[#items + 1] = {
|
||||
kind = "invoke_begin",
|
||||
invocation_id = inv.id,
|
||||
word_index = word_idx,
|
||||
invocation_ids = open_invocation_ids_snapshot(),
|
||||
items [#items + 1] = {
|
||||
kind = "invoke_begin",
|
||||
invocation_id = inv.id,
|
||||
word_index = word_idx,
|
||||
invocation_ids = open_invocation_ids_snapshot(),
|
||||
}
|
||||
invocation_stack[#invocation_stack + 1] = inv
|
||||
return inv
|
||||
end
|
||||
|
||||
local function emit_invoke_end(inv)
|
||||
inv.end_word = #items + 1 -- 1-based items index of invoke_end
|
||||
-- 0-based emitted-word position of the LAST word inside this invocation.
|
||||
-- After the last body word was emitted, `word_idx` was incremented past it, so `word_idx - 1` is the 0-based position of the last word.
|
||||
inv.end_pos = word_idx - 1
|
||||
inv.end_word = #items + 1 -- 1-based items index of invoke_end
|
||||
items[#items + 1] = {
|
||||
kind = "invoke_end",
|
||||
invocation_id = inv.id,
|
||||
word_index = word_idx,
|
||||
invocation_ids = open_invocation_ids_snapshot(),
|
||||
kind = "invoke_end",
|
||||
invocation_id = inv.id,
|
||||
word_index = word_idx,
|
||||
invocation_ids = open_invocation_ids_snapshot(),
|
||||
}
|
||||
for i = #invocation_stack, 1, -1 do
|
||||
if invocation_stack[i] == inv then
|
||||
@@ -2430,7 +2462,7 @@ local function _project_emission_inner(root_body_entry, ctx_table)
|
||||
end
|
||||
inv.word_count = wc_inside
|
||||
-- count_mismatch is a construction error: word_counts["mac_X"] is the declared count populated by the components pass;
|
||||
-- we compare against the measured word count.
|
||||
-- We compare against the measured word count.
|
||||
local declared = ctx_table.word_counts["mac_" .. bare]
|
||||
if declared and wc_inside ~= declared then
|
||||
local err = {
|
||||
@@ -2490,7 +2522,7 @@ local function _project_emission_inner(root_body_entry, ctx_table)
|
||||
}
|
||||
end
|
||||
|
||||
--- Project a body string into the canonical per-atom emission projection.
|
||||
--- Project a body string into the per-atom emission projection.
|
||||
---
|
||||
--- Semantics:
|
||||
--- * Direct one-word tokens (`nop`, `add_ui`, ...): one `word` item, encoder = ident, word_count = 1.
|
||||
@@ -2507,17 +2539,33 @@ end
|
||||
--- Every emitted `word` carries: `i` (0-based word index), `encoder`, `args` (top-level args), `def_path`, `def_line`,
|
||||
--- `call_text` (the immediate token spelling), `root_call_text` (outermost `mac_X(...)` text), `word_count` (always 1),
|
||||
--- `invocation_ids` (innermost last), `outermost_invocation_id`.
|
||||
--- Markers carry: `kind`, `name`, `line`, `word_index`, `target` (only for offset kind), plus `invocation_ids` / `outermost_invocation_id` for the open invocation stack at that word.
|
||||
--- Markers carry: `kind`, `name`, `line`, `word_index`, `target` (only for offset kind), plus `invocation_ids` / `outermost_invocation_id`
|
||||
--- for the open invocation stack at that word.
|
||||
---
|
||||
--- @param body_text string -- the raw atom body string
|
||||
--- @param component_index table -- bare-name → component record (corpus.component_body_index)
|
||||
--- @param word_counts table -- macro name → emitted word count
|
||||
--- @param components table -- bare-name → component definition (corpus.components); REQUIRED — consumed at the invocation-construction site to stamp
|
||||
--- `invocation.debug_skip`. A missing or non-table `components` raises a fail-loud error rather than silently falling back.
|
||||
--- @return EmissionProjection
|
||||
function M.project_emission(body_text, component_index, word_counts)
|
||||
-- Project nested invocation ancestry and construction failures.
|
||||
-- The public surface remains `M.project_emission(body_text, ...)`;
|
||||
-- the recursive walk is delegated to `_project_emission_inner` so that component bodies (which arrive as `{body_tokens, body_off,
|
||||
-- line_of, source, declaration}` records from `corpus.component_body_index`) re-enter the same walker with the same shared output state.
|
||||
function M.project_emission(body_text, component_index, word_counts, components)
|
||||
-- The recursive walk delegates to `_project_emission_inner` so component bodies (which arrive as
|
||||
-- `{body_tokens, body_off, line_of, source, declaration}` records from `corpus.component_body_index`)
|
||||
-- re-enter the same walker with the same shared output state.
|
||||
--
|
||||
-- The walker is body-relative: it builds `line_of` from `body_text` and stamps body-relative line numbers (1..N)
|
||||
-- into `item.line` and `invocation.call_line`. `passes/emission_model.lua::stamp_root_provenance` performs the single
|
||||
-- conversion from body-relative to physical source line at the close site, using the source's `line_of` closure that
|
||||
-- the pass forwarded. One owner of the line state.
|
||||
if type(components) ~= "table" then
|
||||
error("duffle.project_emission: `components` is required "
|
||||
.. "(bare-name -> component definition, e.g. corpus.components); "
|
||||
.. "got " .. type(components) .. ". "
|
||||
.. "The emission pass MUST forward the corpus registry "
|
||||
.. "so the invocation-construction site can stamp `debug_skip` "
|
||||
.. "without a second pass, source parse, or parallel lookup.",
|
||||
0)
|
||||
end
|
||||
|
||||
if type(body_text) ~= "string" or body_text == "" then
|
||||
-- Empty body: still return a valid (empty) projection.
|
||||
@@ -2531,17 +2579,18 @@ function M.project_emission(body_text, component_index, word_counts)
|
||||
}
|
||||
end
|
||||
|
||||
local tokens = M.tokenize_body(body_text)
|
||||
local line_of = M.LineIndex(body_text)
|
||||
local tokens = M.tokenize_body(body_text)
|
||||
return _project_emission_inner({
|
||||
body_tokens = tokens,
|
||||
body_off = 0,
|
||||
line_of = line_of,
|
||||
line_of = M.LineIndex(body_text),
|
||||
source = "",
|
||||
declaration = 0,
|
||||
}, {
|
||||
},
|
||||
{
|
||||
component_index = component_index or {},
|
||||
word_counts = word_counts or {},
|
||||
components = components,
|
||||
})
|
||||
end
|
||||
|
||||
|
||||
@@ -28,11 +28,6 @@ param(
|
||||
|
||||
$ErrorActionPreference = 'Stop'
|
||||
|
||||
$gdbInitPath = [System.IO.Path]::GetFullPath((Join-Path $PSScriptRoot '..\build\gen\hello_gte.gdbinit'))
|
||||
if (-not (Test-Path -LiteralPath $gdbInitPath -PathType Leaf)) {
|
||||
Write-Warning "Generated GDB skip sidecar missing (non-fatal): $gdbInitPath. Run the GTE build to regenerate it; debugger launch will continue without generated skip-over commands."
|
||||
}
|
||||
|
||||
# ── Pre-checks ──
|
||||
foreach ($p in @($PcsxPath, $ExePath, $HelperZip)) {
|
||||
if (-not (Test-Path $p)) {
|
||||
|
||||
@@ -1,21 +1,20 @@
|
||||
--- passes/annotation.lua — Atom-annotation DSL validator.
|
||||
---
|
||||
--- Validates `MipsAtom_(name) atom_info(atom_bind(Binds_X), atom_reads(...), atom_writes(...)) { ... }` declarations in source files.
|
||||
--- Also reads: `Binds_*` struct declarations (`typedef Struct_(Binds_X) { ... };`)
|
||||
--- Also reads `Binds_*` struct declarations (`typedef Struct_(Binds_X) { ... };`).
|
||||
---
|
||||
--- Source scanning: done ONCE upstream by `duffle.scan_source()` (ps1_meta.lua pre-scans each source and stashes the result in `src.scan`).
|
||||
--- `duffle.scan_source()` scans each source once upstream; `ps1_meta.lua` stores that result in `src.scan`.
|
||||
---
|
||||
--- Writes:
|
||||
--- - `<ctx.out_root>/<dir_basename>.errors.h` — one per module, with `#error` directives on findings (the C compile will surface the error)
|
||||
--- - The annotations.txt report is rendered by `passes/report.lua` from the canonical `corpus.sources_by_dir` projection (re-validating each source via `M.validate()`).
|
||||
--- Ownership: the canonical `ctx.shared.corpus` supplies cross-source registries, while each `src.scan` supplies its source's declarations and bodies.
|
||||
--- A context without `ctx.shared.corpus` is rejected with an explicit canonical-corpus message.
|
||||
---
|
||||
--- **Conventions**: tabs (1/level), EmmyLua annotations, no regex, Lua 5.3 compatible
|
||||
--- Writes `<ctx.out_root>/<dir_basename>.errors.h` once per module, with `#error` directives for findings that the C compile surfaces.
|
||||
--- `passes/report.lua` renders annotations.txt from `corpus.sources_by_dir`, re-validating each source through `M.validate()`.
|
||||
---
|
||||
--- **Conventions**: tabs (1/level), EmmyLua annotations, no regex, Lua 5.3 compatible.
|
||||
|
||||
-- Bootstrap: same as entry scripts. See `ps1_meta.lua` for the rationale.
|
||||
-- Bootstrap: load `scripts/duffle_paths.lua` (sets package.path + package.cpath).
|
||||
-- Uses `debug.getinfo` to find this file's own directory, so it works both standalone and when require'd from the orchestrator.
|
||||
-- Bootstrap: load `duffle_paths.lua` via `debug.getinfo(1, "S").source` (works both standalone + when require'd).
|
||||
-- duffle_paths.lua sets package.path then returns `require("duffle")` at the bottom, so the dofile value IS the duffle module.
|
||||
-- Bootstrap follows the entry scripts; `scripts/duffle_paths.lua` sets package.path and package.cpath. See `ps1_meta.lua` for the rationale.
|
||||
-- `debug.getinfo(1, "S").source` locates this file for standalone and orchestrated runs, then `duffle_paths.lua` returns the loaded `duffle` module.
|
||||
local _bootstrap_dir = debug.getinfo(1, "S").source:match("^@?(.*[/\\])") or "./"
|
||||
local duffle = dofile(_bootstrap_dir .. "../duffle_paths.lua")
|
||||
local write_file = duffle.write_file
|
||||
@@ -62,15 +61,15 @@ local ensure_dir = duffle.ensure_dir
|
||||
--- @field writes string[] -- R_* names (write targets)
|
||||
--- @field errors string[]|nil -- parse-time errors from scan_source (atom_info body malformed)
|
||||
|
||||
--- @class SkipOverMarker -- sub-shape of scan_source.lua's @class SkipOverMarker
|
||||
--- @field marker_kind string -- exact marker ident (always "atom_dbg_skip_over")
|
||||
--- @class DebugSkipMarker -- sub-shape of scan_source.lua's @class DebugSkipMarker
|
||||
--- @field marker_kind string -- exact marker ident read from source. Only "atom_dbg_skip" (bare) is positive.
|
||||
--- @field marker_line integer
|
||||
--- @field args string|nil -- trimmed text inside the parens (nil when has_parens is false)
|
||||
--- @field has_parens boolean
|
||||
--- @field is_bare boolean -- true iff marker_kind == "atom_dbg_skip" AND has_parens == false (the only positive form)
|
||||
--- @field pending boolean -- true while awaiting the following declaration
|
||||
--- @field superseded_by_marker_line integer|nil -- set on a marker that was bumped out of the pending slot
|
||||
--- @field target_kind string|nil -- "atom" | "comp_bare" | "comp_proc" | "unrelated" once observed
|
||||
--- @field declaration_line integer|nil
|
||||
|
||||
--- @class Finding
|
||||
--- @field line integer -- source line (or 0 for pass-level)
|
||||
@@ -104,10 +103,8 @@ local ensure_dir = duffle.ensure_dir
|
||||
-- Per-check functions (the CHECK_RULES table's payload)
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
--
|
||||
-- Each check has a uniform `append_to_findings` shape (errors[] / warnings[] / info[]).
|
||||
-- The dispatcher in `validate()` decides which findings list each check writes to — by convention,
|
||||
-- "existence" checks (declaration must exist, struct must exist) write errors[]; "shape" checks (writes/reads must be wave-context) write warnings[].
|
||||
-- The `macro_word_drift` check writes both errors[] (missing/mismatch) and info[] (match).
|
||||
--- The dispatcher in `validate()` routes each result by convention: existence checks write errors[] and shape checks write warnings[].
|
||||
--- `macro_word_drift` writes errors[] for missing or mismatched metadata and info[] for a match.
|
||||
|
||||
--- Check: every annotated atom must have a matching MipsAtom_(name) declaration.
|
||||
--- @param a AtomAnnotation
|
||||
@@ -138,9 +135,7 @@ local function check_unique_annotation(pipe_ctx, findings)
|
||||
end
|
||||
|
||||
--- Check: BIND atoms must reference a real Binds_* struct.
|
||||
--- Emitting a warning here keeps the annotation pass from being stop-on-error for the common test-fixture case,
|
||||
--- while still surfacing the issue in the report.
|
||||
--- The static-analysis report remains the source of truth for build-stopping errors.
|
||||
--- I keep this as a warning so the annotation pass can report the common test-fixture case; `check_abi_handoff` in static analysis supplies the build-stopping error.
|
||||
--- @param a AtomAnnotation
|
||||
--- @param pipe_ctx PipeCtx
|
||||
--- @param findings Findings
|
||||
@@ -182,9 +177,8 @@ local function check_macro_word_drift(m, wc, findings)
|
||||
}
|
||||
end
|
||||
|
||||
--- Check: atom_dbg_reg_default(R_X, <type>) must target a register declared as a debug-visible alias in `pipe_ctx.register_alias_registry`,
|
||||
--- with a type name found in `pipe_ctx.type_name_registry`.
|
||||
--- Pointer depth is still bounded to 0 or 1. Duplicate defaults are still detected.
|
||||
--- Check: atom_dbg_reg_default(R_X, <type>) targets an alias in `pipe_ctx.register_alias_registry` and a type in `pipe_ctx.type_name_registry`.
|
||||
--- Pointer depth remains bounded to 0 or 1, and duplicate defaults remain errors.
|
||||
--- @param _src SourceFile -- unused (kept for the per_source shape)
|
||||
--- @param pipe_ctx PipeCtx
|
||||
--- @param findings Findings
|
||||
@@ -233,10 +227,8 @@ local function check_semantic_reg_defaults(_src, pipe_ctx, findings)
|
||||
end
|
||||
end
|
||||
|
||||
--- Check: atom_reg_types(R_X, <type>) entries must point to a register declared in `pipe_ctx.register_alias_registry`, with a type name found in `pipe_ctx.type_name_registry`.
|
||||
--- The alias ident `R_<n>` now encodes the GPR identity only for entries that are explicitly opted in via the bare `atom_reg` marker.
|
||||
--- R_T0..R_T3 are intentionally NOT auto-included (per the prototype principle: no auto-include of wave-context; explicit opt-in only).
|
||||
--- The check fires for any R_T0..R_T3 reference that hasn't been opted in via `#define atom_reg`.
|
||||
--- Check: atom_reg_types(R_X, <type>) entries target an alias in `pipe_ctx.register_alias_registry` and a type in `pipe_ctx.type_name_registry`.
|
||||
--- A bare `atom_reg` marker opts the `R_<n>` alias into GPR identity; references to R_T0..R_T3 require the same explicit marker.
|
||||
--- @param _src SourceFile
|
||||
--- @param pipe_ctx PipeCtx
|
||||
--- @param findings Findings
|
||||
@@ -267,7 +259,7 @@ local function check_atom_reg_types(_src, pipe_ctx, findings)
|
||||
end
|
||||
end
|
||||
|
||||
--- Check: atom_view(Binds_X) entries must reference a real Binds_* struct and that struct must declare at least one field.
|
||||
--- Check: atom_view(Binds_X) entries reference a Binds_* struct with at least one field.
|
||||
--- @param _src SourceFile
|
||||
--- @param pipe_ctx PipeCtx
|
||||
--- @param findings Findings
|
||||
@@ -296,8 +288,7 @@ local function check_atom_view_layout(_src, pipe_ctx, findings)
|
||||
end
|
||||
end
|
||||
|
||||
--- Check: Binds_* structs may not have duplicate field names
|
||||
--- (they would defeat the typed-field name lookup that atom_view exposes in gdb).
|
||||
--- Check: Binds_* structs require unique field names because atom_view uses those names for typed-field lookup in gdb.
|
||||
--- @param _src SourceFile
|
||||
--- @param pipe_ctx PipeCtx
|
||||
--- @param findings Findings
|
||||
@@ -320,26 +311,30 @@ local function check_binds_no_duplicate_fields(_src, pipe_ctx, findings)
|
||||
end
|
||||
end
|
||||
|
||||
-- Check: skip-over markers must satisfy shape + placement constraints.
|
||||
--- Walks the priority list once; at most one error is appended per marker so that a single source-level defect does not cascade into multiple findings.
|
||||
-- Check: debug-skip markers must satisfy shape + placement constraints.
|
||||
--- Walks the priority list once; each marker produces at most one error, so one source defect yields one finding.
|
||||
--- Priority order (first defect wins):
|
||||
--- 1. has_parens == false -> requires parentheses: marker()
|
||||
--- 2. args ~= "" -> takes no arguments
|
||||
--- 3. superseded_by_marker_line -> duplicate marker (cite superseding line)
|
||||
--- 4. pending + no target_kind -> dangling (no following declaration)
|
||||
--- 5. unsupported target_kind -> marker precedes an unrelated declaration
|
||||
--- Valid markers before whole-atom / bare-component / proc-component declarations emit no error and remain in src.scan.skip_over.atoms / .components.
|
||||
--- @param marker SkipOverMarker
|
||||
--- 1. marker_kind ~= "atom_dbg_skip" -> legacy/renamed spelling (use `atom_dbg_skip`)
|
||||
--- 2. marker_kind == "atom_dbg_skip" AND has_parens -> parenthesized form (the marker is bare-only)
|
||||
--- 3. args ~= "" -> takes no arguments
|
||||
--- 4. superseded_by_marker_line -> duplicate marker (cite superseding line)
|
||||
--- 5. pending + no target_kind -> dangling (no following declaration)
|
||||
--- 6. unsupported target_kind -> marker precedes an unrelated declaration
|
||||
--- Valid markers stamp `debug_skip` on whole-atom, bare-component, and proc-component declaration records in scan_source.lua.
|
||||
--- @param marker DebugSkipMarker
|
||||
--- @param _pipe_ctx PipeCtx -- unused today; kept for plex-shape consistency with per_annot
|
||||
--- @param findings Findings
|
||||
local function check_skip_marker(marker, _pipe_ctx, findings)
|
||||
local kind = marker.marker_kind
|
||||
local line = marker.marker_line
|
||||
|
||||
if not marker.has_parens then
|
||||
-- Tasks 6+7 left `scan.debug_skip_markers` with production records for `atom_dbg_skip` only; other identifiers take the walker's unrelated branch.
|
||||
|
||||
if marker.has_parens then
|
||||
findings.errors[#findings.errors + 1] = {
|
||||
line = line,
|
||||
msg = string.format("%s marker at line %d requires parentheses: marker()", kind, line),
|
||||
msg = string.format("%s marker at line %d must be bare; the parenthesized form is no longer accepted (use `atom_dbg_skip MipsAtom_(name) { ... }`)",
|
||||
kind, line),
|
||||
}
|
||||
return
|
||||
end
|
||||
@@ -384,11 +379,8 @@ end
|
||||
|
||||
--- Warn when a source references an unregistered alias.
|
||||
---
|
||||
--- R_TapePtr / R_AtomJmp / R_PrimCursor / R_FaceCursor / R_VertBase / R_OtBase are the context aliases opted in via `#define atom_reg` in lottes_tape.h.
|
||||
--- A source referencing an unregistered R_X emits one pass-level info entry
|
||||
--- (emitted only when at least one such rejection lands in this source) tells users where to look.
|
||||
---
|
||||
--- This check directs raw C-ABI register names to explicit alias registration.
|
||||
--- R_TapePtr, R_AtomJmp, R_PrimCursor, R_FaceCursor, R_VertBase, and R_OtBase opt in through `#define atom_reg` in lottes_tape.h.
|
||||
--- When a source uses an unregistered R_X, this check emits one pass-level info entry for that source and directs C-ABI register names to explicit alias registration.
|
||||
--- @param _src SourceFile
|
||||
--- @param pipe_ctx PipeCtx
|
||||
--- @param findings Findings
|
||||
@@ -421,7 +413,7 @@ end
|
||||
-- per_annot(annot, pipe_ctx, findings) -- runs once per AtomAnnotation
|
||||
-- post(pipe_ctx, findings) -- runs once after all per_annot calls complete (full-corpus aggregation)
|
||||
-- per_macro(macro, wc, findings) -- runs once per TAPE_WORDS / _Pragma macro declaration
|
||||
-- per_skip_marker(marker, pipe_ctx, findings) -- runs once per src.scan.skip_over.markers entry
|
||||
-- per_skip_marker(marker, pipe_ctx, findings) -- runs once per src.scan.debug_skip_markers entry
|
||||
--
|
||||
-- Adding a new check = 1 row here + 1 function above. The `validate()` dispatch loop never needs editing.
|
||||
|
||||
@@ -444,14 +436,8 @@ local CHECK_RULES = {
|
||||
--
|
||||
-- Pure check: read from src.scan, run validations, emit findings. The scan was done once upstream.
|
||||
|
||||
--- Build the corpus-wide pipe_ctx ONCE per pass run.
|
||||
--- Reads the merged `corpus.*` registries (canonical cross-source lookups),
|
||||
--- and the corpus-wide `atom_infos` list (preserving source order + duplicates).
|
||||
--- The corpus is the source of truth; per-source scans retain body / declaration
|
||||
--- ownership via `src.scan` and the per-source `atoms` / `atom_infos` projections.
|
||||
---
|
||||
--- Canonical ownership: a context without `ctx.shared.corpus` is rejected with an explicit canonical-corpus message.
|
||||
--- No per-source fallback synthesis is performed; callers MUST construct a canonical ctx through `build_ctx`.
|
||||
--- Builds one pass-wide pipe_ctx from the merged `corpus.*` registries and source-ordered `corpus.atom_infos`; per-source declarations and bodies remain in `src.scan`.
|
||||
--- The module ownership contract above requires callers to construct `ctx.shared.corpus` through `build_ctx`; the error message below enforces that gate.
|
||||
--- @param ctx PassCtx
|
||||
--- @return PipeCtx
|
||||
local function build_corpus_pipe_ctx(ctx)
|
||||
@@ -462,9 +448,7 @@ local function build_corpus_pipe_ctx(ctx)
|
||||
.. "no per-source fallback is supported)", 0)
|
||||
end
|
||||
|
||||
-- Corpus atom_infos preserves source-order + duplicates;
|
||||
-- the per-check `check_unique_annotation` post-rule still flags duplicate annotation
|
||||
-- names within this list. We pre-compute the annot_counts map here so the per_source checks can iterate it without re-walking.
|
||||
-- `corpus.atom_infos` preserves source order and duplicates; I precompute counts here for `check_unique_annotation` and the per-source checks.
|
||||
local annot_counts = {}
|
||||
for _, info in ipairs(corpus.atom_infos or {}) do
|
||||
if info and info.atom_name then
|
||||
@@ -472,10 +456,9 @@ local function build_corpus_pipe_ctx(ctx)
|
||||
end
|
||||
end
|
||||
|
||||
-- The pipe_ctx views REFERENCE the corpus tables directly (no copies).
|
||||
-- Every consumer of these fields observes mutations via the canonical corpus without independently mutable registry construction.
|
||||
return {
|
||||
-- Cross-source lookup tables (canonical corpus projections).
|
||||
-- Cross-source lookup tables from corpus.
|
||||
register_alias_registry = corpus.register_alias_registry or {},
|
||||
type_name_registry = corpus.type_name_registry or {},
|
||||
atom_views = corpus.atom_views or {},
|
||||
@@ -489,8 +472,7 @@ local function build_corpus_pipe_ctx(ctx)
|
||||
annot_counts = annot_counts,
|
||||
-- Corpus-wide collisions (recorded by scan_source.merge_corpus_registries).
|
||||
collisions = corpus.collisions or {},
|
||||
-- wc still consumed by check_macro_word_drift; reads from the canonical
|
||||
-- `corpus.word_counts` table (built by word_count_eval.run).
|
||||
-- `check_macro_word_drift` reads `corpus.word_counts`, populated by word_count_eval.run.
|
||||
word_counts = corpus.word_counts or {},
|
||||
}
|
||||
end
|
||||
@@ -498,7 +480,7 @@ end
|
||||
--- Validate one source against its pre-scanned SourceScan payload + the corpus-wide pipe_ctx.
|
||||
--- @param ctx PassCtx
|
||||
--- @param src SourceFile
|
||||
--- @param corpus_pipe_ctx PipeCtx|nil -- built once per pass from corpus registries; nil = self-build (canonical projection).
|
||||
--- @param corpus_pipe_ctx PipeCtx|nil -- built once per pass from corpus registries; nil builds the same projection here.
|
||||
--- @return AnnotatedResult
|
||||
local function validate(ctx, src, corpus_pipe_ctx)
|
||||
corpus_pipe_ctx = corpus_pipe_ctx or build_corpus_pipe_ctx(ctx)
|
||||
@@ -527,11 +509,7 @@ local function validate(ctx, src, corpus_pipe_ctx)
|
||||
}
|
||||
end
|
||||
|
||||
-- Build the per-source pipe_ctx (Fleury: expose structure).
|
||||
-- Cross-source visibility comes from `corpus_pipe_ctx`;
|
||||
-- per-source declaration / body ownership comes from `src.scan`.
|
||||
-- pipe_ctx.types / pipe_ctx.atom_views / pipe_ctx.seen_defaults / pipe_ctx.type_occurrences
|
||||
-- are projected from the per-source scan so the per_source check rules can iterate the source-local occurrences.
|
||||
-- Build a per-source pipe_ctx: shared lookups come from `corpus_pipe_ctx`, while declarations, bodies, types, views, defaults, and occurrences come from `src.scan`.
|
||||
local seen_defaults = {}
|
||||
for reg, _ in pairs(scan.types or {}) do
|
||||
seen_defaults[reg] = (seen_defaults[reg] or 0) + 1
|
||||
@@ -551,8 +529,7 @@ local function validate(ctx, src, corpus_pipe_ctx)
|
||||
seen_defaults = seen_defaults,
|
||||
atom_infos_list = atom_infos_list,
|
||||
binds_list = scan.binds or {},
|
||||
-- Source-derived registries: still populated from the scan payload as a convenience for callers that want source-local visibility.
|
||||
-- The canonical cross-source lookup tables live in corpus_pipe_ctx.
|
||||
-- See the module ownership contract; these shared lookup tables come from corpus_pipe_ctx.
|
||||
register_alias_registry = corpus_pipe_ctx.register_alias_registry,
|
||||
type_name_registry = corpus_pipe_ctx.type_name_registry,
|
||||
}
|
||||
@@ -563,9 +540,7 @@ local function validate(ctx, src, corpus_pipe_ctx)
|
||||
-- Each check writes to the list appropriate for its severity.
|
||||
local findings = { errors = {}, warnings = {}, info = {} }
|
||||
|
||||
-- Propagate parse-time errors from scan_source's atom_info parsing.
|
||||
-- These are errors found in the atom_info(...) body itself (e.g., malformed args).
|
||||
-- They are pre-existing in the scan payload — we just lift them into our findings list.
|
||||
-- Lift parse-time errors already recorded in scan_source's atom_info payload into this pass's findings list.
|
||||
for _, a in ipairs(annots) do
|
||||
if a.errors then
|
||||
for _, msg in ipairs(a.errors) do
|
||||
@@ -589,11 +564,9 @@ local function validate(ctx, src, corpus_pipe_ctx)
|
||||
if rule.post then rule.post(pipe_ctx, findings) end
|
||||
end
|
||||
|
||||
-- Per-skip-marker rules.
|
||||
-- Each raw marker recorded by scan_source (in scan.skip_over.markers) is validated independently;
|
||||
-- the check emits at most one error per marker.
|
||||
-- Valid markers stay attached to scan.skip_over.atoms /.components for dwarf_injection.lua consumer.
|
||||
local skip_markers = scan.skip_over and scan.skip_over.markers or {}
|
||||
-- scan_source records each marker in scan.debug_skip_markers; this loop validates each record independently and emits at most one error per marker.
|
||||
-- Valid markers stamp `debug_skip = true` on the following atom or component declaration, which downstream consumers read directly.
|
||||
local skip_markers = scan.debug_skip_markers or {}
|
||||
for _, marker in ipairs(skip_markers) do
|
||||
for _, rule in ipairs(CHECK_RULES) do
|
||||
if rule.per_skip_marker then rule.per_skip_marker(marker, pipe_ctx, findings) end
|
||||
@@ -684,15 +657,12 @@ function M.run(ctx)
|
||||
local errors = {}
|
||||
local warnings = {}
|
||||
|
||||
-- Build the corpus-wide pipe_ctx ONCE per pass run.
|
||||
-- Build the shared pipe_ctx once for this run; every validate() call sees the same cross-source registries.
|
||||
-- The corpus owns the canonical cross-source registries; per-source scans retain body / declaration ownership.
|
||||
-- The pipe_ctx is shared across every validate() invocation in this M.run so cross-source visibility is constant.
|
||||
local corpus_pipe_ctx = build_corpus_pipe_ctx(ctx)
|
||||
local corpus = ctx.shared.corpus
|
||||
|
||||
-- Per-DIRECTORY (per-module) aggregation.
|
||||
-- Group sources by `src.dir`, validate every source in the dir, then emit ONE errors.h per dir.
|
||||
-- The corpus owns `sources_by_dir`; this pass reads the corpus bucket directly.
|
||||
-- Group `corpus.sources_by_dir` by module, validate every source in each bucket, and emit one errors.h per directory.
|
||||
local by_dir = (corpus and corpus.sources_by_dir) or {}
|
||||
|
||||
for dir, dir_sources in pairs(by_dir) do
|
||||
|
||||
@@ -1,23 +1,23 @@
|
||||
--- passes/atoms_source_map.lua — Per-.word source-line map emitter for tape atoms.
|
||||
---
|
||||
--- Reads the canonical `atom.paths` projection produced by the upstream `emission_model` pass.
|
||||
--- The ordered `items` stream, dense `word_events`, and `invocations` views are the only semantic inputs to this pass;
|
||||
--- it emits one `WORD N LINE L TEXT T` line per emitted `.word`.
|
||||
--- Writer: this pass, given `atom.paths` (the per-atom mutable surface owned by `emission_model`). Readers:
|
||||
--- `passes/dwarf_injection.lua` (synthesizes DW_TAG_inlined_subroutine + per-word line program rows) and the gdb-runtime
|
||||
--- wrapper at `scripts/gdb/gdb_tape_atoms.gdb` (loads the source map via `source <path>`).
|
||||
---
|
||||
--- **Two output forms** (per the workspace's per-emission-form pattern from
|
||||
--- `guide_metaprogram_ssdl.md`):
|
||||
--- 1. **Canonical text form** — `<out_root>/<basename>.atoms.sourcemap.txt`.
|
||||
--- Format-version-tagged for forward-compat.
|
||||
--- Lives in `<out_root>/` (build/gen).
|
||||
--- Matches the convention used by `annotation.lua` (`<out_root>/<basename>.errors.h`) + `static_analysis.lua` (`<out_root>/<basename>.static_analysis.txt`).
|
||||
--- Inputs from `atom.paths`: the ordered `items` stream, dense `word_events`, `invocations` views. Outputs: one
|
||||
--- `WORD N LINE L TEXT T` line per emitted `.word`, plus the per-word provenance form that DWARF synthesis consumes.
|
||||
---
|
||||
--- **Two output forms** (per the workspace's per-emission-form pattern from `guide_metaprogram_ssdl.md`):
|
||||
--- 1. **Sourcemap.txt form** — `<out_root>/<basename>.atoms.sourcemap.txt`. Format-version-tagged for forward-compat.
|
||||
--- Lives in `<out_root>/` (build/gen). Mirrors the convention used by `annotation.lua`
|
||||
--- (`<out_root>/<basename>.errors.h`) and `static_analysis.lua` (`<out_root>/<basename>.static_analysis.txt`).
|
||||
--- Compile artifacts (`*.macs.h`, `*.offsets.h`) stay in `<source_dir>/gen/`.
|
||||
--- 2. **gdb-runtime form** — `<ctx.out_root>/gdb_tape_atoms_runtime.gdb`
|
||||
--- (pure gdb command script; addresses pre-computed via `nm`; the 9 user commands defined as `define ... end` blocks).
|
||||
--- Emitted ONLY when `ctx.flags.gdb_runtime` is true AND `ctx.flags.elf_path` points to an existing ELF.
|
||||
--- The gdb runtime form lets `gdb-multiarch --without-python` users (the common case on Windows MinGW builds)
|
||||
--- load the source-map data via `source <path>` — no Python/Tcl/Guile required.
|
||||
--- 2. **gdb-runtime form** — `<ctx.out_root>/gdb_tape_atoms_runtime.gdb`. A pure gdb command script — addresses come
|
||||
--- from `nm`, the 9 user commands are static `define ... end` blocks. Emitted when `ctx.flags.gdb_runtime` is true
|
||||
--- AND `ctx.flags.elf_path` points to an existing ELF. Useful for `gdb-multiarch --without-python` users
|
||||
--- (the common case on Windows MinGW builds) — `source <path>` loads it with no Python / Tcl / Guile required.
|
||||
---
|
||||
--- **Output format** (canonical text form):
|
||||
--- **Output format** (sourcemap.txt form):
|
||||
--- ```
|
||||
--- # FORMAT_VERSION 1
|
||||
--- # auto-generated by ps1_meta.lua (passes/atoms_source_map.lua) — DO NOT EDIT
|
||||
@@ -31,11 +31,9 @@
|
||||
--- ENDATOM
|
||||
--- ```
|
||||
---
|
||||
--- Marker records are zero-width in `atom.paths.items`; they do not appear in
|
||||
--- the dense word view and therefore emit no WORD rows.
|
||||
--- Marker records are zero-width in `atom.paths.items`, so they emit no WORD rows in the dense word view.
|
||||
---
|
||||
--- **Conventions:** tabs (1/level), EmmyLua annotations, no regex,
|
||||
--- Lua 5.3 compatible.
|
||||
--- **Conventions:** tabs (1/level), EmmyLua annotations, no regex, Lua 5.3 compatible.
|
||||
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
-- Module-scope requires + package.path setup
|
||||
@@ -62,17 +60,16 @@ local FORMAT_VERSION = 1
|
||||
|
||||
--- @class AtomSourceMapCtx
|
||||
--- @field shared table -- `ctx.shared`
|
||||
--- @field shared.corpus table -- canonical source-order corpus
|
||||
--- @field shared.corpus table -- source-order registry; single writer is build_ctx
|
||||
--- @field shared.word_counts table
|
||||
--- @field out_root string -- output root (e.g. "build/gen")
|
||||
--- @field flags table -- `ctx.flags`; reads `flags.gdb_runtime` + `flags.elf_path`
|
||||
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
-- Canonical atom-path renderers
|
||||
-- Atom-path renderers
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
--- Join canonical words to canonical word items. `items` supplies the ordered
|
||||
--- word boundaries, while `word_events` supplies call text and source lines.
|
||||
--- Join word boundaries (from `items`) to per-word call text + source lines (from `word_events`).
|
||||
--- @param atom table
|
||||
--- @return table[], integer
|
||||
local function canonical_word_entries(atom)
|
||||
@@ -99,11 +96,11 @@ local function canonical_word_entries(atom)
|
||||
return entries, #events
|
||||
end
|
||||
|
||||
--- Render one atom's provenance stanza. Format 1 remains:
|
||||
--- `WORD N CALL <src-path>:<src-line> MACRO <name> "<def-path>:<def-line>" BODY <line>`
|
||||
--- `WORD N CALL <src-path>:<src-line> RAW`
|
||||
--- Component identity comes from the canonical outermost invocation record;
|
||||
--- the count-table lookup is the canonical component declaration witness.
|
||||
--- Render one atom's provenance stanza. Format 1 line shapes:
|
||||
--- `WORD N CALL <src-path>:<src-line> MACRO <name> "<def-path>:<def-line>" BODY <line>` (component invocation)
|
||||
--- `WORD N CALL <src-path>:<src-line> RAW` (raw `.word` outside any mac_* component)
|
||||
--- Component identity comes from the outermost invocation record; the count-table lookup confirms the component was
|
||||
--- declared in `corpus.word_counts` (populated by word_count_eval + components passes).
|
||||
--- @param src table
|
||||
--- @param atom table
|
||||
--- @param wc table -- identity alias of corpus.word_counts
|
||||
@@ -162,7 +159,7 @@ local function render_provenance(src, wc)
|
||||
return table.concat(lines, "\n") .. "\n"
|
||||
end
|
||||
|
||||
--- Render one atom's stanza for the canonical text form (ATOM header line, N WORD lines, ENDATOM marker).
|
||||
--- Render one atom's stanza for the sourcemap.txt form (ATOM header line, N WORD lines, ENDATOM marker).
|
||||
--- Returns (lines, total_words).
|
||||
--- @param src table
|
||||
--- @param atom table
|
||||
@@ -183,8 +180,8 @@ local function emit_atom_stanza(src, atom)
|
||||
return lines, total
|
||||
end
|
||||
|
||||
--- Render the full source map file content for one source (one .atoms.sourcemap.txt per source).
|
||||
--- Mirrors offsets.lua's `project_atoms` shape: scan.atoms + scan.raw_atoms, no kind filter.
|
||||
--- Render the full source map file content for one source (one .atoms.sourcemap.txt per source). Mirrors offsets.lua's
|
||||
--- `project_atoms` shape: scan.atoms + scan.raw_atoms, no kind filter.
|
||||
--- @param src table
|
||||
--- @param wc table
|
||||
--- @return string
|
||||
@@ -219,8 +216,7 @@ local function gdb_escape(s)
|
||||
return (s:gsub("\\", "\\\\"):gsub('"', '\\"'))
|
||||
end
|
||||
|
||||
--- Build the list of atoms with addresses + word entries.
|
||||
--- Shared helper for the gdb-runtime file emission.
|
||||
--- Build the list of atoms with addresses + word entries. Shared helper for the gdb-runtime file emission.
|
||||
--- @param ctx PassCtx
|
||||
--- @return table[] -- list of {idx, name, src_path, file_base, addr, size_bytes, words, entries}
|
||||
local function build_atom_table(ctx)
|
||||
@@ -256,12 +252,12 @@ local function build_atom_table(ctx)
|
||||
return matched
|
||||
end
|
||||
|
||||
--- Append the 9 gdb command definitions to `lines`. Pure gdb scripting no Python, no Tcl, no Guile required.
|
||||
--- **Fully hardcoded per-atom** because gdb doesn't do nested `$` substitution in var names
|
||||
--- `$__atom_name_$__i` inside a `while` loop is treated as one literal identifier, not a concat.
|
||||
--- Append the 9 gdb command definitions to `lines`. Pure gdb scripting — addresses come from `nm`, the convenience
|
||||
--- vars set in `emit_gdb_runtime` provide printf args, and each command is a static sequence of `printf` / `tbreak` /
|
||||
--- `if ... end` blocks. The Lua pass emits N atoms' worth of lines; runtime iteration is gdb's job.
|
||||
---
|
||||
--- Each command is a static sequence of `printf` / `tbreak` / `if ... end` blocks.
|
||||
--- The Lua pass emits N atoms' worth of lines — no runtime iteration.
|
||||
--- Why hardcoded per-atom: gdb's `$` substitution doesn't concat inside var names — `$__atom_name_$__i` in a `while`
|
||||
--- loop resolves to one literal identifier, not `name_i`. Compile-time emission is the only path.
|
||||
--- @param lines table -- output line buffer (mutated in place)
|
||||
--- @param matched table -- list of atom records from `build_atom_table`
|
||||
local function append_gdb_commands(lines, matched)
|
||||
@@ -402,9 +398,9 @@ local function append_gdb_commands(lines, matched)
|
||||
lines[#lines + 1] = "end"
|
||||
end
|
||||
|
||||
--- Emit the gdb-runtime file (post-link). Pure gdb scripting — no Python.
|
||||
--- Reads ELF addresses via `mipsel-none-elf-nm -S`, embeds them in `<ctx.out_root>/gdb_tape_atoms_runtime.gdb`
|
||||
--- so gdb loads the data via `set $var = ...` + `define ... end` blocks at source-time.
|
||||
--- Emit the gdb-runtime file (post-link). Pure gdb scripting — addresses come from `mipsel-none-elf-nm -S`, get embedded
|
||||
--- in `<ctx.out_root>/gdb_tape_atoms_runtime.gdb`, and load via `set $var = ...` + `define ... end` blocks at gdb
|
||||
--- source-time.
|
||||
--- @param ctx PassCtx
|
||||
local function emit_gdb_runtime(ctx)
|
||||
if not (ctx.flags and ctx.flags.gdb_runtime) then return end
|
||||
@@ -465,8 +461,7 @@ local function emit_gdb_runtime(ctx)
|
||||
local out_path = ctx.out_root .. "/gdb_tape_atoms_runtime.gdb"
|
||||
duffle.ensure_dir(duffle.dirname(out_path))
|
||||
duffle.write_file_lf(out_path, table.concat(lines, "\n") .. "\n")
|
||||
io.stderr:write(string.format(
|
||||
"[atoms_source_map] wrote %s (%d atoms)\n", out_path, #matched))
|
||||
-- io.stderr:write(string.format("[atoms_source_map] wrote %s (%d atoms)\n", out_path, #matched))
|
||||
end
|
||||
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
@@ -475,10 +470,10 @@ end
|
||||
|
||||
local M = {}
|
||||
|
||||
--- Pass entry: emit one `<out_root>/<basename>.atoms.sourcemap.txt` per source file that contains at least one `MipsAtom_(name)` / `MipsCode code_<name>` declaration.
|
||||
--- Also emits `<out_root>/<basename>.atoms.provenance.txt`:
|
||||
--- per-.word provenance with `mac_X(...)` component resolution back to the component's definition file:line + the per-word body line.
|
||||
--- Optionally also emit `<ctx.out_root>/gdb_tape_atoms_runtime.gdb` when `ctx.flags.gdb_runtime` is true.
|
||||
--- Pass entry. For each source that declares at least one `MipsAtom_(name)` / `MipsCode code_<name>`, emit two files
|
||||
--- in `<out_root>/`: `<basename>.atoms.sourcemap.txt` (per-word call-site map) and `<basename>.atoms.provenance.txt`
|
||||
--- (per-word definition + body line, resolved via the outermost `mac_X(...)` invocation). When `ctx.flags.gdb_runtime`
|
||||
--- is true and `ctx.flags.elf_path` exists, also emit the post-link gdb script `<ctx.out_root>/gdb_tape_atoms_runtime.gdb`.
|
||||
--- @param ctx PassCtx
|
||||
--- @return PassResult
|
||||
function M.run(ctx)
|
||||
@@ -491,8 +486,7 @@ function M.run(ctx)
|
||||
error("atoms_source_map.run requires ctx.shared.corpus.source_order (canonical corpus).", 0)
|
||||
end
|
||||
|
||||
-- Word counts are owned by `corpus.word_counts`.
|
||||
-- The canonical owner is `corpus.word_counts` (populated by `passes/word_count_eval.lua` + `passes/components.lua`).
|
||||
-- Word counts come from `corpus.word_counts` (populated by word_count_eval + components passes).
|
||||
local wc = corpus.word_counts or {}
|
||||
if not next(wc) then
|
||||
warnings[#warnings + 1] = {
|
||||
@@ -501,7 +495,7 @@ function M.run(ctx)
|
||||
}
|
||||
end
|
||||
|
||||
-- Always emit the canonical text form (per-source).
|
||||
-- Always emit the text form (per-source).
|
||||
for _, src in ipairs(corpus.source_order) do
|
||||
local has_projection = false
|
||||
for _, atom in ipairs((src.scan or {}).atoms or {}) do
|
||||
|
||||
+66
-104
@@ -1,12 +1,12 @@
|
||||
--- passes/components.lua — Component-macro header generator.
|
||||
---
|
||||
--- Reads the pre-scanned SourceScan payload (produced once upstream by `duffle.scan_source`)
|
||||
--- for `MipsAtomComp_(ac_X)` and `MipsAtomComp_Proc_(ac_X, { body })` declarations, then does per-source backward lookups
|
||||
--- for the function-args string (from the preceding `FI_ MipsAtom ac_X(...)` function declaration)
|
||||
--- and the preceding comment block (for LSP/IntelliSense signature docs).
|
||||
--- Ownership: `corpus.word_counts`, `corpus.components`, and `corpus.component_body_index`.
|
||||
--- Scanner owns `declaration_comment` and `debug_skip` on each declaration record; this pass projects both forward.
|
||||
---
|
||||
--- Emits a per-directory `<dir_basename>.macs.h` containing one `#define mac_X(sig) \` macro per component + `WORD_COUNT(mac_X, N)`
|
||||
--- entries for downstream offset computation.
|
||||
--- Reads the pre-scanned SourceScan payload from `duffle.scan_source` for `MipsAtomComp_(ac_X)` and `MipsAtomComp_Proc_(ac_X, { body })` declarations,
|
||||
--- then resolves the function-args string from the preceding `FI_ MipsAtom ac_X(...)` declaration via a backward walk.
|
||||
---
|
||||
--- Emits one `<dir_basename>.macs.h` per source with `#define mac_X(sig) \` macros plus `WORD_COUNT(mac_X, N)` entries for downstream offset computation.
|
||||
---
|
||||
--- **Conventions**: tabs (1/level), EmmyLua annotations, no regex,
|
||||
--- Lua 5.3 compatible.
|
||||
@@ -59,7 +59,6 @@ local GEN_SUBDIR = "gen"
|
||||
--- @field sources SourceFile[] -- all source files in the build
|
||||
--- @field metadata_path string -- path to word_count.metadata.h
|
||||
--- @field shared table -- cross-pass shared state
|
||||
--- @field shared.word_counts table<string, integer> -- populated by word-counts + components
|
||||
--- @field out_root string -- output root (e.g. "build/gen")
|
||||
--- @field project_root string -- project root (e.g. "code/")
|
||||
--- @field upstream table<string, table> -- per-pass upstream outputs
|
||||
@@ -72,11 +71,13 @@ local GEN_SUBDIR = "gen"
|
||||
--- @field warnings table[] -- {line=, msg=} entries; build-succeeds
|
||||
|
||||
--- @class Component
|
||||
--- @field name string -- atom name (without `ac_` prefix)
|
||||
--- @field body string -- brace-delimited body (without the braces)
|
||||
--- @field args string|nil -- function-args string (function form only)
|
||||
--- @field line integer -- source line of the declaration
|
||||
--- @field comment string|nil -- preceding `/* */` or `//` comment block (signature doc)
|
||||
--- @field name string -- atom name (without `ac_` prefix)
|
||||
--- @field body string -- brace-delimited body (without the braces)
|
||||
--- @field args string|nil -- function-args string (function form only)
|
||||
--- @field line integer -- source line of the declaration
|
||||
--- @field comment string|nil -- scanner-owned `declaration_comment`; the components pass reads it from the scanner record
|
||||
--- @field kind string -- "comp_bare" | "comp_proc"
|
||||
--- @field debug_skip boolean -- mirror of `a.debug_skip` (scanner-owned); true iff a bare `atom_dbg_skip` marker immediately preceded the declaration
|
||||
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
-- Local helpers (file I/O + path normalization)
|
||||
@@ -85,7 +86,11 @@ local GEN_SUBDIR = "gen"
|
||||
local M = {}
|
||||
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
-- Back-walk helpers (composed into the 2 entry points below: find_function_args_for + preceding_comment_block)
|
||||
-- Back-walk helpers (composed into the entry point below: find_function_args_for)
|
||||
--
|
||||
-- Only the function-args lookup for proc components occurs here.
|
||||
-- The preceding-comment walk occur in `scan_source.lua` — `a.declaration_comment` carries the resolved comment,
|
||||
-- so this file reads it forward rather than re-walking the source.
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
--- Find the args of the function declaration that immediately precedes a `MipsAtomComp_Proc_` invocation of the given name.
|
||||
@@ -132,69 +137,6 @@ local function find_function_args_for(source, name, before_pos)
|
||||
return inner
|
||||
end
|
||||
|
||||
--- Find the contiguous comment block immediately preceding `pos` in `source`.
|
||||
--- Returns the comment text (with the `/* */` or `//` markers preserved) or an empty string if no comment is adjacent.
|
||||
---
|
||||
--- Used to copy signature comments from the source declaration (`MipsAtomComp_` / `MipsAtomComp_Proc_` / function decl)
|
||||
--- over to the generated `mac_X` macro, so LSP/IntelliSense displays the args doc.
|
||||
--- @param source string
|
||||
--- @param pos integer
|
||||
--- @return string
|
||||
local function preceding_comment_block(source, pos)
|
||||
local scan_pos = pos
|
||||
local pieces = {}
|
||||
while true do
|
||||
-- skip whitespace backward; land on the next non-ws character.
|
||||
local non_ws = scan_pos - 1
|
||||
while non_ws > 0 do
|
||||
local ch = source:sub(non_ws, non_ws)
|
||||
if ch == " " or ch == "\t" or ch == "\n" or ch == "\r" then
|
||||
non_ws = non_ws - 1
|
||||
else
|
||||
break
|
||||
end
|
||||
end
|
||||
if non_ws == 0 then break end
|
||||
if non_ws >= 2 and source:sub(non_ws - 1, non_ws) == "*/" then
|
||||
-- block comment close: find the opening /* by walking back over /* candidates
|
||||
-- in source[1..non_ws-1].
|
||||
local prefix = source:sub(1, non_ws - 1)
|
||||
local open_at = nil
|
||||
for scan = #prefix - 1, 1, -1 do
|
||||
if prefix:sub(scan, scan + 1) == "/*" then
|
||||
open_at = scan
|
||||
break
|
||||
end
|
||||
end
|
||||
if not open_at then break end
|
||||
-- include the indentation before the /* by walking back over leading spaces + tabs.
|
||||
local block_start = open_at
|
||||
while block_start > 1 do
|
||||
local ch = source:sub(block_start - 1, block_start - 1)
|
||||
if ch ~= " " and ch ~= "\t" then break end
|
||||
block_start = block_start - 1
|
||||
end
|
||||
table.insert(pieces, 1, source:sub(block_start, non_ws))
|
||||
scan_pos = block_start
|
||||
else
|
||||
-- line comment path: must end in newline, must start with //.
|
||||
local ch = source:sub(non_ws, non_ws)
|
||||
if ch ~= "\n" and ch ~= "\r" then break end
|
||||
-- walk back from non_ws to the start of the source line (most recent \n or position 1).
|
||||
local line_start = non_ws
|
||||
while line_start > 1 and source:sub(line_start - 1, line_start - 1) ~= "\n" do
|
||||
line_start = line_start - 1
|
||||
end
|
||||
local line = source:sub(line_start, non_ws)
|
||||
if line:sub(1, 2) ~= "//" then break end
|
||||
table.insert(pieces, 1, line)
|
||||
scan_pos = line_start - 1
|
||||
end
|
||||
end
|
||||
if #pieces == 0 then return "" end
|
||||
return table.concat(pieces, "\n")
|
||||
end
|
||||
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
-- Argument-name extraction
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
@@ -246,8 +188,12 @@ end
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
--- Project pre-scanned MipsAtomComp_ / MipsAtomComp_Proc_ entries into Component shape.
|
||||
--- Does per-source backward lookups for args (preceding function decl) and comment (preceding comment block).
|
||||
--- Reads the scanner-owned `declaration_comment` (resolved by scan_source.lua, skipping backward across an associated bare `atom_dbg_skip` marker when present).
|
||||
--- Per-source backward lookups remain in place only for the function `args` of proc components.
|
||||
--- That lookup is unique to components.lua and stays separate from the declaration-comment walk.
|
||||
--- Carries `body_tokens` forward from scan-source so word_count_rec reads from the precomputed table instead of calling duffle.tokenize_body again.
|
||||
--- Carries the scanner-owned `debug_skip` flag forward so the generated projection can emit `/* atom_dbg_skip */`
|
||||
--- before the authored comment and so `update_canonical_components` can mirror the same field onto `corpus.components[name]`.
|
||||
--- @param source string -- the full source text (needed for backward lookups)
|
||||
--- @param scan table -- SourceScan from duffle.scan_source
|
||||
--- @return Component[]
|
||||
@@ -255,8 +201,10 @@ local function project_components(source, scan)
|
||||
local out = {}
|
||||
for _, a in ipairs(scan.atoms) do
|
||||
if a.kind == "comp_bare" or a.kind == "comp_proc" then
|
||||
local args = find_function_args_for(source, a.raw_name, a.ident_pos)
|
||||
local comment = preceding_comment_block(source, a.ident_pos)
|
||||
local args = find_function_args_for(source, a.raw_name, a.ident_pos)
|
||||
-- Comment ownership: scan_source.lua stamps `declaration_comment` on the record by walking backward past any associated bare marker.
|
||||
-- The pass reads `declaration_comment` directly.
|
||||
local comment = a.declaration_comment or ""
|
||||
out[#out + 1] = {
|
||||
line = a.line,
|
||||
name = a.name,
|
||||
@@ -266,6 +214,7 @@ local function project_components(source, scan)
|
||||
args = args,
|
||||
comment = comment,
|
||||
kind = a.kind, -- "comp_bare" | "comp_proc"; provenance emitter reads this.
|
||||
debug_skip = a.debug_skip == true,
|
||||
}
|
||||
end
|
||||
end
|
||||
@@ -450,6 +399,9 @@ end
|
||||
|
||||
--- Build the list of lines for one component
|
||||
--- (signature comment, `#define mac_X(...)` line with backslash-continued tokens, then `WORD_COUNT(mac_X, N)` entry).
|
||||
--- For skipped components, a `/* atom_dbg_skip */` marker comment is emitted immediately before the authored comment block.
|
||||
--- The marker is a single line, the comment comes next, and the `#define` line follows. The `debug_skip` stamp is scanner-owned
|
||||
--- (`a.debug_skip == true` on the declaration record); the components pass projects it directly.
|
||||
--- @param c Component
|
||||
--- @param components Component[]
|
||||
--- @param wc table<string, integer>
|
||||
@@ -457,6 +409,13 @@ end
|
||||
local function build_component_lines(c, counts)
|
||||
local lines = {}
|
||||
|
||||
-- Marker comment: emitted once for every skipped component.
|
||||
-- The marker is scanner-owned (declared by `atom_dbg_skip` immediately before the declaration in the source);
|
||||
-- the components pass projects `c.debug_skip` and emits the marker as a generated comment.
|
||||
if c.debug_skip then
|
||||
lines[#lines + 1] = "/* atom_dbg_skip */"
|
||||
end
|
||||
|
||||
if c.comment and c.comment ~= "" then
|
||||
for _, line in ipairs(split_comment_lines(c.comment)) do
|
||||
lines[#lines + 1] = line
|
||||
@@ -551,9 +510,9 @@ end
|
||||
-- Pass entry
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
--- (internal) Extend the canonical `corpus.word_counts` with this source's component macros so offsets sees them without re-reading the file.
|
||||
--- (internal) Extend `corpus.word_counts` with this source's component macros so offsets sees them without re-reading the file.
|
||||
--- First declaration wins: a later caller's count is dropped (the existing entry from the first source is preserved).
|
||||
--- @param corpus table -- the canonical corpus
|
||||
--- @param corpus table -- the corpus
|
||||
--- @param components Component[]
|
||||
--- @param counts table<string, integer> -- precomputed word counts (from count_all_components)
|
||||
local function update_canonical_word_counts(corpus, components, counts)
|
||||
@@ -567,33 +526,37 @@ local function update_canonical_word_counts(corpus, components, counts)
|
||||
end
|
||||
|
||||
--- @class ComponentDef
|
||||
--- @field name string -- bare name (without ac_/mac_ prefix)
|
||||
--- @field line integer -- definition source line (line of `MipsAtomComp_(ac_X)` / `MipsAtomComp_Proc_(ac_X, ...)`)
|
||||
--- @field path string -- absolute source path of the definition
|
||||
--- @field kind string -- "comp_bare" | "comp_proc"
|
||||
--- @field name string -- bare name (without ac_/mac_ prefix)
|
||||
--- @field line integer -- definition source line (line of `MipsAtomComp_(ac_X)` / `MipsAtomComp_Proc_(ac_X, ...)`)
|
||||
--- @field path string -- absolute source path of the definition
|
||||
--- @field kind string -- "comp_bare" | "comp_proc"
|
||||
--- @field debug_skip boolean -- mirror of the scanner-owned `a.debug_skip`; consumers read this directly
|
||||
|
||||
--- (internal) Populate the canonical `corpus.components` projection with this source's components-by-name map.
|
||||
--- (internal) Populate `corpus.components` with this source's components-by-name map.
|
||||
--- First declaration wins; later declarations of the same bare name are dropped and recorded as a collision via `corpus.collisions` (kind = "component").
|
||||
--- The pass does NOT write to `ctx.shared.components` (ownership follows the canonical contract).
|
||||
--- @param corpus table -- the canonical corpus
|
||||
--- The `debug_skip` field mirrors the scanner-owned declaration record (`c.debug_skip`).
|
||||
--- No parallel skip map is built here; consumers that need the per-component skip state read `corpus.components[name].debug_skip` directly.
|
||||
--- @param corpus table -- the corpus
|
||||
--- @param src SourceFile
|
||||
--- @param components Component[]
|
||||
local function update_canonical_components(corpus, src, components)
|
||||
local rel_path = src.path:gsub("\\", "/")
|
||||
for _, c in ipairs(components) do
|
||||
-- Keyed by bare name (e.g. `yield`, `load_tri_indices`).
|
||||
-- The atoms_source_map pass looks up components by bare name from the canonical corpus;
|
||||
-- The atoms_source_map pass looks up components by bare name from the corpus;
|
||||
-- `mac_` prefix lives at the call-site identifier and is stripped before lookup.
|
||||
if corpus.components[c.name] == nil then
|
||||
corpus.components[c.name] = {
|
||||
name = c.name,
|
||||
line = c.line,
|
||||
path = rel_path,
|
||||
kind = c.kind or "comp_bare",
|
||||
name = c.name,
|
||||
line = c.line,
|
||||
path = rel_path,
|
||||
kind = c.kind or "comp_bare",
|
||||
debug_skip = c.debug_skip == true,
|
||||
}
|
||||
else
|
||||
-- A second declaration of the same bare name: record a typed collision so static-analysis + the report can surface it.
|
||||
-- Identical-shape declarations (same path + line) do NOT record a collision (the first-wins entry already covers the case).
|
||||
-- Identical-shape declarations (same path + line) reuse the first-wins entry without a collision record.
|
||||
local existing = corpus.components[c.name]
|
||||
if existing.path ~= rel_path or existing.line ~= c.line then
|
||||
local kind = c.kind or "comp_bare"
|
||||
@@ -611,10 +574,10 @@ local function update_canonical_components(corpus, src, components)
|
||||
end
|
||||
end
|
||||
|
||||
--- (internal) Populate the canonical `corpus.component_body_index` projection with this source's body index entries.
|
||||
--- (internal) Populate `corpus.component_body_index` with this source's body index entries.
|
||||
--- First declaration wins; later declarations are dropped (no separate collision record: the components collision is already surfaced by `update_canonical_components`).
|
||||
--- The pass does NOT write to `ctx.shared.component_body_index` (the corpus owns this projection).
|
||||
--- @param corpus table -- the canonical corpus
|
||||
--- The pass writes to `corpus.component_body_index` only (the corpus owns this projection).
|
||||
--- @param corpus table -- the corpus
|
||||
--- @param src SourceFile
|
||||
--- @param components Component[]
|
||||
--- @param scan table -- the SourceScan payload (for line_of)
|
||||
@@ -641,13 +604,13 @@ function M.run(ctx)
|
||||
local errors = {}
|
||||
local warnings = {}
|
||||
|
||||
-- Canonical-corpus ownership gate.
|
||||
-- Corpus ownership gate.
|
||||
local corpus = ctx.shared and ctx.shared.corpus
|
||||
if type(corpus) ~= "table" then
|
||||
error("components.run requires ctx.shared.corpus (canonical corpus).", 0)
|
||||
error("components.run requires ctx.shared.corpus.", 0)
|
||||
end
|
||||
if type(corpus.source_order) ~= "table" then
|
||||
error("components.run requires ctx.shared.corpus.source_order (canonical corpus).", 0)
|
||||
error("components.run requires ctx.shared.corpus.source_order.", 0)
|
||||
end
|
||||
if type(corpus.word_counts) ~= "table" then
|
||||
error("components.run requires ctx.shared.corpus.word_counts; "
|
||||
@@ -655,25 +618,24 @@ function M.run(ctx)
|
||||
.. "(see PASSES deps).", 0)
|
||||
end
|
||||
|
||||
-- Canonical projection ownership:
|
||||
-- Projection ownership:
|
||||
-- * `corpus.word_counts["mac_"..name]` — current component count
|
||||
-- * `corpus.components[name]` — bare-name component definition
|
||||
-- * `corpus.component_body_index[name]` — body / line_of / source index
|
||||
-- The pass does NOT mutate `ctx.shared.components` or `ctx.shared.component_body_index`
|
||||
-- (ownership follows the canonical corpus; consumers read from the corpus directly).
|
||||
-- The pass writes to the corpus only; consumers read from the corpus directly.
|
||||
|
||||
for _, src in ipairs(corpus.source_order) do
|
||||
-- project_components reads from src.scan + does backward lookups on src.text
|
||||
local components = project_components(src.text, src.scan)
|
||||
if #components > 0 then
|
||||
-- Compute all component word counts once per source.
|
||||
-- Use `corpus.word_counts` (the canonical count table) so the recursive lookup sees both authored-metadata entries
|
||||
-- Use `corpus.word_counts` so the recursive lookup sees both authored-metadata entries
|
||||
-- (loaded by word_count_eval.run) AND same-source component entries (populated earlier in this loop by `update_canonical_word_counts`).
|
||||
local counts = count_all_components(components, corpus.word_counts)
|
||||
local macs_path = emit_component_macros_h(ctx, src, components, counts)
|
||||
if macs_path then
|
||||
outputs[#outputs + 1] = { macs_h = macs_path }
|
||||
-- Populate the canonical projections AFTER disk emission (so the byte-identical `.macs.h` contract is preserved before any current-count mutation).
|
||||
-- Populate the projections AFTER disk emission (so the byte-identical `.macs.h` contract is preserved before any current-count mutation).
|
||||
update_canonical_word_counts(corpus, components, counts)
|
||||
update_canonical_components(corpus, src, components)
|
||||
update_canonical_component_body_index(corpus, src, components, src.scan)
|
||||
|
||||
+314
-443
File diff suppressed because it is too large
Load Diff
@@ -1,30 +1,30 @@
|
||||
--- passes/emission_model.lua: Per-atom emission projection.
|
||||
---
|
||||
--- The `emission-model` pass owns `atom.paths` (the canonical per-atom mutable surface)
|
||||
--- for every atom-with-body and every raw atom-with-body declared in `ctx.shared.corpus.source_order`.
|
||||
--- For each such atom, the pass invokes `duffle.project_emission(body_text, component_index, word_counts)`
|
||||
--- and stores the ordered `items` stream plus the dense `word_events` / `markers` / `invocations` views on `atom.paths`.
|
||||
--- The `emission-model` pass owns `atom.paths`, the canonical per-atom mutable surface for atoms and raw atoms with bodies in `ctx.shared.corpus.source_order`.
|
||||
--- For each atom, the pass invokes `duffle.project_emission(body_text, component_index, word_counts, components)`.
|
||||
--- It stores the ordered `items` stream plus the dense `word_events` / `markers` / `invocations` views on `atom.paths`.
|
||||
---
|
||||
--- Public boundary:
|
||||
--- * `M.run(ctx)` is the only entry point.
|
||||
--- * The pass returns `{outputs = {}, errors = ..., warnings = ...}`.
|
||||
--- Pass kind = `validation` → `PASS_KIND_STOP_ON_ERROR.validation` keeps build-stopping semantics (no policy change in this task).
|
||||
--- Pass kind = `validation` → `PASS_KIND_STOP_ON_ERROR.validation` preserves the existing build-stopping policy.
|
||||
---
|
||||
--- Source-order discipline:
|
||||
--- * `corpus.source_order` is the canonical ordering of source records.
|
||||
--- * For each source, the pass iterates `src.scan.atoms` and `src.scan.raw_atoms` IN SOURCE ORDER, preserving declaration order.
|
||||
--- * `corpus.source_order` sets the source-record order.
|
||||
--- * Within each source, the pass visits `src.scan.atoms` and `src.scan.raw_atoms` in declaration order.
|
||||
---
|
||||
--- Per-atom projection fields on `atom.paths`:
|
||||
--- `tokens`, `line_in_body`, `items`, `word_events`, `markers`, `invocations`, `errors`, `warnings`.
|
||||
--- The dense views are built from `items` only; the pass never re-walks source text or tokens.
|
||||
--- The construction walk appends `items` and derives each dense view from that ordered stream.
|
||||
---
|
||||
--- Component expansion and construction validation:
|
||||
--- * known `mac_X(...)` calls recursively expand component bodies;
|
||||
--- * invocation records retain monotonic IDs, parent IDs, immediate call text, and the immutable outermost root call text;
|
||||
--- * component cycles retain balanced invocation boundaries and emit a `cycle` construction error without recursing indefinitely;
|
||||
--- * invocation construction stamps `debug_skip` from `corpus.components[name].debug_skip` at the construction site (no second pass, no source parse, no parallel lookup);
|
||||
--- * component cycles close balanced invocation boundaries and emit a `cycle` construction error at the recursive edge;
|
||||
--- * declared-vs-measured component word counts emit `count_mismatch` construction errors; opaque uncounted macros emit warnings.
|
||||
---
|
||||
--- The pass does NOT consult `_code_macros` / `_code_macro_bodies`. Those private tables are owned by `passes.scan_source` and stripped before this pass runs.
|
||||
--- `passes.scan_source` strips its private `_code_macros` / `_code_macro_bodies` tables before this pass runs.
|
||||
|
||||
local M = {}
|
||||
|
||||
@@ -39,10 +39,27 @@ local duffle = dofile(_bootstrap_dir .. "../duffle_paths.lua")
|
||||
-- ─────────────────────────────────────────────────────────────────────────
|
||||
|
||||
-- Convert the recursive walk's body-relative line numbers into physical source lines once.
|
||||
-- Consumers read these canonical fields rather than rebuilding line state or tokenizing source again.
|
||||
-- The walker builds `line_of` from `body_text` and stamps body-relative line numbers (1..N) into `item.line` and `invocation.call_line`.
|
||||
-- This function converts those values to physical source lines at the close site with the forwarded source `line_of` closure.
|
||||
--
|
||||
-- `call_line` discipline:
|
||||
-- * ROOT invocations (`inv.parent_id == 0`) receive body-relative `call_line` values directly from `M.LineIndex(body_text)` in the walker.
|
||||
-- The source `line_of` closure supplies physical lines at the close site, so this function converts each root value exactly once.
|
||||
-- * INNER invocations (`inv.parent_id ~= 0`) receive physical `call_line` values directly from the COMPONENT's `line_of` in the walker.
|
||||
-- Recursive descent forwards that closure through `corpus.component_body_index[name].line_of`; those values arrive physical and remain unchanged.
|
||||
--
|
||||
-- After this function, every `inv.call_line` is physical. DWARF and provenance output read it directly.
|
||||
-- The word-event loop forwards the already-physical `outer_inv.call_line` into `we.call_line` for words inside an invocation.
|
||||
local function stamp_root_provenance(projection, atom_record, src, corpus)
|
||||
local root_line_of = src.scan and src.scan.line_of
|
||||
local root_body_line = root_line_of and root_line_of((atom_record.body_off or 1) - 1)
|
||||
assert(type(root_line_of) == "function"
|
||||
, "emission_model: src.scan.line_of is required (canonical LineIndex closure over the source text) to stamp physical provenance")
|
||||
assert(type(atom_record.body_off) == "number"
|
||||
, "emission_model: atom_record.body_off (byte offset of the body's first byte in source) is required to derive `root_body_line`. The scanner must populate body_off for every atom record.")
|
||||
-- `root_body_line` is the physical source line of the ATOM HEADER byte containing the opening `{`; that byte is one byte BEFORE `atom_record.body_off`.
|
||||
-- The walker assigns line 2 to the body's first content line because line 1 is the trailing `\n` after `{`. Body-text line k therefore maps to `root_body_line + (k - 1)`.
|
||||
-- `body_off - 1` points at the opening `{`, whose line index identifies the header line. `body_off` points after `{` and would shift every word row forward by one line.
|
||||
local root_body_line = root_line_of(atom_record.body_off - 1)
|
||||
or atom_record.line or 0
|
||||
local component_index = corpus.component_body_index or {}
|
||||
local word_items = {}
|
||||
@@ -51,27 +68,32 @@ local function stamp_root_provenance(projection, atom_record, src, corpus)
|
||||
if item.kind == "word" then word_items[#word_items + 1] = item end
|
||||
end
|
||||
|
||||
-- Resolve one word's physical body line, where the byte containing that word appears in source.
|
||||
-- * Component expansions carry `invocation_ids`; the component's full-file `line_of` leaves `item.line` physical.
|
||||
-- * Raw tokens in the root atom body carry an empty `invocation_ids` list and a body-relative `item.line`; convert them here.
|
||||
local function body_line_for(event, item)
|
||||
local body_line_of = root_line_of
|
||||
local body_off = atom_record.body_off or 0
|
||||
local ids = event.invocation_ids or {}
|
||||
local inner_id = ids[#ids]
|
||||
local inner_inv = inner_id and projection.invocations[inner_id]
|
||||
if inner_inv then
|
||||
local component = component_index[inner_inv.component_name]
|
||||
if component and component.line_of then
|
||||
-- Component-body walkers already receive the declaration source's full line index, so their item.line is physical.
|
||||
return item.line or 0
|
||||
local ids = event.invocation_ids or {}
|
||||
-- The innermost open invocation identifies which line index the walker used.
|
||||
-- A component `line_of` makes `item.line` physical; the atom's `body_text` line index makes it body-relative.
|
||||
if ids and #ids > 0 then
|
||||
local inner_id = ids[#ids]
|
||||
local inner_inv = inner_id and projection.invocations[inner_id]
|
||||
if inner_inv then
|
||||
local component = component_index[inner_inv.component_name]
|
||||
if component and component.line_of then
|
||||
-- Walker used `comp.line_of`, which is the source's physical LineIndex. item.line is already physical.
|
||||
return item.line or 0
|
||||
end
|
||||
end
|
||||
end
|
||||
local first_line = body_line_of and body_line_of(math.max(1, body_off - 1)) or root_body_line
|
||||
return (first_line or 0) + (item.line or 1) - 1
|
||||
-- RAW root-body word: item.line is body-text's 1-based line number (the first content line is line 2 because line 1 is the trailing `\n` after `{`).
|
||||
-- Convert body-text-relative → physical using `root_body_line + (item.line - 1)`.
|
||||
return (root_body_line or 0) + (item.line or 1) - 1
|
||||
end
|
||||
|
||||
-- Stamp root-source path onto invocation records whose `call_path` was left empty by the walker.
|
||||
-- The walker passes `body_entry.source` to `emit_invoke_begin` as the call_path argument; for the root body_entry created by `M.project_emission` that source is ""
|
||||
-- (the caller passes only the body text).
|
||||
-- After this stamp every invocation record has a physical call_path that matches what `passes/atoms_source_map.lua` matches the in-memory provenance projection.
|
||||
-- Stamp the root source path onto invocation records whose `call_path` the walker left empty.
|
||||
-- The walker passes `body_entry.source` to `emit_invoke_begin`; `M.project_emission` creates the root `body_entry` with source `""`, leaving its `call_path` empty.
|
||||
-- This stamp gives every invocation a physical `call_path` matching `passes/atoms_source_map.lua`'s in-memory provenance projection.
|
||||
local root_path = src.path or ""
|
||||
for _, inv in ipairs(projection.invocations) do
|
||||
if inv.call_path == nil or inv.call_path == "" then
|
||||
@@ -79,6 +101,35 @@ local function stamp_root_provenance(projection, atom_record, src, corpus)
|
||||
end
|
||||
end
|
||||
|
||||
-- Normalize `inv.call_line` to a physical source line.
|
||||
-- * ROOT invocations (`parent_id == 0`) carry body-relative `call_line` values from `M.LineIndex(body_text)`; convert them once with `root_body_line`.
|
||||
-- * INNER invocations (`parent_id ~= 0`) carry physical `call_line` values from the component's `line_of`; retain them unchanged.
|
||||
for _, inv in ipairs(projection.invocations) do
|
||||
if inv.parent_id == 0 then
|
||||
inv.call_line = (root_body_line or 0) + (inv.call_line or 1) - 1
|
||||
end
|
||||
end
|
||||
|
||||
-- Build `body_lines` for each invocation.
|
||||
-- `atoms_source_map` and `dwarf_injection` read `inv.body_lines[k]` directly from the invocation record created here.
|
||||
-- Component words already carry physical `item.line` values from the walker's COMPONENT line index, so `body_line_for` returns them unchanged.
|
||||
for _, inv in ipairs(projection.invocations) do
|
||||
local sw = inv.start_word
|
||||
local ew = inv.end_word
|
||||
local bls = {}
|
||||
for i = sw, ew do
|
||||
local it = projection.items and projection.items[i]
|
||||
if it and it.kind == "word" then
|
||||
local fake_event = { invocation_ids = { inv.id } }
|
||||
bls[#bls + 1] = body_line_for(fake_event, it) or 0
|
||||
end
|
||||
end
|
||||
inv.body_lines = bls
|
||||
end
|
||||
|
||||
-- Resolve each `word_event`'s physical `body_line` and `call_line`.
|
||||
-- For words inside an invocation, `we.call_line` identifies the OUTER atom source line containing the `mac_X(...)` token that triggered expansion.
|
||||
-- The root-invocation conversion above makes every `inv.call_line` physical; forward it directly and use each raw word's `body_line` as the fallback.
|
||||
for index, we in ipairs(projection.word_events) do
|
||||
local item = word_items[index] or {}
|
||||
local body_line = body_line_for(we, item)
|
||||
@@ -88,7 +139,10 @@ local function stamp_root_provenance(projection, atom_record, src, corpus)
|
||||
local call_line = body_line
|
||||
local outer_id = we.outermost_invocation_id or 0
|
||||
local outer_inv = projection.invocations[outer_id]
|
||||
if outer_inv then call_line = (root_body_line or 0) + (outer_inv.call_line or 1) - 1 end
|
||||
if outer_inv then
|
||||
-- `outer_inv.call_line` is physical after the conversion loop above, so use it directly.
|
||||
call_line = outer_inv.call_line
|
||||
end
|
||||
we.call_line = call_line
|
||||
|
||||
if we.def_path == nil or we.def_path == "" then we.def_path = src.path or "" end
|
||||
@@ -103,7 +157,10 @@ local function project_atom(atom_record, src, corpus)
|
||||
local body = atom_record.body or ""
|
||||
local wc = corpus.word_counts or {}
|
||||
local cbi = corpus.component_body_index or {}
|
||||
local proj = duffle.project_emission(body, cbi, wc)
|
||||
-- Stamp invocation-level `debug_skip` during construction (Task 5 of atom_component_skip_semantics_20260725).
|
||||
-- I found the declaration metadata in `corpus.components`; the walker forwards that registry to `emit_invoke_begin` in `duffle.lua`.
|
||||
-- That construction site stamps `invocation.debug_skip` while appending each record to `proj.invocations`.
|
||||
local proj = duffle.project_emission(body, cbi, wc, corpus.components)
|
||||
local paths = {
|
||||
tokens = atom_record.body_tokens or {},
|
||||
line_in_body = duffle.build_body_line_index(body),
|
||||
@@ -144,7 +201,7 @@ function M.run(ctx)
|
||||
end
|
||||
local proj = project_atom(atom, src, corpus)
|
||||
for _, e in ipairs(proj.errors) do
|
||||
-- Preserve `kind` (cycle / count_mismatch / unbalanced) so readers can dispatch on the diagnostic class without re-parsing the message string.
|
||||
-- Preserve `kind` (cycle / count_mismatch / unbalanced) so readers dispatch on the diagnostic class and leave the message string as display text.
|
||||
errors[#errors + 1] = {
|
||||
kind = e.kind,
|
||||
line = e.line,
|
||||
@@ -161,7 +218,7 @@ function M.run(ctx)
|
||||
end
|
||||
end
|
||||
|
||||
-- Walk every source in canonical order; for each source, iterate atoms + raw_atoms.
|
||||
-- Walk `corpus.source_order`; within each source, visit atoms followed by raw_atoms.
|
||||
-- Recognized kinds (atom | raw_atom | comp_bare | comp_proc) each receive the atom.paths projection via duffle.project_emission.
|
||||
-- Components are macros inlined into atom bodies; focused tests and isolated component analyses consume atom.paths directly.
|
||||
for _, src in ipairs(corpus.source_order) do
|
||||
|
||||
@@ -5,10 +5,8 @@
|
||||
--- - `build/gen/<dir_basename>.annotations.txt` — one per source-directory containing atoms; aggregates across all sources in the directory.
|
||||
--- - `build/gen/annotation_validation.txt` — the project summary.
|
||||
---
|
||||
--- The annotation pass emits `errors.h` files per module and the canonical
|
||||
--- `corpus.sources_by_dir` projection groups sources by directory. This pass
|
||||
--- iterates the canonical dir projection directly and re-validates each source
|
||||
--- via `annotation.validate()` to get the detailed per-source results.
|
||||
--- The annotation pass emits `errors.h` files per module and the canonical `corpus.sources_by_dir` projection groups sources by directory.
|
||||
--- This pass iterates the canonical dir projection directly and re-validates each source via `annotation.validate()` to get the detailed per-source results.
|
||||
---
|
||||
--- **Conventions**: tabs (1/level), EmmyLua annotations, no regex,
|
||||
--- Lua 5.3 compatible.
|
||||
@@ -27,11 +25,9 @@
|
||||
local _bootstrap_dir = debug.getinfo(1, "S").source:match("^@?(.*[/\\])") or "./"
|
||||
local duffle = dofile(_bootstrap_dir .. "../duffle_paths.lua")
|
||||
|
||||
-- Load the annotation pass so we can re-validate each source against the
|
||||
-- canonical corpus projection. The annotation pass exposes `M.validate`,
|
||||
-- which returns the per-source AnnotationResult (atoms / annots / macros /
|
||||
-- binds / errors / warnings) that the report pass renders into the
|
||||
-- per-module `<dir_basename>.annotations.txt` output.
|
||||
-- Load the annotation pass so we can re-validate each source against the canonical corpus projection.
|
||||
-- The annotation pass exposes `M.validate`, which returns the per-source AnnotationResult (atoms / annots / macros / binds / errors / warnings)
|
||||
-- that the report pass renders into the per-module `<dir_basename>.annotations.txt` output.
|
||||
local annotation = dofile(_bootstrap_dir .. "annotation.lua")
|
||||
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
@@ -382,11 +378,9 @@ end
|
||||
-- Orchestration helpers
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
--- (internal) Re-validate every source in a directory against the canonical
|
||||
--- corpus projection. Calls `annotation.validate()` per source to produce the
|
||||
--- per-source AnnotationResult (atoms / annots / macros / binds / errors /
|
||||
--- warnings) that the report renderer consumes. This is the canonical path —
|
||||
--- no private stash; each report pass run is reproducible from the corpus.
|
||||
--- (internal) Re-validate every source in a directory against the canonical corpus projection.
|
||||
--- Calls `annotation.validate()` per source to produce the per-source AnnotationResult (atoms / annots / macros / binds / errors / warnings)
|
||||
--- that the report renderer consumes. Eeach report pass run is reproducible from the corpus.
|
||||
--- Returns the list of module results + the flat list of all results (for the project-wide summary).
|
||||
--- @param ctx PassCtx
|
||||
--- @param dir_sources SourceFile[]
|
||||
|
||||
+264
-181
@@ -1,12 +1,12 @@
|
||||
--- passes/scan_source.lua — Source pre-scan pass (the "mega entity" pass).
|
||||
---
|
||||
--- Single source-walk pass that produces the fat `SourceScan` payload consumed by all downstream passes. Walks each `ctx.sources` entry once,
|
||||
--- Single source-walk pass that produces the fat `SourceScan` payload consumed by all downstream passes. Walks each corpus source record once,
|
||||
--- extracting every construct type the metaprograms need:
|
||||
---
|
||||
--- MipsAtom_ (kind = "atom", with optional atom_info inner)
|
||||
--- MipsAtomComp_ (kind = "comp_bare")
|
||||
--- MipsAtomComp_Proc_ (kind = "comp_proc", body inside last {})
|
||||
--- atom_dbg_skip_over (whole-atom/component debug-step marker; following declaration disambiguates)
|
||||
--- atom_dbg_skip — bare whole-atom/component debug-step marker; following declaration disambiguates
|
||||
--- MipsCode code_<name> (kind = "raw_atom", offsets pass only)
|
||||
--- typedef Struct_(Binds_X) { fields }
|
||||
--- #pragma mac_X tape_atom words=N + _Pragma("...")
|
||||
@@ -37,38 +37,29 @@ local parse_enum_int_literal
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
--- @class SourceScan
|
||||
--- @field atoms AtomEntry[] -- MipsAtom_ + MipsAtomComp_ + MipsAtomComp_Proc_
|
||||
--- @field raw_atoms AtomEntry[] -- MipsCode code_<name> { body } (offsets pass only)
|
||||
--- @field binds BindsEntry[] -- typedef Struct_(Binds_X) { fields } (fields pre-parsed)
|
||||
--- @field atom_infos AtomInfoEntry[] -- MipsAtom_(name) atom_info(...) (sub-calls pre-parsed)
|
||||
--- @field macros MacroEntry[] -- #pragma mac_X tape_atom words=N + _Pragma("...")
|
||||
--- @field skip_over SkipOverScan -- atom/component debug-step markers + resolved declaration associations
|
||||
--- @field types table<string, RegTypeDefault> -- atom_dbg_reg_default(R_X, <type>) declarations
|
||||
--- @field atom_views table<string, AtomViewEntry> -- MipsAtom_(name) -> {binds_name, reg_type_overrides, info_line}
|
||||
--- @field atom_ctxs table<string, AtomCtxEntry> -- MipsAtom_(name) -> {rbind_atom, info_line, source} (atom_ctx(...) call sites)
|
||||
--- @field atom_phases table<string, AtomPhaseGroup> -- phase_label -> {atoms = {atom_name1, atom_name2, ...}} (atom_phase(...) tags)
|
||||
--- @field line_of fun(pos: integer): integer -- shared LineIndex closure
|
||||
--- @field atoms AtomEntry[] -- MipsAtom_ + MipsAtomComp_ + MipsAtomComp_Proc_
|
||||
--- @field raw_atoms AtomEntry[] -- MipsCode code_<name> { body } (offsets pass only)
|
||||
--- @field binds BindsEntry[] -- typedef Struct_(Binds_X) { fields } (fields pre-parsed)
|
||||
--- @field atom_infos AtomInfoEntry[] -- MipsAtom_(name) atom_info(...) (sub-calls pre-parsed)
|
||||
--- @field macros MacroEntry[] -- #pragma mac_X tape_atom words=N + _Pragma("...")
|
||||
--- @field debug_skip_markers DebugSkipMarker[] -- raw marker evidence for annotation validation; `debug_skip` lives on the declaration record itself
|
||||
--- @field types table<string, RegTypeDefault> -- atom_dbg_reg_default(R_X, <type>) declarations
|
||||
--- @field atom_views table<string, AtomViewEntry> -- MipsAtom_(name) -> {binds_name, reg_type_overrides, info_line}
|
||||
--- @field atom_ctxs table<string, AtomCtxEntry> -- MipsAtom_(name) -> {rbind_atom, info_line, source} (atom_ctx(...) call sites)
|
||||
--- @field atom_phases table<string, AtomPhaseGroup> -- phase_label -> {atoms = {atom_name1, atom_name2, ...}} (atom_phase(...) tags)
|
||||
--- @field line_of fun(pos: integer): integer -- shared LineIndex closure
|
||||
|
||||
--- @class SkipOverScan
|
||||
--- @field atoms table<string, SkipOverAssociation>
|
||||
--- @field components table<string, SkipOverAssociation>
|
||||
--- @field markers SkipOverMarker[]
|
||||
|
||||
--- @class SkipOverMarker
|
||||
--- @field marker_kind string -- exact marker ident (always "atom_dbg_skip_over")
|
||||
--- @field marker_line integer
|
||||
--- @field marker_pos integer
|
||||
--- @field after_paren integer
|
||||
--- @field args string|nil -- trimmed marker args""
|
||||
--- @field has_parens boolean
|
||||
--- @field pending boolean
|
||||
--- @field superseded_by_marker_line integer|nil
|
||||
--- @field target_name string|nil -- stripped declaration name once observed
|
||||
--- @field target_raw_name string|nil -- source-written declaration name once observed
|
||||
--- @field target_kind string|nil -- "atom" | "comp_bare" | "comp_proc" | "unrelated" once observed
|
||||
--- @field declaration_line integer|nil
|
||||
--- @field declaration_pos integer|nil
|
||||
--- @field proc_prelude boolean|nil -- marker has crossed FI_ and awaits MipsAtomComp_Proc_
|
||||
--- @class DebugSkipMarker
|
||||
--- @field marker_kind string -- exact marker ident read from source. Only "atom_dbg_skip" (bare) is positive; any other ident reaches the unrelated fallback and is never associated with a declaration.
|
||||
--- @field marker_line integer -- line of the marker ident start
|
||||
--- @field marker_pos integer -- byte position of the marker ident start (the comment walker anchors here)
|
||||
--- @field is_bare boolean -- true iff marker_kind == "atom_dbg_skip" AND has_parens == false (the only positive form)
|
||||
--- @field has_parens boolean -- true iff a `(...)` follows the marker ident (diagnostic-only)
|
||||
--- @field args string|nil -- trimmed args inside the `(...)` (nil when has_parens is false)
|
||||
--- @field pending boolean -- true while awaiting the following declaration
|
||||
--- @field superseded_by_marker_line integer|nil -- set when a newer marker bumped this one out of the pending slot
|
||||
--- @field target_kind string|nil -- "atom" | "comp_bare" | "comp_proc" | "unrelated" once observed (nil if no declaration ever followed)
|
||||
--- @field proc_prelude boolean|nil -- true after the marker crossed an `FI_` prelude and awaits `MipsAtomComp_Proc_`
|
||||
|
||||
--- @class RegTypeDefault
|
||||
--- @field name string -- "R_TapePtr" (the register ident; without the value part)
|
||||
@@ -96,12 +87,6 @@ local parse_enum_int_literal
|
||||
--- @field reg_type_overrides table<string, RegTypeOverride> -- "R_T0" -> override
|
||||
--- @field info_line integer -- line of the atom_info call
|
||||
|
||||
--- @class SkipOverAssociation
|
||||
--- @field marker_line integer
|
||||
--- @field declaration_line integer
|
||||
--- @field kind string
|
||||
--- @field marker SkipOverMarker
|
||||
|
||||
--- @class SourceFile
|
||||
--- @field path string -- absolute path to the source file
|
||||
--- @field text string -- the full source text
|
||||
@@ -125,16 +110,16 @@ local parse_enum_int_literal
|
||||
--- @field warnings table[]
|
||||
|
||||
--- @class AtomEntry
|
||||
--- @field line integer
|
||||
--- @field name string -- atom name (for components: without ac_ prefix)
|
||||
--- @field body string -- brace-delimited body (without the braces)
|
||||
--- @field body_off integer -- char offset of body[1] in source
|
||||
--- @field kind string -- "atom" | "comp_bare" | "comp_proc" | "raw_atom"
|
||||
--- @field raw_name string -- un-stripped name (for components: with ac_ prefix)
|
||||
--- @field ident_pos integer -- position of the MipsAtom_/MipsAtomComp_ ident start
|
||||
--- @field after_paren integer -- position past the closing paren
|
||||
--- @field args string|nil -- populated by components pass (backward lookup)
|
||||
--- @field comment string|nil -- populated by components pass (backward lookup)
|
||||
--- @field line integer
|
||||
--- @field name string -- atom name (for components: without ac_ prefix)
|
||||
--- @field body string -- brace-delimited body (without the braces)
|
||||
--- @field body_off integer -- char offset of body[1] in source
|
||||
--- @field kind string -- "atom" | "comp_bare" | "comp_proc" | "raw_atom"
|
||||
--- @field raw_name string -- un-stripped name (for components: with ac_ prefix)
|
||||
--- @field ident_pos integer -- position of the MipsAtom_/MipsAtomComp_ ident start
|
||||
--- @field after_paren integer -- position past the closing paren
|
||||
--- @field debug_skip boolean -- true when an `atom_dbg_skip` bare marker immediately precedes this declaration (sole-owner stamp; see push_debug_skip_marker)
|
||||
--- @field declaration_comment string|nil -- populated by the scanner (backward walk past the marker, captures contiguous `/* */` or `//` block)
|
||||
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
-- Local helpers (shared by per-form parsers)
|
||||
@@ -166,8 +151,10 @@ local function strip_ac_prefix(raw_name)
|
||||
end
|
||||
|
||||
-- Preserve a source marker until the following declaration parser observes it.
|
||||
local function push_skip_over_marker(out, marker)
|
||||
local markers = out.skip_over.markers
|
||||
-- The scanner is the sole owner of marker recognition, placement association, declaration comment attachment, and canonical `debug_skip` fields.
|
||||
-- Raw marker evidence lives in `out.debug_skip_markers` for annotation validation; the declaration record carries the resolved `debug_skip` boolean directly.
|
||||
local function push_debug_skip_marker(out, marker)
|
||||
local markers = out.debug_skip_markers
|
||||
local prior = markers[#markers]
|
||||
if prior and prior.pending then
|
||||
prior.pending = false
|
||||
@@ -197,55 +184,150 @@ local function find_body_braces(source, after_paren, fallback)
|
||||
return body, after_brace, brace + 1
|
||||
end
|
||||
|
||||
-- Attach the pending marker to the next declaration.
|
||||
-- The declaration form disambiguates whole atoms from components;
|
||||
-- unsupported declarations retain placement evidence for annotation.lua and populate neither lookup table.
|
||||
local function associate_skip_over_marker(out, target_name, target_raw_name, target_kind, declaration_line, declaration_pos)
|
||||
local markers = out.skip_over.markers
|
||||
local marker = markers[#markers]
|
||||
if not (marker and marker.pending) then return end
|
||||
-- Walk backward from `start_pos` capturing contiguous `/* */` block(s) and
|
||||
-- `//` line(s) that immediately precede it. The caller (preceding_declaration_comment)
|
||||
-- supplies `start_pos` so the walker does not need to detect marker shape or prelude layout.
|
||||
-- The scanner already knows the marker_pos + decl ident_pos and threads that knowledge forward.
|
||||
--
|
||||
-- The walker captures:
|
||||
-- - Block comment close `*/` followed by walking back to `/*`.
|
||||
-- - `//` line comments (the line containing the current non-ws position starts with `//`).
|
||||
-- It stops at the first non-ws char that does not begin a comment block or line.
|
||||
-- Empty string if no comment is adjacent.
|
||||
-- @param source string
|
||||
-- @param start_pos integer -- exclusive upper bound for the captured block
|
||||
-- @return string
|
||||
local function preceding_comment_walk_backward(source, start_pos)
|
||||
local pieces = {}
|
||||
local scan_pos = start_pos
|
||||
while scan_pos > 0 do
|
||||
local non_ws = scan_pos - 1
|
||||
while non_ws > 0 do
|
||||
local ch = source:sub(non_ws, non_ws)
|
||||
if ch == " " or ch == "\t" or ch == "\n" or ch == "\r" then
|
||||
non_ws = non_ws - 1
|
||||
else
|
||||
break
|
||||
end
|
||||
end
|
||||
if non_ws == 0 then break end
|
||||
|
||||
marker.pending = false
|
||||
marker.target_name = target_name
|
||||
marker.target_raw_name = target_raw_name
|
||||
marker.target_kind = target_kind
|
||||
marker.declaration_line = declaration_line
|
||||
marker.declaration_pos = declaration_pos
|
||||
|
||||
if not (marker.has_parens and marker.args == "") then return end
|
||||
|
||||
local association = {
|
||||
marker_line = marker.marker_line,
|
||||
declaration_line = declaration_line,
|
||||
kind = target_kind,
|
||||
marker = marker,
|
||||
}
|
||||
if target_kind == "atom" then
|
||||
out.skip_over.atoms[target_name] = association
|
||||
elseif target_kind == "comp_bare" or target_kind == "comp_proc" then
|
||||
out.skip_over.components[target_name] = association
|
||||
if non_ws >= 2 and source:sub(non_ws - 1, non_ws) == "*/" then
|
||||
-- Block comment close: walk back over `/*` candidates.
|
||||
local prefix = source:sub(1, non_ws - 1)
|
||||
local open_at = nil
|
||||
for scan = #prefix - 1, 1, -1 do
|
||||
if prefix:sub(scan, scan + 1) == "/*" then
|
||||
open_at = scan
|
||||
break
|
||||
end
|
||||
end
|
||||
if not open_at then break end
|
||||
local block_start = open_at
|
||||
while block_start > 1 do
|
||||
local ch = source:sub(block_start - 1, block_start - 1)
|
||||
if ch ~= " " and ch ~= "\t" then break end
|
||||
block_start = block_start - 1
|
||||
end
|
||||
table.insert(pieces, 1, source:sub(block_start, non_ws))
|
||||
scan_pos = block_start
|
||||
else
|
||||
-- Line comment check: walk back from non_ws to the most recent `\n`
|
||||
-- (or position 1) and inspect the resulting line. This handles both
|
||||
-- `// foo\n<marker>` (non_ws ends on `o`) and `// foo\r\n<marker>`.
|
||||
local line_start = non_ws
|
||||
while line_start > 1 and source:sub(line_start - 1, line_start - 1) ~= "\n" do
|
||||
line_start = line_start - 1
|
||||
end
|
||||
local line = source:sub(line_start, non_ws)
|
||||
if line:sub(1, 2) ~= "//" then break end
|
||||
table.insert(pieces, 1, line)
|
||||
scan_pos = line_start - 1
|
||||
end
|
||||
end
|
||||
if #pieces == 0 then return "" end
|
||||
return table.concat(pieces, "\n")
|
||||
end
|
||||
|
||||
-- Register a parsed atom entry in `out.atoms` and link its skip-over marker.
|
||||
-- Captures the shared 8-field shape used by MipsAtom_, MipsAtomComp_, MipsAtomComp_Proc_.
|
||||
local function register_atom(out, kind, declaration_line, name, body, body_off, raw_name, pos, after_paren)
|
||||
-- Resolve the start position for the declaration-comment walk.
|
||||
-- When a debug-skip marker is pending, the walker must start from the position immediately before the marker ident
|
||||
-- (so it walks backward past the marker text and any `FI_ MipsAtom ac_X(args)` proc-prelude layout — neither of which is visible if we start from the declaration ident_pos).
|
||||
-- When no marker is pending, the walker starts from the declaration ident_pos directly.
|
||||
-- @param pending_marker DebugSkipMarker|nil
|
||||
-- @param ident_pos integer -- declaration ident position
|
||||
-- @return integer
|
||||
local function comment_walk_start(pending_marker, ident_pos)
|
||||
if pending_marker then
|
||||
return pending_marker.marker_pos - 1
|
||||
end
|
||||
return ident_pos - 1
|
||||
end
|
||||
|
||||
-- Attach the pending marker to the next declaration.
|
||||
-- The declaration form disambiguates whole atoms from components; the resolved `debug_skip` is stamped directly on the declaration record
|
||||
-- (sole-owner discipline; see push_debug_skip_marker).
|
||||
--
|
||||
-- A marker is POSITIVE (stamps `debug_skip = true` on the declaration) iff:
|
||||
-- marker_kind == "atom_dbg_skip" AND is_bare == true
|
||||
-- Any other spelling or shape (parenthesized form, legacy name) is recorded as a raw marker for annotation validation but never stamps `debug_skip`.
|
||||
-- @param out SourceScan
|
||||
-- @param target_kind string|nil -- "atom" | "comp_bare" | "comp_proc" | "unrelated" once observed
|
||||
-- @return boolean|nil -- true iff the marker is the positive bare form
|
||||
local function attach_debug_skip_marker(out, target_kind)
|
||||
local markers = out.debug_skip_markers
|
||||
local marker = markers[#markers]
|
||||
if not (marker and marker.pending) then return nil end
|
||||
|
||||
marker.pending = false
|
||||
marker.target_kind = target_kind
|
||||
|
||||
if marker.marker_kind == "atom_dbg_skip" and marker.is_bare then
|
||||
return true
|
||||
end
|
||||
return nil
|
||||
end
|
||||
|
||||
-- Register a parsed atom entry in `out.atoms`. Stamps the resolved `debug_skip` boolean
|
||||
-- on the record when a positive bare `atom_dbg_skip` marker is pending.
|
||||
-- Captures the shared shape used by MipsAtom_, MipsAtomComp_, MipsAtomComp_Proc_.
|
||||
local function register_atom(out, kind, declaration_line, name, body, body_off, raw_name, pos, after_paren, source)
|
||||
-- Capture the pending marker BEFORE attaching so the walker can anchor the backward comment walk on the marker's marker_pos
|
||||
-- (which is the correct anchor even when an `FI_ MipsAtom ac_X(args)` proc-prelude separates the marker from the declaration).
|
||||
local pending_marker = nil
|
||||
local markers = out.debug_skip_markers
|
||||
local m = markers[#markers]
|
||||
if m and m.pending then pending_marker = m end
|
||||
|
||||
local positive = attach_debug_skip_marker(out, kind)
|
||||
local comment = ""
|
||||
if kind == "comp_bare" or kind == "comp_proc" then
|
||||
-- Scanner-owned declaration-comment attachment.
|
||||
-- The walker does not need to detect marker shape.
|
||||
-- A pending_marker record (or the declaration ident_pos fallback) supplies the anchor position.
|
||||
local start_pos = comment_walk_start(pending_marker, pos)
|
||||
comment = preceding_comment_walk_backward(source, start_pos)
|
||||
end
|
||||
out.atoms[#out.atoms + 1] = {
|
||||
line = declaration_line, name = name, body = body, body_off = body_off,
|
||||
kind = kind, raw_name = raw_name,
|
||||
ident_pos = pos, after_paren = after_paren,
|
||||
line = declaration_line,
|
||||
name = name,
|
||||
body = body,
|
||||
body_off = body_off,
|
||||
kind = kind,
|
||||
raw_name = raw_name,
|
||||
ident_pos = pos,
|
||||
after_paren = after_paren,
|
||||
debug_skip = positive == true,
|
||||
declaration_comment = comment,
|
||||
}
|
||||
associate_skip_over_marker(out, name, raw_name, kind, declaration_line, pos)
|
||||
end
|
||||
|
||||
-- Register a parsed raw-atom entry in `out.raw_atoms` and link its skip-over marker.
|
||||
-- Register a parsed raw-atom entry in `out.raw_atoms`.
|
||||
-- Captures the 5-field shape used by MipsCode (the raw-atom form; offsets pass only).
|
||||
local function register_raw_atom(out, declaration_line, name, body, body_off, raw_name, pos, marker_kind)
|
||||
local function register_raw_atom(out, declaration_line, name, body, body_off, raw_name, pos)
|
||||
out.raw_atoms[#out.raw_atoms + 1] = {
|
||||
line = declaration_line, name = name, body = body, body_off = body_off,
|
||||
kind = "raw_atom", raw_name = raw_name,
|
||||
}
|
||||
associate_skip_over_marker(out, name, raw_name, marker_kind, declaration_line, pos)
|
||||
end
|
||||
|
||||
-- Parse a `Type*` chain (zero or more `*` separated by optional whitespace) followed by the type ident.
|
||||
@@ -354,7 +436,7 @@ end
|
||||
|
||||
-- Parse the `Enum_(<underlying>, <name>) { <body> }` body for entries.
|
||||
-- Captures one field per named enumerator with the shape { name, value }.
|
||||
-- The value is the integer literal parsed from the source via the canonical `parse_enum_int_literal`.
|
||||
-- The value is the integer literal parsed from the source via `parse_enum_int_literal`.
|
||||
local function parse_enum_body_fields(body)
|
||||
return walk_body_fields(body, function(entry_name, name_end, after_name)
|
||||
local value
|
||||
@@ -652,20 +734,13 @@ local function scan_atom_info_subcalls(info_inner, info_line)
|
||||
}
|
||||
end
|
||||
local SUBCALL_HANDLERS = {
|
||||
-- scan: atom_bind(<Binds_X>)
|
||||
atom_bind = function(sub_inner) binds = duffle.trim(sub_inner) end,
|
||||
-- scan: atom_reads(<R_X [atom_type(<T>)], ...>)
|
||||
atom_reads = function(sub_inner, info_line) rw_handler(sub_inner, info_line, "atom_reads") end,
|
||||
-- scan: atom_writes(<R_X [atom_type(<T>)], ...>)
|
||||
atom_writes = function(sub_inner, info_line) rw_handler(sub_inner, info_line, "atom_writes") end,
|
||||
-- scan: atom_view(<Binds_X>)
|
||||
atom_view = function(sub_inner) view_binds = duffle.trim(sub_inner) end,
|
||||
-- scan: atom_reg_types(<R_X>, <T>)
|
||||
atom_reg_types = reg_types_handler,
|
||||
-- scan: atom_ctx(<atom_name>)
|
||||
atom_ctx = function(sub_inner, info_line) ident_handler(sub_inner, info_line, "ctx_atom_name") end,
|
||||
-- scan: atom_phase(<label>)
|
||||
atom_phase = function(sub_inner, info_line) ident_handler(sub_inner, info_line, "phase_label") end,
|
||||
atom_bind = function(sub_inner) binds = duffle.trim(sub_inner) end, -- scan: atom_bind(<Binds_X>)
|
||||
atom_reads = function(sub_inner, info_line) rw_handler(sub_inner, info_line, "atom_reads") end, -- scan: atom_reads(<R_X [atom_type(<T>)], ...>)
|
||||
atom_writes = function(sub_inner, info_line) rw_handler(sub_inner, info_line, "atom_writes") end, -- scan: atom_writes(<R_X [atom_type(<T>)], ...>)
|
||||
atom_view = function(sub_inner) view_binds = duffle.trim(sub_inner) end, -- scan: atom_view(<Binds_X>)
|
||||
atom_reg_types = reg_types_handler, -- scan: atom_reg_types(<R_X>, <T>)
|
||||
atom_ctx = function(sub_inner, info_line) ident_handler(sub_inner, info_line, "ctx_atom_name") end, -- scan: atom_ctx(<atom_name>)
|
||||
atom_phase = function(sub_inner, info_line) ident_handler(sub_inner, info_line, "phase_label") end, -- scan: atom_phase(<label>)
|
||||
}
|
||||
|
||||
local sub_pos = 1
|
||||
@@ -1006,7 +1081,7 @@ end
|
||||
-- pos -- position of the construct's leading ident (e.g., `M` of `MipsAtom_`)
|
||||
-- ident_end -- position past the leading ident (where the `(` should be)
|
||||
-- line_of -- closure over LineIndex(source) for 1-based line lookups
|
||||
-- out -- the SourceScan out table (mutated in place: out.atoms / out.raw_atoms / out.binds / out.atom_infos / out.macros / out.skip_over)
|
||||
-- out -- the SourceScan out table (mutated in place: out.atoms / out.raw_atoms / out.binds / out.atom_infos / out.macros / out.debug_skip_markers)
|
||||
-- returns -- new position after the construct
|
||||
--
|
||||
-- All parsers read source-as-written via the duffle primitives (skip_ws_and_cmt / read_parens / read_braces / read_balanced).
|
||||
@@ -1015,34 +1090,45 @@ end
|
||||
--
|
||||
-- Adding a new construct = 1 row in DECL_PARSERS + 1 parser function. The scan_source() loop never needs editing.
|
||||
|
||||
--- Parse an empty debug-skip marker and retain its raw placement evidence.
|
||||
--- The marker_kind is the source ident itself (e.g. `atom_dbg_skip_over`).
|
||||
--- The dispatch table maps each ident to this same function;
|
||||
--- the marker_kind is derived from the source so future idents route through the same row.
|
||||
--- Parse a `atom_dbg_skip` marker and record its raw placement evidence.
|
||||
---
|
||||
--- Positive path: the BARE form (`atom_dbg_skip` followed by whitespace + a supported declaration)
|
||||
--- stamps the `debug_skip` field on the immediately-following declaration record via `attach_debug_skip_marker`. `is_bare`
|
||||
--- is set true only when `marker_kind == "atom_dbg_skip"` and there are no parens.
|
||||
---
|
||||
--- Diagnostic-only path: a following `(...)` is recorded as an invalid parenthesized-form marker so the annotation rule can emit a precise "parenthesized form" diagnostic.
|
||||
--- The parenthesized form stays diagnostic; the bare form alone carries the runtime stamp.
|
||||
--- @param source string
|
||||
--- @param pos integer
|
||||
--- @param ident_end integer
|
||||
--- @param line_of fun(pos: integer): integer
|
||||
--- @param out SourceScan
|
||||
--- @return integer
|
||||
local function parse_skip_over_marker(source, pos, ident_end, line_of, out)
|
||||
local open_paren = duffle.skip_ws_and_cmt(source, ident_end)
|
||||
local marker = {
|
||||
marker_kind = source:sub(pos, ident_end - 1),
|
||||
marker_line = line_of(pos),
|
||||
marker_pos = pos,
|
||||
after_paren = ident_end,
|
||||
args = nil,
|
||||
has_parens = false,
|
||||
}
|
||||
--- @return integer -- source cursor position to resume from
|
||||
local function parse_dbg_skip_marker(source, pos, ident_end, line_of, out)
|
||||
local marker_kind = source:sub(pos, ident_end - 1)
|
||||
|
||||
-- Diagnostic-only detection of an invalid following `(...)`.
|
||||
-- The cursor is advanced past the `()` either way to keep token order coherent for the next scan iteration.
|
||||
local marker_end = ident_end
|
||||
local open_paren = duffle.skip_ws_and_cmt(source, ident_end)
|
||||
local has_parens = false
|
||||
local args = nil
|
||||
if source:sub(open_paren, open_paren) == "(" then
|
||||
local inner, after_paren = duffle.read_parens(source, open_paren)
|
||||
marker.after_paren = after_paren
|
||||
marker.args = duffle.trim(inner)
|
||||
marker.has_parens = true
|
||||
marker_end = after_paren
|
||||
has_parens = true
|
||||
args = duffle.trim(inner)
|
||||
end
|
||||
push_skip_over_marker(out, marker)
|
||||
return marker.after_paren
|
||||
|
||||
push_debug_skip_marker(out, {
|
||||
marker_kind = marker_kind,
|
||||
marker_line = line_of(pos),
|
||||
marker_pos = pos,
|
||||
is_bare = (marker_kind == "atom_dbg_skip") and (not has_parens),
|
||||
has_parens = has_parens,
|
||||
args = args,
|
||||
})
|
||||
return marker_end
|
||||
end
|
||||
|
||||
-- Parse `atom_dbg_reg_default(R_X, <type>...)`;
|
||||
@@ -1138,7 +1224,7 @@ local function parse_mips_atom(source, pos, ident_end, line_of, out)
|
||||
local body, after_brace, body_off = find_body_braces(source, brace_search_pos, open_paren + 1)
|
||||
if not body then return after_brace end
|
||||
if raw_name and raw_name ~= "" then
|
||||
register_atom(out, "atom", line_of(pos), raw_name, body, body_off, raw_name, pos, after_paren)
|
||||
register_atom(out, "atom", line_of(pos), raw_name, body, body_off, raw_name, pos, after_paren, source)
|
||||
end
|
||||
|
||||
return after_brace
|
||||
@@ -1161,7 +1247,7 @@ local function parse_mips_atom_comp(source, pos, ident_end, line_of, out)
|
||||
local body, after_brace, body_off = find_body_braces(source, after_paren, open_paren + 1)
|
||||
if not body then return after_brace end
|
||||
local name = strip_ac_prefix(raw_name)
|
||||
register_atom(out, "comp_bare", line_of(pos), name, body, body_off, raw_name, pos, after_paren)
|
||||
register_atom(out, "comp_bare", line_of(pos), name, body, body_off, raw_name, pos, after_paren, source)
|
||||
|
||||
return after_brace
|
||||
end
|
||||
@@ -1195,7 +1281,7 @@ local function parse_mips_atom_comp_proc(source, pos, ident_end, line_of, out)
|
||||
-- Position of body[1] in source = open_paren + 1 (start of inner) + last_brace_pos + 1 (past '{').
|
||||
local body_off = open_paren + 2 + last_brace_pos
|
||||
|
||||
register_atom(out, "comp_proc", line_of(pos), name, body, body_off, raw_name, pos, after_paren)
|
||||
register_atom(out, "comp_proc", line_of(pos), name, body, body_off, raw_name, pos, after_paren, source)
|
||||
|
||||
return after_paren
|
||||
end
|
||||
@@ -1217,7 +1303,7 @@ local function parse_mips_code(source, pos, ident_end, line_of, out)
|
||||
local atom_name = next_ident:sub(6)
|
||||
local body, after_brace, body_off = find_body_braces(source, next_after, ident_end)
|
||||
if not body then return after_brace end
|
||||
register_raw_atom(out, line_of(pos), atom_name, body, body_off, atom_name, pos, "unrelated")
|
||||
register_raw_atom(out, line_of(pos), atom_name, body, body_off, atom_name, pos)
|
||||
|
||||
return after_brace
|
||||
end
|
||||
@@ -1316,7 +1402,7 @@ end
|
||||
--- 4. `typedef <type> TSet_(<name>);` duffle TSet_ convention.
|
||||
--- Strips TSet_ wrapper; adds to type_name_registry (kind="typedef") with underlying_type=<type>.
|
||||
---
|
||||
--- All four shapes also associate an "unrelated" skip-over marker (the existing behavior — typedef declarations don't carry atom_dbg_skip_over).
|
||||
--- All four shapes also attach an "unrelated" debug-skip marker (the existing behavior — typedef declarations don't carry atom_dbg_skip).
|
||||
--- @param source string
|
||||
--- @param pos integer
|
||||
--- @param ident_end integer
|
||||
@@ -1337,7 +1423,7 @@ local function parse_typedef_binds(source, pos, ident_end, line_of, out)
|
||||
local body, after_brace = find_body_braces(source, after_paren, open_paren + 1)
|
||||
if not body then return after_brace end
|
||||
register_struct_type(body, name, pos, line_of, out)
|
||||
associate_skip_over_marker(out, name, name, "unrelated", line_of(pos), pos)
|
||||
attach_debug_skip_marker(out, "unrelated")
|
||||
return after_brace
|
||||
|
||||
-- ── Shape 2: `typedef Enum_(<underlying>, <name>) { <body> } <alias>;`
|
||||
@@ -1353,7 +1439,7 @@ local function parse_typedef_binds(source, pos, ident_end, line_of, out)
|
||||
local body, after_brace = find_body_braces(source, after_paren, open_paren + 1)
|
||||
if not body then return after_brace end
|
||||
register_enum_type(underlying, name, body, pos, line_of, out)
|
||||
associate_skip_over_marker(out, name, name, "unrelated", line_of(pos), pos)
|
||||
attach_debug_skip_marker(out, "unrelated")
|
||||
return after_brace
|
||||
end
|
||||
|
||||
@@ -1389,7 +1475,7 @@ local function parse_typedef_binds(source, pos, ident_end, line_of, out)
|
||||
-- Empty underlying span is acceptable; the TSet_ wrapper itself
|
||||
-- encodes the alias identity (per the duffle TSet_ convention).
|
||||
register_typedef_alias("", tset_name, pos, line_of, out)
|
||||
associate_skip_over_marker(out, tset_name, tset_name, "unrelated", line_of(pos), pos)
|
||||
attach_debug_skip_marker(out, "unrelated")
|
||||
return after_paren
|
||||
end
|
||||
|
||||
@@ -1432,7 +1518,7 @@ local function parse_typedef_binds(source, pos, ident_end, line_of, out)
|
||||
local underlying_span = source:sub(after_typedef, tset_pos - 1)
|
||||
local underlying = duffle.trim(underlying_span)
|
||||
register_typedef_alias(underlying, tset_arg, pos, line_of, out)
|
||||
associate_skip_over_marker(out, tset_arg, tset_arg, "unrelated", line_of(pos), pos)
|
||||
attach_debug_skip_marker(out, "unrelated")
|
||||
return tset_arg_end or (semi_pos + 1)
|
||||
end
|
||||
|
||||
@@ -1441,7 +1527,7 @@ local function parse_typedef_binds(source, pos, ident_end, line_of, out)
|
||||
local underlying_span = source:sub(after_typedef, last_ident_pos - 1)
|
||||
local underlying = duffle.trim(underlying_span)
|
||||
register_typedef_alias(underlying, last_ident, pos, line_of, out)
|
||||
associate_skip_over_marker(out, last_ident, last_ident, "unrelated", line_of(pos), pos)
|
||||
attach_debug_skip_marker(out, "unrelated")
|
||||
return last_ident_end
|
||||
end
|
||||
|
||||
@@ -1629,7 +1715,9 @@ local DECL_PARSERS = {
|
||||
MipsAtom_ = parse_mips_atom,
|
||||
MipsAtomComp_ = parse_mips_atom_comp,
|
||||
MipsAtomComp_Proc_ = parse_mips_atom_comp_proc,
|
||||
atom_dbg_skip_over = parse_skip_over_marker,
|
||||
-- `atom_dbg_skip` is the only debug-skip parser entry. Every other
|
||||
-- identifier follows the ordinary unrelated-token path; there is no alias.
|
||||
atom_dbg_skip = parse_dbg_skip_marker,
|
||||
atom_dbg_reg_default = parse_atom_dbg_reg_default,
|
||||
MipsCode = parse_mips_code,
|
||||
typedef = parse_typedef_binds,
|
||||
@@ -1638,6 +1726,9 @@ local DECL_PARSERS = {
|
||||
enum = parse_enum,
|
||||
}
|
||||
|
||||
-- Only the bare `atom_dbg_skip` marker reaches `parse_dbg_skip_marker`.
|
||||
-- Unknown identifiers follow the same unrelated-token path as every other unsupported source token.
|
||||
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
-- The single source walker
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
@@ -1648,33 +1739,31 @@ local DECL_PARSERS = {
|
||||
--- @param source_file string|nil -- absolute source path (forwarded into AliasEntry.source_file)
|
||||
--- @param code_macros table|nil -- cross-source `R_*_Code` registry; nil = local-only
|
||||
--- @param code_macro_bodies table|nil -- cross-source raw RHS body table; nil = local-only
|
||||
--- @return table -- SourceScan { atoms, raw_atoms, binds, atom_infos, macros, skip_over, line_of, register_alias_registry, _code_macros, _code_macro_bodies }
|
||||
--- @return table -- SourceScan { atoms, raw_atoms, binds, atom_infos, macros, debug_skip_markers, line_of, register_alias_registry, _code_macros, _code_macro_bodies }
|
||||
local function scan_source(source, source_file, code_macros, code_macro_bodies)
|
||||
local line_of = duffle.LineIndex(source)
|
||||
local out = {
|
||||
atoms = {},
|
||||
raw_atoms = {},
|
||||
binds = {},
|
||||
atom_infos = {},
|
||||
macros = {},
|
||||
skip_over = {
|
||||
atoms = {},
|
||||
components = {},
|
||||
markers = {},
|
||||
},
|
||||
types = {},
|
||||
atom_views = {},
|
||||
line_of = line_of,
|
||||
atoms = {},
|
||||
raw_atoms = {},
|
||||
binds = {},
|
||||
atom_infos = {},
|
||||
macros = {},
|
||||
-- Raw marker evidence for annotation validation. The `debug_skip` boolean
|
||||
-- is stamped on the declaration record itself; the projection lives on AtomEntry.debug_skip.
|
||||
debug_skip_markers = {},
|
||||
types = {},
|
||||
atom_views = {},
|
||||
line_of = line_of,
|
||||
-- Source-derived register-alias registry (atom_reg opt-in entries).
|
||||
-- Keys are full R_* idents (never stripped); see parse_enum / parse_enum_body.
|
||||
register_alias_registry = {},
|
||||
-- Source-derived type-name registry.
|
||||
-- Populated from `typedef Struct_(...)`, `typedef Enum_(...)`, `typedef <type> <alias>`, and `typedef <type> TSet_(<name>)` declarations.
|
||||
-- The propagation pass at the end of `scan_source()` resolves byte_size via the builtin map,
|
||||
-- Source-derived type-name registry.
|
||||
-- Populated from `typedef Struct_(...)`, `typedef Enum_(...)`, `typedef <type> <alias>`, and `typedef <type> TSet_(<name>)` declarations.
|
||||
-- The propagation pass at the end of `scan_source()` resolves byte_size via the builtin map,
|
||||
-- typedef chain walking (cycle-guarded, depth <= 8), and struct field sums.
|
||||
-- See `propagate_type_sizes()` below.
|
||||
type_name_registry = {},
|
||||
-- Shared `R_*_Code -> integer code` registry
|
||||
-- Shared `R_*_Code -> integer code` registry
|
||||
-- (passed in from M.run pass 1; same reference so preprocessor intercept writes are visible to the enum-value resolver).
|
||||
-- Stripped from `src.scan` before return.
|
||||
_code_macros = code_macros or {},
|
||||
@@ -1708,26 +1797,27 @@ local function scan_source(source, source_file, code_macros, code_macro_bodies)
|
||||
if parser then
|
||||
pos = parser(source, pos, ident_end, line_of, out)
|
||||
else
|
||||
-- A component-procedure declaration has an FI_ signature before MipsAtomComp_Proc_; keep the marker pending across that prelude.
|
||||
-- Any other identifier begins an unrelated declaration/construct and consumes the marker so it cannot drift to a later atom.
|
||||
local markers = out.skip_over.markers
|
||||
-- Unsupported identifiers follow the unrelated-token path. If a
|
||||
-- pending marker is still open, consume it so it cannot drift to a
|
||||
-- later declaration. Unsupported identifiers never create marker records.
|
||||
local markers = out.debug_skip_markers
|
||||
local marker = markers[#markers]
|
||||
if marker and marker.pending then
|
||||
if ident == "FI_" then
|
||||
marker.proc_prelude = true
|
||||
elseif not marker.proc_prelude then
|
||||
associate_skip_over_marker(out, ident, ident, "unrelated", line_of(pos), pos)
|
||||
attach_debug_skip_marker(out, "unrelated")
|
||||
end
|
||||
end
|
||||
pos = ident_end
|
||||
end
|
||||
else
|
||||
local markers = out.skip_over.markers
|
||||
local markers = out.debug_skip_markers
|
||||
local marker = markers[#markers]
|
||||
if marker and marker.pending and marker.proc_prelude then
|
||||
local c = source:sub(pos, pos)
|
||||
if c == "{" or c == ";" then
|
||||
associate_skip_over_marker(out, c, c, "unrelated", line_of(pos), pos)
|
||||
attach_debug_skip_marker(out, "unrelated")
|
||||
end
|
||||
end
|
||||
pos = pos + 1
|
||||
@@ -1748,8 +1838,8 @@ end
|
||||
-- Corpus merge — first-wins lookup identity + typed collisions
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
-- These helpers run ONCE per `M.run` invocation, after every per-source scan has attached `src.scan`.
|
||||
-- They merge per-source scans into the canonical `ctx.shared.corpus.*` registries.
|
||||
-- The corpus is the source of truth; `src.scan` keeps the source-local projection for the duration of the run but the cross-source visibility lives on `corpus`.
|
||||
-- They merge per-source scans into the `ctx.shared.corpus.*` registries.
|
||||
-- `src.scan` keeps the source-local projection for the duration of the run; the cross-source visibility lives on `corpus`.
|
||||
|
||||
-- Build a deterministic site record (path + line) from a per-source entry.
|
||||
-- Falls back to the placeholder when an entry lacks a recorded source file or line.
|
||||
@@ -1858,8 +1948,8 @@ local function phase_shape(entry)
|
||||
end
|
||||
|
||||
-- Merge a new declaration site into a registry following the first-wins discipline.
|
||||
-- * first declaration: Entry becomes the canonical corpus entry (entry.sites initialized).
|
||||
-- * identical subsequent: Append the new site to entry.sites (no collision).
|
||||
-- * first declaration: Entry becomes the corpus entry (entry.sites initialized).
|
||||
-- * identical subsequent: Append the new site to entry.sites.
|
||||
-- * conflicting shape: Keep first entry, append ONE typed collision record with shape diff.
|
||||
local function merge_named_with_sites(registry, name, new_entry, site, collisions, kind, shape_fn)
|
||||
if registry[name] == nil then
|
||||
@@ -1888,8 +1978,8 @@ local function merge_named_with_sites(registry, name, new_entry, site, collision
|
||||
}
|
||||
end
|
||||
|
||||
-- Merge per-source scans into the canonical corpus registries.
|
||||
-- Iterates `corpus.source_order` (not `ctx.sources`) — the corpus is the source of truth.
|
||||
-- Merge per-source scans into the corpus registries.
|
||||
-- Iterates `corpus.source_order` (not `ctx.sources`).
|
||||
-- Each source owns only its `src.scan`; the corpus owns the cross-source lookup tables.
|
||||
local function merge_corpus_registries(corpus)
|
||||
-- Ensure every expected corpus table exists (the fixture_ctx seeds most of these,
|
||||
@@ -2002,8 +2092,7 @@ local M = {}
|
||||
--- No output files; this is a pure in-memory pre-processing pass.
|
||||
---
|
||||
--- Runs in 5 phases.
|
||||
--- Resolve: Resolve the canonical source order from `ctx.shared.corpus.source_order`.
|
||||
--- The canonical corpus is the SOLE source of truth; no `ctx.sources` alias is consulted and no per-source fallback synthesis is performed.
|
||||
--- Resolve: Source order from `ctx.shared.corpus.source_order` (the corpus owns it; the check below enforces the invariant).
|
||||
--- Pass 1a: `scan_source_pre_pass` over every source, populating LOCAL `code_macros` AND LOCAL `code_macro_bodies` tables.
|
||||
--- The bodies table holds the raw post-`=` text of every `#define R_*_Code` line (cross-source)
|
||||
--- so the chain walker can fall back when the defining `#define` lives in a different source than the chain call site.
|
||||
@@ -2012,8 +2101,8 @@ local M = {}
|
||||
--- Pass 2: The full `scan_source(source, source_file, code_macros, code_macro_bodies)` walk per source. The per-source `src.scan` payload includes the source-local registries
|
||||
--- (register_alias_registry, type_name_registry, atom_views, atom_ctxs, atom_phases, binds, atoms, atom_infos, ...).
|
||||
--- Strip: Strip `src.scan._code_macros`, `src.scan._code_macro_bodies`, and the `_source_file` pointer.
|
||||
--- The LOCAL tables `code_macros` and `code_macro_bodies` go out of scope here; they MUST NOT appear on `ctx.shared`, `ctx.shared.corpus`, or any `src.scan` after this point.
|
||||
--- Merge: Iterate `ctx.shared.corpus.source_order` in declared order. For every source's local registry, first-wins lookup identity (entry from the first declaration site becomes the canonical corpus entry);
|
||||
--- The LOCAL tables `code_macros` and `code_macro_bodies` stay confined to this function; they go out of scope on return.
|
||||
--- Merge: Iterate `ctx.shared.corpus.source_order` in declared order. For every source's local registry, first-wins lookup identity (entry from the first declaration site becomes the corpus entry);
|
||||
--- identical shapes coalesce by appending the declaration site; conflicting shapes keep the first lookup entry and append ONE typed collision record with shape diff.
|
||||
--- Populate `register_alias_registry`, `type_name_registry`, `binds_by_name`, `atoms_by_name`, `atom_views`, `atom_ctxs`, `atom_phases`.
|
||||
--- `atom_infos` ALWAYS appends every record (preserving source order + duplicates for annotation evidence).
|
||||
@@ -2022,14 +2111,12 @@ local M = {}
|
||||
--- @return PassResult
|
||||
function M.run(ctx)
|
||||
-- The cross-source _code_macros / _code_macro_bodies tables are LOCAL to this run.
|
||||
-- They are shared across source scans ONLY long enough to resolve cross-source R_*_Code chains, then DISCARDED.
|
||||
-- They MUST NOT appear on ctx.shared, ctx.shared.corpus, or any src.scan after
|
||||
-- this function returns.
|
||||
-- They live across source scans only long enough to resolve cross-source R_*_Code chains, then go out of scope on M.run return.
|
||||
-- The Lua GC reclaims them; nothing here survives onto ctx.shared, ctx.shared.corpus, or any src.scan.
|
||||
local code_macros = {}
|
||||
local code_macro_bodies = {}
|
||||
|
||||
-- Resolve the canonical source list. The corpus owns the authoritative source_order.
|
||||
-- A context without `ctx.shared.corpus` is rejected with an explicit canonical-corpus error.
|
||||
-- Canonical-corpus check (see the docstring Resolve phase). The corpus is the only source of source_order.
|
||||
ctx.shared = ctx.shared or {}
|
||||
local corpus = ctx.shared.corpus
|
||||
if not corpus or type(corpus.source_order) ~= "table" then
|
||||
@@ -2069,20 +2156,16 @@ function M.run(ctx)
|
||||
src.scan._code_macro_bodies = nil
|
||||
src.scan._source_file = nil
|
||||
end
|
||||
-- Pre-tokenize each atom body once (plex: single source of truth).
|
||||
-- Downstream passes (offsets, word-counts, components, static-analysis) read from `atom.body_tokens` instead of calling `split_top_level_commas` / `tokenize_body` independently.
|
||||
-- The tokens are memoized in duffle.lua's cache, so re-access is O(1).
|
||||
-- Pre-tokenize each atom body once (plex: cache lives in duffle.lua; downstream passes read from `atom.body_tokens` instead of calling `split_top_level_commas` / `tokenize_body` independently).
|
||||
-- Re-access is O(1) thanks to the memoization.
|
||||
for _, atom in ipairs(src.scan.atoms) do atom.body_tokens = duffle.tokenize_body(atom.body) end
|
||||
for _, atom in ipairs(src.scan.raw_atoms or {}) do atom.body_tokens = duffle.tokenize_body(atom.body) end
|
||||
end
|
||||
|
||||
-- Merge per-source scans into the canonical corpus registries.
|
||||
-- First-wins lookup identity + collision discipline (see merge_corpus_registries).
|
||||
-- The corpus is always present; no conditional / fallback path.
|
||||
-- Merge per-source scans into the corpus registries (see merge_corpus_registries for first-wins + collision discipline).
|
||||
merge_corpus_registries(corpus)
|
||||
|
||||
-- code_macros and code_macro_bodies go out of scope here; their references are not captured on corpus, ctx.shared, or any src.scan.
|
||||
-- The Lua GC reclaims them on M.run return.
|
||||
-- code_macros and code_macro_bodies are function-local; the GC reclaims them on M.run return.
|
||||
return { outputs = {}, errors = {}, warnings = {} }
|
||||
end
|
||||
|
||||
|
||||
@@ -1,12 +1,16 @@
|
||||
--- passes/static_analysis.lua — Per-atom static-analysis checks.
|
||||
---
|
||||
--- Ownership: `ctx.shared.corpus` is the canonical merged registry; per-source fallback synthesis is rejected.
|
||||
--- `atom.paths` supplies the emitted and analysis projections consumed by this pass.
|
||||
---
|
||||
--- Per-atom rules:
|
||||
--- 1. transfer_hazards: A single forward walker (`analyze_hardware_relations`) reads `atom.paths.word_events` once per atom.
|
||||
--- For each emitted word event it (a) inspects pending CPU/COP0/COP2/GTE relations against the event as CONSUMER
|
||||
--- (recording a hazard on `atom.paths.hazards` when the producer→consumer gap is below the required retire-slot count),
|
||||
--- (b) applies the event's GPR value effects (`duffle.INSTRUCTION_GPR_EFFECTS`) to `atom.paths.forward_state.gpr_values`,
|
||||
--- applies bounded constant propagation, and stages matching relation rows as PRODUCERS (with `destination_match` filters, e.g. for the IRGB fan-out).
|
||||
--- The `transfer_hazards` CHECK_RULES reader projects `atom.paths.hazards` into per-atom findings without re-walking source.
|
||||
--- The `transfer_hazards` CHECK_RULES reader projects `atom.paths.hazards` into per-atom findings.
|
||||
--- The reader does NOT re-walk source; this is the per-check purity contract.
|
||||
--- The walker runs once per atom before the per-atom dispatch; the reader runs inside the same dispatch.
|
||||
--- 2. control_transfer_delay_slot_use: For every emitted branch/jump/call encoder in `duffle.CONTROL_TRANSFER_DELAY_SLOT_POLICIES`
|
||||
--- (the six `branch_*` encoders plus `jump` / `jump_reg` / `jump_link` / `call_reg` / `call_addr`),
|
||||
@@ -26,7 +30,7 @@
|
||||
--- must be in `corpus.register_alias_registry`.
|
||||
--- 9. atom_type_consistency: Every `reg_type_overrides[R_X].type_name` must resolve in `corpus.type_name_registry`.
|
||||
--- 10. binds_no_substruct_deref: Every `load_word(R_A, R_B, O_(Type, Field))` and `store_word(...)` in every atom body must reference a leaf scalar
|
||||
--- (pointer-to-struct counts as leaf; nested struct members do NOT).
|
||||
--- (pointer-to-struct counts as leaf; nested struct members fail the leaf test).
|
||||
---
|
||||
---
|
||||
--- Findings carry an explicit `kind` ("error" / "warning" / "info").
|
||||
@@ -297,7 +301,7 @@ end
|
||||
-- Only the matching row stages (the non-matching row is ignored for that event).
|
||||
--
|
||||
-- After the walker runs, the `transfer_hazards` CHECK_RULES reader (`check_transfer_hazards`) copies every entry on `atom.paths.hazards` into the per-atom findings list.
|
||||
-- The reader does NOT re-walk source or re-classify tokens; it is a pure projection of the walker's output.
|
||||
-- The first `transfer_hazards` reader comment above records the projection contract.
|
||||
--
|
||||
-- The walker is called once before the CHECK_RULES per-atom dispatch (see `validate()`);
|
||||
-- The reader runs as part of the same CHECK_RULES dispatch so its findings land in `findings` alongside the other checks.
|
||||
@@ -319,7 +323,7 @@ local function is_cop2_consumer_of(consumer_event, destination, producer_rel)
|
||||
for _, pos in ipairs(args) do
|
||||
if pos == destination then return true end
|
||||
end
|
||||
-- Match via the command's input set: the consumer's encoder resolves to a canonical `gte_cmdw_*`
|
||||
-- Match via the command's input set: the consumer encoder resolves to a `gte_cmdw_*`
|
||||
-- short form whose `duffle.GTE_COMMAND_INPUTS` entry includes the destination (or a fan-out target).
|
||||
local aliases = duffle.GTE_COMMAND_ALIASES or {}
|
||||
local canonical = aliases[consumer_token] or consumer_token
|
||||
@@ -528,7 +532,7 @@ local function apply_gpr_effects(ev_ident, ev_args, forward_state)
|
||||
end
|
||||
end
|
||||
|
||||
-- Look up the canonical alias of a GTE command ident.
|
||||
-- Look up the alias of a GTE command ident.
|
||||
-- Defaults to the input ident so unknown idents surface rather than silently inheriting a 0-cycle command input set.
|
||||
local function canonical_command(ident)
|
||||
local aliases = duffle.GTE_COMMAND_ALIASES or {}
|
||||
@@ -932,7 +936,7 @@ end
|
||||
--
|
||||
-- The single forward walker `analyze_hardware_relations` (defined above) has already populated `atom.paths.hazards`.
|
||||
-- This check copies every entry on that list into the per-atom `findings` table.
|
||||
-- It does NOT re-walk source / re-classify tokens; it is a pure projection of the walker's output.
|
||||
-- The first `transfer_hazards` reader comment above records the projection contract.
|
||||
--
|
||||
-- The walker also populates `atom.paths.relations` (one entry per satisfied-or-violated relation touch) and `atom.paths.forward_state` (the GPR-value lattice).
|
||||
-- Neither of those is rendered as a finding here; bounded-value rules and LWC2 unknown edges share on top of the same forward walker and adds additional readers.
|
||||
@@ -954,7 +958,7 @@ end
|
||||
--
|
||||
-- The forward walker stages post-command latch relations on `atom.paths.hazards` with `relation_id = "command_latch_input"`.
|
||||
-- This reader filters those entries and re-emits them under the `gte_input_latch` check name so the test contract can target them independently of the transfer_hazards check.
|
||||
-- The reader does NOT re-walk source tokens or build its own pending state; it is a pure projection of the walker's output.
|
||||
-- The first `transfer_hazards` reader comment above records the projection contract.
|
||||
-- ─────────────────────────────────────────────────────────────────────────
|
||||
|
||||
local function check_gte_input_latch(atom, _pipe_ctx, findings)
|
||||
@@ -984,7 +988,7 @@ end
|
||||
-- A subsequent MFC2 (or any encoder that reads a C2 register) that picks the WRONG register for the active role emits a `result_role_mismatch` warning.
|
||||
-- For example, reading `C2_SXY0` after RTPS is wrong: the `latest_screen_xy` role is `C2_SXY2`.
|
||||
--
|
||||
-- The reader does NOT re-walk source tokens; it consumes `forward_state.post_command_roles` and `atom.paths.word_events` only.
|
||||
-- The first `transfer_hazards` reader comment above records the projection contract.
|
||||
-- ─────────────────────────────────────────────────────────────────────────
|
||||
|
||||
local function check_gte_result_position(atom, _pipe_ctx, findings)
|
||||
@@ -1029,7 +1033,7 @@ local function check_gte_result_position(atom, _pipe_ctx, findings)
|
||||
local args = ev.args or {}
|
||||
local reg = args[2]
|
||||
-- Find any post-command `latest_screen_xy` role entry recorded by a prior command.
|
||||
-- The newest projected screen coordinate is recorded under the command's canonical name.
|
||||
-- The newest projected screen coordinate is recorded under the command name.
|
||||
-- Reading from C2_SXY0 (the older projection slot) when a `latest_screen_xy` role was set to C2_SXY2 by RTPS / RTPT is a semantic mismatch.
|
||||
local latest_screen_xy_entry = nil
|
||||
for r, e in pairs(forward.post_command_roles or {}) do
|
||||
@@ -1077,10 +1081,10 @@ end
|
||||
-- (the nop is needed to retire the relation, even if it can be replaced by independent useful work).
|
||||
-- * `modeled-redundant`: no modeled relation is pending immediately before the nop (the nop is a redundant hazard).
|
||||
--
|
||||
-- Branch/jump delay-slot NOPs are NOT classified by this check (they are exclusively owned by `control_transfer_delay_slot_use`).
|
||||
-- Branch/jump delay-slot NOPs belong to `control_transfer_delay_slot_use`, so this check leaves them unclassified.
|
||||
-- The fixed `mac_yield()` handshake (`jump_reg(R_AtomJmp), nop`) is preserved as suppressed.
|
||||
--
|
||||
-- The reader does NOT re-walk source tokens; it consumes `forward_state.pending` snapshots and `atom.paths.word_events`.
|
||||
-- The first `transfer_hazards` reader comment above records the projection contract.
|
||||
-- ─────────────────────────────────────────────────────────────────────────
|
||||
|
||||
local function check_hazard_nop_use(atom, _pipe_ctx, findings)
|
||||
@@ -1238,10 +1242,8 @@ end
|
||||
-- ─────────────────────────────────────────────────────────────────────────
|
||||
-- Check #1c: control-transfer delay-slot use.
|
||||
--
|
||||
-- Reads `atom.paths.word_events` (the semantic emitted-word stream from
|
||||
-- `passes/emission_model.lua`). For each event whose `encoder` is in
|
||||
-- `duffle.CONTROL_TRANSFER_DELAY_SLOT_POLICIES`, inspect the next emitted
|
||||
-- event in the SAME `events` array.
|
||||
-- Reads `atom.paths.word_events` (the semantic emitted-word stream from `passes/emission_model.lua`).
|
||||
-- For each event whose `encoder` is in `duffle.CONTROL_TRANSFER_DELAY_SLOT_POLICIES`, inspect the next emitted event in the SAME `events` array.
|
||||
-- The next event is the hardware delay-slot word (the duffle pipeline already absorbs the BD-slot into the branch's cost in `analyze_atom_paths`.
|
||||
-- This check observes, it does not reschedule.
|
||||
--
|
||||
@@ -1252,7 +1254,7 @@ end
|
||||
-- Suppress the finding when `policy.suppress_arg1[first_arg]` is non-nil.
|
||||
-- The only current suppression is `jump_reg(R_AtomJmp)`, the fixed `mac_yield()` handshake.
|
||||
--
|
||||
-- `pipe_ctx` is unused; the uniform `(atom, pipe_ctx, findings)` signature is preserved so the check plugs into
|
||||
-- `pipe_ctx` is unused; the uniform `(atom, pipe_ctx, findings)` signature is preserved so the check plugs into
|
||||
-- the existing CHECK_RULES dispatch without modifying the per-atom loop or analyze_atom_paths.
|
||||
-- `passes/emission_model` already normalizes `nop2` to two `nop` events and `atom_label` to zero events, so no special-case branching is needed for either.
|
||||
-- ─────────────────────────────────────────────────────────────────────────
|
||||
@@ -1402,8 +1404,8 @@ end
|
||||
--- 2. Body MUST contain an `add_ui_self(R_TapePtr, S_(Binds_X))` (or equivalent advance by the struct's byte count). Missing = error.
|
||||
--- 3. atom_bind(Binds_X) where Binds_X doesn't exist = error.
|
||||
--- Per-atom: Verify the atom body reads every field of its `Binds_X` from R_TapePtr and advances R_TapePtr by S_(Binds_X).
|
||||
--- Takes `(atom, pipe_ctx, findings)`; `pipe_ctx` carries the cross-atom
|
||||
--- `info_by_atom` + `binds_index` tables (built once by validate() before the per-atom loop).
|
||||
--- Takes `(atom, pipe_ctx, findings)`; `pipe_ctx` carries the cross-atom `info_by_atom` + `binds_index` tables
|
||||
--- (built once by validate() before the per-atom loop).
|
||||
--- `validate()` owns per-atom iteration; this function evaluates one atom.
|
||||
local function check_abi_handoff(atom, pipe_ctx, findings)
|
||||
local info = pipe_ctx.info_by_atom[atom.name]
|
||||
@@ -1658,10 +1660,9 @@ local function analyze_atom_paths(atom)
|
||||
|
||||
local succ, term = successors(tok_idx)
|
||||
if term then
|
||||
-- Terminator: record the path's cycle sum.
|
||||
-- We do NOT add the terminator token to `visited` a path ends here, so a different path that
|
||||
-- ALSO reaches this terminator is a legitimate new path (not a loop).
|
||||
-- If we marked it visited, subsequent paths that reach the same terminator would be incorrectly flagged as loops.
|
||||
-- Terminator: record the path's cycle sum.
|
||||
-- The terminator token stays out of `visited`, so another path reaching the same terminator remains a distinct path.
|
||||
-- Marking it visited would flag those legitimate paths as loops.
|
||||
path_count = path_count + 1
|
||||
if new_acc < cycles_min then cycles_min = new_acc end
|
||||
if new_acc > cycles_max then cycles_max = new_acc end
|
||||
@@ -1704,7 +1705,7 @@ end
|
||||
--- (deduplicated across atoms so the warning section doesn't get spammed with N copies of "macro X not in duffle.INSTRUCTION_LATENCY").
|
||||
--- Per-atom: emit one finding per unknown macro seen, deduplicated across atoms
|
||||
--- (so the warning section doesn't get spammed with N copies of "macro X not in duffle.INSTRUCTION_LATENCY").
|
||||
--- Reuses `analyze_atom_paths`'s per-atom unknown_macros discovery (it's the canonical place that walks tokens and computes per-token cycle costs).
|
||||
--- Reuses `analyze_atom_paths`'s per-atom unknown_macros discovery, which walks tokens and computes per-token cycle costs.
|
||||
local function check_per_atom_cycle_budget(atom, pipe_ctx, findings)
|
||||
local p = atom.paths or {}
|
||||
for _, name in ipairs(p.unknown_macros or {}) do
|
||||
@@ -1736,9 +1737,9 @@ end
|
||||
-- The rule is intentionally permissive because the production `code/duffle/` and `code/gte_hello/`
|
||||
-- sources use R_* aliases in atom_reads / atom_writes that may not yet be opted in via the bare `atom_reg` marker.
|
||||
-- R_TapePtr / R_AtomJmp / R_PrimCursor / R_FaceCursor / R_VertBase / R_OtBase ARE opted in.
|
||||
-- Raw C-ABI aliases like R_T0..R_T3 are intentionally NOT auto-included (per the prototype principle:
|
||||
-- no auto-include of wave-context; explicit opt-in only). Warnings keep the build green
|
||||
-- and report aliases that need explicit registration.
|
||||
-- Raw C-ABI aliases like R_T0..R_T3 require explicit opt-in; the prototype keeps wave-context registration explicit.
|
||||
-- no auto-include of wave-context; explicit opt-in only).
|
||||
-- Warnings keep the build green and report aliases that need explicit registration.
|
||||
local function check_enum_alias_membership(_src, pipe_ctx, findings)
|
||||
local reg_registry = pipe_ctx.register_alias_registry or {}
|
||||
|
||||
@@ -1836,7 +1837,7 @@ end
|
||||
-- the `<Field>` MUST resolve to a leaf scalar of `<Type>`. A "leaf scalar" is:
|
||||
-- * a non-struct field with `pointer_depth >= 1` (pointer-to-struct IS a leaf — the field is a pointer; the pointee is unrelated), OR
|
||||
-- * a non-struct field whose type_name resolves to a typedef / enum / builtin in `type_name_registry`.
|
||||
-- A nested struct member (pointer_depth == 0 AND type_name resolves to a `kind = "struct"` registry entry) is NOT a leaf scalar and is flagged.
|
||||
-- A nested struct member (pointer_depth == 0 and type_name resolves to a `kind = "struct"` registry entry) fails the leaf-scalar test.
|
||||
-- The check also flags fields whose Type has no `fields` table (typedefs and enums don't have fields — any Field reference against them is bogus)
|
||||
-- and fields whose name doesn't appear in the resolved Type's fields array.
|
||||
--
|
||||
@@ -1858,7 +1859,7 @@ local function find_field_by_name(type_entry, field_name)
|
||||
end
|
||||
|
||||
-- True iff a (field, type_registry) pair is a leaf scalar (safe to dereference as a tape-payload field).
|
||||
-- Pointer-to-X is always leaf; non-pointer struct members are NOT leaf.
|
||||
-- Pointer-to-X is always a leaf; non-pointer struct members fail the leaf test.
|
||||
local function is_field_leaf(field, type_registry)
|
||||
if field.pointer_depth and field.pointer_depth > 0 then
|
||||
return true
|
||||
@@ -1925,7 +1926,7 @@ end
|
||||
-- per_atom(atom, pipe_ctx, findings) — runs once per atom inside validate()'s single loop
|
||||
-- post(pipe_ctx, findings) — runs once after all per-atom calls complete
|
||||
-- per_macro(macro, wc, findings) — runs once per TAPE_WORDS / _Pragma macro declaration
|
||||
-- per_skip_marker(marker, pipe_ctx, findings) — runs once per src.scan.skip_over.markers entry
|
||||
-- per_skip_marker(marker, pipe_ctx, findings) — runs once per src.scan.debug_skip_markers entry
|
||||
-- per_source(src, pipe_ctx, findings) — runs once per source AFTER the per-atom loop completes
|
||||
-- (registry-driven rule; same CHECK_RULES table)
|
||||
-- Each check is one table row and one `check_*` function.
|
||||
@@ -1951,11 +1952,11 @@ local CHECK_RULES = {
|
||||
-- ════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
--- Build the corpus-wide pipe_ctx ONCE per pass run.
|
||||
--- Reads the merged `corpus.*` registries (canonical cross-source lookups), and the corpus-wide `atom_infos` list (preserving source order + duplicates).
|
||||
--- The corpus is the source of truth; per-source scans retain body / declaration ownership via `src.scan` and the per-source `atoms` / `atom_infos` projections.
|
||||
--- Reads the merged `corpus.*` registries and the corpus-wide `atom_infos` list (preserving source order + duplicates).
|
||||
--- The corpus supplies shared registries; `src.scan` and per-source projections retain body and declaration ownership.
|
||||
---
|
||||
--- Ownership: A context without `ctx.shared.corpus` is rejected with an explicit canonical-corpus message.
|
||||
--- No per-source fallback synthesis is performed; callers MUST construct a canonical ctx through `build_ctx`.
|
||||
--- A context without `ctx.shared.corpus` is rejected with an explicit corpus message.
|
||||
--- Callers construct the context through `build_ctx`.
|
||||
--- @param ctx PassCtx
|
||||
--- @return PipeCtx
|
||||
local function build_corpus_pipe_ctx(ctx)
|
||||
@@ -1966,9 +1967,9 @@ local function build_corpus_pipe_ctx(ctx)
|
||||
.. "no per-source fallback is supported)", 0)
|
||||
end
|
||||
-- The pipe_ctx views REFERENCE the corpus tables directly (no copies).
|
||||
-- Every consumer of these fields observes mutations via the canonical corpus without independently mutable registry construction.
|
||||
-- Every consumer observes mutations through the corpus tables directly.
|
||||
return {
|
||||
-- Cross-source lookup tables (canonical corpus projections).
|
||||
-- Cross-source lookup tables.
|
||||
register_alias_registry = corpus.register_alias_registry or {},
|
||||
type_name_registry = corpus.type_name_registry or {},
|
||||
atom_views = corpus.atom_views or {},
|
||||
@@ -1985,8 +1986,8 @@ end
|
||||
|
||||
local function validate(ctx, src, corpus_pipe_ctx)
|
||||
local scan = src.scan
|
||||
-- Read the canonical corpus word_counts for the per-atom pipeline
|
||||
-- (atom.paths.word_events is the canonical projection).
|
||||
-- Read the corpus word_counts for the per-atom pipeline
|
||||
-- (`atom.paths.word_events` is the emitted projection).
|
||||
|
||||
local corpus = (ctx.shared and ctx.shared.corpus) or {}
|
||||
|
||||
@@ -2028,7 +2029,7 @@ local function validate(ctx, src, corpus_pipe_ctx)
|
||||
register_alias_registry = corpus_pipe_ctx.register_alias_registry,
|
||||
type_name_registry = corpus_pipe_ctx.type_name_registry,
|
||||
}
|
||||
-- Shared cross-source component-body index is owned by the canonical corpus
|
||||
-- Shared cross-source component-body index is owned by the corpus
|
||||
-- (`corpus.component_body_index`, populated by `passes/components.lua`).
|
||||
-- Per-atom checks consume the corpus-owned index directly.
|
||||
pipe_ctx.component_body_index = (corpus and corpus.component_body_index) or {}
|
||||
@@ -2041,13 +2042,11 @@ local function validate(ctx, src, corpus_pipe_ctx)
|
||||
---
|
||||
--- Body, token, and emission projections come from here (`paths.tokens = body_tokens`, `paths.line_in_body = build_body_line_index` `paths.word_events`
|
||||
--- and related fields are owned by `passes/emission_model.lua` pass (per-atom emission projection).
|
||||
--- This pass reads: `paths.tokens`, `paths.line_in_body` ` paths.items`, `paths.word_events` from the canonical projection,
|
||||
--- then computes `paths.tok_class`, `paths.cycles_min/max`, `paths.branches`, `paths.paths`, `paths.has_loops`, `paths.unknown_macros`
|
||||
--- via `classify_tokens` + `analyze_atom_paths`.
|
||||
--- This pass reads: `paths.tokens`, `paths.line_in_body`, `paths.items`, `paths.word_events` from the emitted projection,
|
||||
--- then computes `paths.tok_class`, `paths.cycles_min/max`, `paths.branches`, `paths.paths`, `paths.has_loops`, `paths.unknown_macros` via `classify_tokens` + `analyze_atom_paths`.
|
||||
--- No re-walk of body text or body_tokens happens here.
|
||||
---
|
||||
--- Canonical contract: `atom.paths` and `atom.paths.word_events` MUST be
|
||||
--- populated by `passes/emission_model.run(ctx)` before this pass runs.
|
||||
--- Canonical contract: `atom.paths` and `atom.paths.word_events` MUST be populated by `passes/emission_model.run(ctx)` before this pass runs.
|
||||
--- The `atom.paths.word_events` projection is owned by the emission-model pass; static-analysis reads it directly.
|
||||
local findings = {}
|
||||
for _, a in ipairs(atoms) do
|
||||
@@ -2354,7 +2353,7 @@ local function emit_module_static_analysis_txt(ctx, dir, dir_sources, atoms, fin
|
||||
end
|
||||
|
||||
-- Module-level findings summary (across all sources).
|
||||
-- Info is its own count; it is NOT lumped into warnings.
|
||||
-- Info has its own count; it remains separate from warnings.
|
||||
local total_errs = #errors
|
||||
local total_warns = #warnings
|
||||
local total_infos = #info
|
||||
@@ -2393,13 +2392,11 @@ function M.run(ctx)
|
||||
local warnings = {}
|
||||
-- `info` aggregates finding-level info across every source (the per-source validate() also
|
||||
-- returns a `summaries` collection for scan/cycle rollups;
|
||||
-- those are NOT finding-level and never enter `info`).
|
||||
-- those are summary rows and never enter `info`).
|
||||
local info = {}
|
||||
|
||||
-- Build the corpus-wide pipe_ctx ONCE per pass run.
|
||||
-- The corpus owns the canonical cross-source registries; per-source scans
|
||||
-- retain body / declaration ownership. The pipe_ctx is shared across every
|
||||
-- validate() invocation in this M.run so cross-source visibility is constant.
|
||||
-- The pipe_ctx is shared across every validate() invocation in this M.run so cross-source visibility is constant.
|
||||
local corpus_pipe_ctx = build_corpus_pipe_ctx(ctx)
|
||||
local corpus = ctx.shared.corpus
|
||||
|
||||
|
||||
@@ -8,9 +8,9 @@
|
||||
--- AFTER computing each current count from the just-built body + `corpus.word_counts`).
|
||||
---
|
||||
--- **Canonical contract**:
|
||||
--- * `ctx.shared.corpus.word_counts` is the canonical count table.
|
||||
--- * `ctx.shared.corpus.word_counts` is the count table.
|
||||
--- * `corpus.word_counts` is the sole count table. Consumers read `corpus.word_counts` directly.
|
||||
--- * `ctx.shared.components` and `ctx.shared.component_body_index` are NOT created by this pass (canonical projections only).
|
||||
--- * `ctx.shared.components` and `ctx.shared.component_body_index` are NOT created by this pass (projections only).
|
||||
--- * No `.macs.h` recursive discovery (no `scan_dir`, no scan cache, no `_invalidate_scan_cache`).
|
||||
---
|
||||
--- **Conventions**: tabs (1/level), EmmyLua annotations, no regex,
|
||||
@@ -116,10 +116,10 @@ function M.run(ctx)
|
||||
end
|
||||
|
||||
-- 3. Load authored metadata. Generated .macs.h files are NOT scanned
|
||||
-- (the canonical pass computes their counts from the just-built bodies after disk emission; see passes/components.lua).
|
||||
-- (the pass computes their counts from the just-built bodies after disk emission; see passes/components.lua).
|
||||
local wc = duffle.load_word_counts(ctx.metadata_path)
|
||||
|
||||
-- 4. Assign the canonical count table. ONE assignment, no copy. The assignment creates no secondary alias.
|
||||
-- 4. Assign the count table. ONE assignment, no copy. The assignment creates no secondary alias.
|
||||
corpus.word_counts = wc
|
||||
|
||||
return { outputs = {}, errors = {}, warnings = {} }
|
||||
|
||||
+9
-16
@@ -265,14 +265,11 @@ USAGE:
|
||||
|
||||
PASS_FLAGS:
|
||||
Pick a phase or one-or-more individual passes:
|
||||
--pre-link [phase; default] Run the pre-link group + transitive deps.
|
||||
The root set is data-driven from each PASSES row's
|
||||
`groups` field; no parallel name list is maintained.
|
||||
--post-link [phase] Run the post-link group + transitive deps.
|
||||
Requires --elf. Sets --gdb-runtime and --dwarf-injection
|
||||
opt-in flags as well.
|
||||
--all Select every row of the PASSES table. Pass-local opt-in
|
||||
guards remain active, so --dwarf-injection still requires
|
||||
--pre-link [phase; default] Run the pre-link group + transitive deps.
|
||||
The root set is data-driven from each PASSES row's groups` field; no parallel name list is maintained.
|
||||
--post-link [phase] Run the post-link group + transitive deps.
|
||||
Requires --elf. Sets --gdb-runtime and --dwarf-injection opt-in flags as well.
|
||||
--all Select every row of the PASSES table. Pass-local opt-in guards remain active, so --dwarf-injection still requires
|
||||
--elf and --gdb-runtime still requires a runtime emission.
|
||||
Or pick any subset:
|
||||
--scan-source Scan sources into the fat SourceScan payload
|
||||
@@ -281,20 +278,16 @@ PASS_FLAGS:
|
||||
--validate Run atom annotation DSL validation
|
||||
--offsets Generate <module>/gen/<basename>.offsets.h
|
||||
--atoms-source-map Generate <basename>.atoms.sourcemap.txt per source
|
||||
--dwarf-injection [opt-in] Select the post-link dwarf-injection pass + set the
|
||||
opt-in flag. Requires --elf.
|
||||
--dwarf-injection [opt-in] Select the post-link dwarf-injection pass + set the opt-in flag. Requires --elf.
|
||||
--static-analysis Static analysis: GTE pipeline-fill, mac_yield, ABI handoff, cycle budget
|
||||
--report Render per-project summary
|
||||
|
||||
COMMON_FLAGS:
|
||||
--unity-root FILE Unity source root: load root + direct quoted authored
|
||||
includes only. Mutually exclusive with --source.
|
||||
--source FILE Exact source file to process (repeatable, never expands
|
||||
includes). Mutually exclusive with --unity-root.
|
||||
--unity-root FILE Unity source root: load root + direct quoted authored includes only. Mutually exclusive with --source.
|
||||
--source FILE Exact source file to process (repeatable, never expands includes). Mutually exclusive with --unity-root.
|
||||
--metadata PATH Path to metadata.h (required)
|
||||
--out-root DIR Output root for reports (default: build/gen)
|
||||
--project-root DIR PS1 repository root (default: derived from
|
||||
<repo>/code/duffle/word_count.metadata.h)
|
||||
--project-root DIR PS1 repository root (default: derived from <repo>/code/duffle/word_count.metadata.h)
|
||||
--gdb-runtime Also emit <out_root>/gdb_tape_atoms_runtime.gdb (post-link, requires --elf)
|
||||
--elf PATH Path to linked .elf (for --gdb-runtime / --dwarf-injection)
|
||||
--verbose Print per-pass debug output
|
||||
|
||||
Reference in New Issue
Block a user