Commit Graph
100 Commits
Author SHA1 Message Date
ed b193df100f feat(directives): scavenge sweep 3/5 (docs/reports/2026-06-08): 6 directives
Lifted:
- test_instantiation_not_mock_away: tests must exercise actual instantiation,
  not mock away the constructor (caught the MiniMax 401 regression)
- preserve_prior_versions_of_review_docs: when iterating on a review,
  preserve prior versions as separate files (v2/v2.1/v2.2/v2.3)
- neutral_language_for_doc_drift: doc-drift fixes use 'predates/stale/
  outdated', not 'fictional' (value judgment, not technical description)
- preserve_before_compact_archive: at 80%+ context, write a comprehensive
  session-synthesis archive before compaction
- user_corrections_log_in_state_toml: per-track state.toml has a
  user_corrections_log section for any reviewed-by-user track
- surface_dirty_state_in_test_runner: when subprocess is dead/degraded,
  print a clear [BATCH-WARN] rather than silent timeout
2026-07-04 01:05:54 -04:00
ed a89d0cb30e feat(directives): scavenge sweep 2/5 (docs/reports/2026-05-11 + 2026-06-01): 2 directives
Lifted:
- profile_first_optimize_second: profile and measure the actual bottleneck
  before any architectural change (RAG init, not AI SDKs, was the bottleneck)
- surface_gaps_at_discovery_not_checkpoint: surface scope gaps and
  architectural deviations the moment they are discovered, not at a
  checkpoint (the 'all good!' footnote pattern is bad UX)
2026-07-04 01:05:07 -04:00
ed f130cc6813 feat(directives): scavenge sweep 1/5 (docs/reports/2026-03-02 through 2026-06-08): 1 directive
Lifted from docs/reports/2026-03-02/MCP_BUGFIX_20260306.md:
- pathlib_read_write_no_newline_kwarg: pathlib read_text/write_text must omit the
  newline kwarg (unsupported pre-3.10; corrupts line endings on Windows)
2026-07-04 01:04:44 -04:00
ed 6664086968 conductor(state): record scavenge pass phase_5 + s_1..s_6 tasks (15 new directives, 81 total) 2026-07-04 00:20:29 -04:00
ed b2ebe25d71 test(directives): add scavenge_lift contract tests (15 directives, 79 parametrized cases) 2026-07-04 00:18:51 -04:00
ed 9656bf2e88 feat(directives): update current_baseline preset with 15 scavenge directives (81 total) 2026-07-04 00:17:12 -04:00
ed 883f7ec5c1 feat(directives): scavenge from intent_dsl_survey + handoffs/: 5 directives 2026-07-04 00:15:27 -04:00
ed ebca201d39 feat(directives): scavenge from nagent_review_20260608/: 5 directives 2026-07-04 00:13:24 -04:00
ed bea5d6b151 feat(directives): scavenge from docs/MMA_Support/: 5 directives 2026-07-04 00:11:15 -04:00
ed 3155518305 test(aggregate_directives): assert clean bodies + pollution-fix + max_chars cap + MCP impl
8 new importable-function tests in tests/test_aggregate_directives.py:
- test_importable_function_returns_clean_body
- test_importable_function_no_meta_md_pollution
- test_importable_function_max_chars_truncates_with_suffix
- test_importable_function_max_chars_zero_is_unlimited
- test_importable_function_max_chars_above_size_is_unlimited
- test_importable_function_missing_preset_raises
- test_importable_function_missing_v1_in_preset_raises
- test_importable_function_accepts_relative_preset_path

7 new MCP-server tests in tests/test_mcp_aggregate_directives.py:
- test_mcp_server_exposes_aggregate_directives_spec
- test_mcp_server_impl_default_preset
- test_mcp_server_impl_explicit_preset_path
- test_mcp_server_impl_max_chars_truncates
- test_mcp_server_impl_missing_preset_returns_error_string
- test_mcp_server_impl_zero_max_chars_is_unlimited
- test_mcp_server_list_tools_includes_aggregate_directives

Tests load modules via importlib.util.spec_from_file_location so the
scripts/mcp_server.py main guard is not triggered during testing.
sandbox-policy compliance: all write paths use tmp_path.

22 tests pass (15 aggregate_directives + 7 MCP); 7 pre-existing CLI
tests still pass byte-identical.
2026-07-03 10:52:16 -04:00
ed a527f6224f feat(mcp-server): add aggregate_directives tool to scripts/mcp_server.py
Adds a new MCP tool that exposes scripts/aggregate_directives.py as an
LLM-callable endpoint. The tool reads a presets markdown file, resolves
each directive's v1.md, and returns the concatenated clean directive
bodies. metadata (meta.md) is never read.

Tool contract:
  name: aggregate_directives
  params: preset_path (string, default current_baseline.md), max_chars (int, default 0)
  result: text content (clean bodies) or 'ERROR: ...' string on failure

Dispatch runs the impl in asyncio.to_thread so the 66+ v1.md reads do not
block the MCP stdio loop. Errors from aggregate_directives are caught and
formatted as 'ERROR: aggregate_directives(<path>) failed: <type>: <msg>' so
the LLM can read the failure mode inline. Tool count: 45 (MCP_TOOL_SPECS)
+ 2 (run_powershell, aggregate_directives) = 47.

Added scripts/ to sys.path so 'from aggregate_directives import ...' works
inside the MCP server process; idempotent with existing project_root and src
path inserts.
2026-07-03 10:48:44 -04:00
ed 5df938805e refactor(aggregate_directives): expose importable aggregate_directives(preset_path, max_chars, project_root) function
Wraps the body-building core in a public aggregate_directives() function that
returns the rendered string and raises FileNotFoundError/ValueError on errors.
The existing aggregate() CLI wrapper catches those and preserves the prior
stdout/file output behavior. parse_preset now accepts an optional root param so
directive paths inside the preset resolve against a caller-supplied
project_root instead of always REPO_ROOT. No behavior change for the CLI; the
new function is the API surface used by the upcoming MCP server tool.
2026-07-03 10:46:19 -04:00
ed 7d7f88f823 test(aggregate_directives): assert every v1.md has top-level # header 2026-07-03 09:46:33 -04:00
ed 4d2179c38f feat(directives): add rule-statement header to 57-63 of 66 v1.md files 2026-07-03 09:45:50 -04:00
ed 17e3a37b41 feat(directives): add rule-statement header to 49-56 of 66 v1.md files 2026-07-03 09:45:49 -04:00
ed 8cf4b57a74 feat(directives): add rule-statement header to 41-48 of 66 v1.md files 2026-07-03 09:45:49 -04:00
ed ead7cf2869 feat(directives): add rule-statement header to 33-40 of 66 v1.md files 2026-07-03 09:45:48 -04:00
ed e9275236f2 feat(directives): add rule-statement header to 25-32 of 66 v1.md files 2026-07-03 09:45:47 -04:00
ed 4692d01c54 feat(directives): add rule-statement header to 17-24 of 66 v1.md files 2026-07-03 09:45:46 -04:00
ed 85ae99c6a5 feat(directives): add rule-statement header to 9-16 of 66 v1.md files 2026-07-03 09:45:46 -04:00
ed cf8eab9f79 feat(directives): add rule-statement header to 1-8 of 66 v1.md files 2026-07-03 09:45:45 -04:00
ed cd272e5c8e conductor(plan): mark Phase 4 expansion complete (E.1-E.5; 66 directives; aggregation script) 2026-07-03 00:05:42 -04:00
ed 8ef66e0287 chore(artifacts): keep throwaway expansion helpers in scripts/tier2/artifacts/ (per Tier 2 convention) 2026-07-03 00:05:08 -04:00
ed 465433e067 feat(directives): add 15 new directives (Phase A expansion) to current_baseline preset
Adds: ast_parse_insufficient, ast_verify_class_methods_after_edit,
contract_change_audit, convention_enforcement_4_mechanisms,
core_value_read_first, decorator_orphan_pitfall,
defer_not_catch_for_native_crashes, edit_small_incremental,
live_gui_session_scoped_no_restart, no_real_io_during_tests,
preserve_line_endings, reset_session_preserves_project_path,
test_narrow_not_kitchen_sink, undo_redo_100_snapshot_capacity,
verify_before_editing.

Total: 66 directives in the preset (51 original + 15 expansion).
2026-07-03 00:04:54 -04:00
ed 9d3222ddc0 feat(scripts): add scripts/aggregate_directives.py (presets -> clean body aggregation)
Aggregates v1.md bodies from a preset markdown into a single output stream.
NEVER reads meta.md (pollution fix). Stdlib-only; supports stdout and -o.

5 tests in tests/test_aggregate_directives.py cover: success path,
no-meta-md pollution, missing-file error, -o flag, no provenance leakage.
2026-07-03 00:03:51 -04:00
ed 454fac1bff feat(directives): harvest 2 directives from docs/guide_state_lifecycle.md (undo/redo 100-snapshot, reset preserves project path) 2026-07-02 23:56:33 -04:00
ed a758f0a4c9 feat(directives): harvest 5 directives from docs/guide_testing.md (sandbox overview, live_gui session-scoped, defer-not-catch, narrow tests, AST method visibility) 2026-07-02 23:56:26 -04:00
ed 782530ba6d feat(directives): harvest 6 directives from conductor/edit_workflow.md (Rules 1, 2, 6, 7, 8 contract-change, 8 EOL) 2026-07-02 23:56:18 -04:00
ed 8407742a14 feat(directives): harvest 2 directives from docs/AGENTS.md (core_value_read_first, convention_enforcement_4_mechanisms) 2026-07-02 23:56:11 -04:00
ed 559db09ce4 refactor(directives): strip metadata header from v1.md (45-51 of 51; meta extracted to meta.md) 2026-07-02 23:47:20 -04:00
ed 5b0f932c4b refactor(directives): strip metadata header from v1.md (36-44 of 51; meta extracted to meta.md) 2026-07-02 23:47:12 -04:00
ed 68352ee206 refactor(directives): strip metadata header from v1.md (27-35 of 51; meta extracted to meta.md) 2026-07-02 23:46:41 -04:00
ed 831499622d refactor(directives): strip metadata header from v1.md (18-26 of 51; meta extracted to meta.md) 2026-07-02 23:46:33 -04:00
ed 71e01dfee7 refactor(directives): strip metadata header from v1.md (9-17 of 51; meta extracted to meta.md) 2026-07-02 23:46:16 -04:00
ed e9a19523d6 refactor(directives): strip metadata header from v1.md (1-8 of 51; meta extracted to meta.md) 2026-07-02 23:46:04 -04:00
ed c9f30abffc docs(reports): TRACK_COMPLETION_directive_hotswap_harness_20260627 2026-07-02 23:27:25 -04:00
ed bbbfbd3947 conductor(verify): Phase 3 §3.1 directory structure verification — all 5 criteria match 2026-07-02 23:26:13 -04:00
ed 6ba4bdde48 feat(role-prompts): add 5 .warm.md duplicates (warm-with: bootstrap; originals untouched as rollback target) 2026-07-02 23:24:17 -04:00
ed 7b0d116479 feat(role-prompts): add tier2-autonomous.warm.md duplicate with warm-with: bootstrap 2026-07-02 23:23:34 -04:00
ed b2aebbc97f feat(role-prompts): add tier4-qa.warm.md duplicate with warm-with: bootstrap 2026-07-02 23:22:47 -04:00
ed 40764252f4 feat(role-prompts): add tier3-worker.warm.md duplicate with warm-with: bootstrap 2026-07-02 23:21:58 -04:00
ed b082cb158c feat(role-prompts): add tier2-tech-lead.warm.md duplicate with warm-with: bootstrap 2026-07-02 23:21:14 -04:00
ed 35831084f5 feat(role-prompts): add tier1-orchestrator.warm.md duplicate with warm-with: bootstrap 2026-07-02 23:20:07 -04:00
ed 2ddaeb52c1 feat(directives): add current_baseline preset (51 directives, all v1) 2026-07-02 23:14:36 -04:00
ed e13d8b9b8a docs(harness): Phase 2 makes .warm.md duplicates not in-place edits (per user 2026-07-02)
The 'update role prompts' step in Phase 2 was specified as in-place
edits to the 5 .opencode/agents/*.md role prompts. The user clarified
2026-07-02 that this should be making duplicates with a .warm.md
suffix so the originals stay as an explicit rollback target.

Changed:
- spec.md: added a USER DIRECTIVE block above 'The role-prompt
  bootstrap' section explaining the .warm.md duplicate convention.
- plan.md: the Phase 2 preamble references the spec directive; Steps
  2.3-2.7 rewritten as 'Create duplicate <name>.warm.md (do NOT
  modify the original)'. Output paths now include .warm.md suffix.
- dispatch_tier3_phase1.md: appended a USER DIRECTIVE section at the
  end + a NEVER-run-chronology-regenerate note (the script corrupts
  Unicode, per the 2026-07-02 revert of bfebb718).

Phase 1 work (51 v1.md directives + 11 commits ee36eaed..8162f629)
is unaffected. Master is at 068411ee (after the chronology revert).
This commit brings master to this point.
2026-07-02 22:53:27 -04:00
ed 068411ee0b Revert "docs(chronology): regenerate after directive_hotswap_harness_20260627 phase-1 commit"
This reverts commit bfebb71871.
2026-07-02 22:46:07 -04:00
ed bfebb71871 docs(chronology): regenerate after directive_hotswap_harness_20260627 phase-1 commit 2026-07-02 22:27:04 -04:00
ed 8162f629c2 conductor(plan): Mark t1_11 complete + phase_1 done (51 directives, current_phase=2) 2026-07-02 22:23:18 -04:00
ed ce0564fef6 feat(directives): commit Phase 1 harvest summary (51 v1.md files)
Systematic extraction of every directive-like statement from the entire doc tree
into conductor/directives/<name>/v1.md files. 51 v1 files lifted verbatim from
production docs.

Per-task atomic commits (t1_1..t1_10) provide the per-directive provenance.
This meta-commit captures the harvest-level summary.

Sources combed: AGENTS.md, conductor/workflow.md, conductor/product-guidelines.md,
conductor/tech-stack.md, all 14 conductor/code_styleguides/*.md,
.opencode/commands/*.md.

Original docs remain untouched as canonical source. The conductor/directives/
tree is a parallel structure, not a replacement. Future v2+ variants can
test alternative encodings (rationale-first, before/after, tabular) against
this baseline.
2026-07-02 22:21:31 -04:00
ed 9d6a30467f conductor(plan): Mark t1_10 complete 2026-07-02 22:17:57 -04:00
ed cdc0f1405c feat(directives): harvest 8 directives from feature_flags/RAG/cache/knowledge styleguides + 4 new styleguides 2026-07-02 22:17:20 -04:00
ed 09d6a9c7c8 conductor(plan): Mark t1_9 complete 2026-07-02 22:11:56 -04:00
ed fa3e5381fe feat(directives): harvest 5 GUI/architecture directives from product-guidelines.md + python.md 2026-07-02 22:11:19 -04:00
ed 2d07df594d conductor(plan): Mark t1_8 complete 2026-07-02 22:09:46 -04:00
ed 77ee0c68fd feat(directives): harvest 6 process anti-pattern directives from AGENTS.md + workflow.md 2026-07-02 22:09:05 -04:00
ed c8094ec225 conductor(plan): Mark t1_7 complete 2026-07-02 22:07:38 -04:00
ed 412494d205 feat(directives): harvest 10 process/workflow directives from AGENTS.md + workflow.md 2026-07-02 22:07:08 -04:00
ed 02320a13ea conductor(plan): Mark t1_6 complete 2026-07-02 22:03:00 -04:00
ed fa488ccfc6 feat(directives): harvest 3 file/taxonomy directives from AGENTS.md + workflow.md 2026-07-02 22:02:35 -04:00
ed 462860de54 conductor(plan): Mark t1_5 complete 2026-07-02 22:01:25 -04:00
ed b5baaaaab7 feat(directives): harvest 5 code style directives from python.md + workflow.md + product-guidelines.md + AGENTS.md 2026-07-02 22:00:37 -04:00
ed a3e587280e conductor(plan): Mark t1_4 complete 2026-07-02 21:58:09 -04:00
ed 62fc04b172 feat(directives): harvest 3 type/data-structure directives + update boundary_layer_exception 2026-07-02 21:57:36 -04:00
ed 5ba69aaa6c conductor(plan): Mark t1_3 complete 2026-07-02 21:55:22 -04:00
ed 0340925d3e feat(directives): harvest 2 directives from error_handling.md (Result pattern + nil-sentinel) 2026-07-02 21:54:48 -04:00
ed ee36eaed9a conductor(plan): Mark t1_2 complete 2026-07-02 21:52:19 -04:00
ed 545ccee118 feat(directives): harvest 3 directives from tier3-worker.md §17.9 (import/aliasing/from_dict bans) 2026-07-02 21:51:48 -04:00
ed 64485a7859 conductor(plan): Mark t1_1 complete 2026-07-02 21:50:26 -04:00
ed f4dfb84681 feat(directives): harvest 7 directives from python.md §17.1-17.7 (banned patterns + boundary exception) 2026-07-02 21:49:13 -04:00
ed 41c8678b28 conductor(track): state.toml phase 0->1 + add Tier 3 dispatch prompt for Phase 1
state.toml: current_phase 0 -> 1; phase_1.status 'pending' -> 'in_progress';
last_updated 2026-06-27 -> 2026-07-02 (matches the drift-audit batch that
updated the plan/spec).

dispatch_tier3_phase1.md: surgical prompt for the Tier 3 worker that
will execute Phase 1 (lift 48 directives verbatim from the doc tree
into conductor/directives/ per the plan). Captures the 2026-07-02 drift-
audit findings so the harvester doesn't propagate stale refs (e.g.,
the corrected audit_optional_in_3_files.py naming; the corrected §17
line ranges; the corrected send() is canonical, not send_result()).
Includes STOP-AFTER-PHASE-1 rule so Phase 2 is NOT auto-dispatched.

Per conductor/workflow.md Tier 1 Orchestrator role: this is the
maximum responsibility I (Tier 1) take in this session — actual code
generation is delegated to Tier 2/3.
2026-07-02 21:35:18 -04:00
ed fcba9e8935 Merge remote-tracking branch 'origin/master' 2026-07-02 20:53:31 -04:00
ed 6f4832b6a7 docs(skill): rewrite mma-orchestrator SKILL.md for OpenCode Task tool
The mma-orchestrator skill is what the meta-tooling Tier 1/2 agents
load. The previous version was entirely built around the deprecated
scripts/mma_exec.py / claude_mma_exec.py bridge scripts — every
example used 'uv run python scripts/mma_exec.py --role tierN-X ...'
which was deprecated 2026-06-27 in favor of the OpenCode Task tool.
Rewrote the skill to use the OpenCode Task tool's subagent_type
parameter (tier3-worker / tier4-qa / tier1-orchestrator /
tier2-tech-lead) as the canonical mechanism, with explicit
deprecation notes for mma_exec.py.

Also updated: tool count (26 -> 45, now in src/mcp_tool_specs.py);
data locations (Ticket/Track/WorkerContext now in src/mma.py; the
src/models.py shim note).

The 8 mma_exec.py invocation examples in the previous version would
have caused Tier 2 Tech Lead agents to literally invoke deprecated
scripts. This is the highest-impact drift of the session — the user
explicitly said the deprecated invocation was wrong, and this skill
is what loaded the wrong pattern into agent context.
2026-07-02 20:42:48 -04:00
ed 524bff6eb9 docs(guides): file-size drift + FileItem/ContextPreset location drift
- guide_ai_client.md: ~116KB -> ~166KB (src/ai_client.py actual size);
  '46 tools' clarified to '45 MCP tools + the PowerShell shell tool
  defined here in ai_client.py'.
- guide_gui_2.md: '~260KB, ~5400 lines' -> '~437KB, ~8970 lines
  (as of 2026-07-02)'.
- guide_context_curation.md: 'src/models.py:510 + :909' FileItem
  + ContextPreset line refs -> src/project_files.py +
  src/context_presets.py (per module_taxonomy_refactor_20260627).
2026-07-02 20:42:00 -04:00
ed 444ee13f7f docs(guides): fix vendor_capabilities.py, mma_exec.py, audit_optional_returns drift
- Readme.md AI Client row: '5 providers' -> 8; added VendorCapabilities
  inlining note + Result[str] send() API note.
- guide_ai_client.md: src/vendor_capabilities.py refs -> src/ai_client.py
  #region: Vendor Capabilities (2 sites: the capabilities param
  comment + the V2 Capability Matrix section).
- docs/AGENTS.md: audit_optional_returns.py -> audit_optional_in_3_files.py
  (the live script; the successor is not yet built).
- guide_mma.md SubConversationRunner sketch: 'Reuses mma_exec.py' ->
  'Would reuse the WorkerPool internal subprocess template (NOT the
  deprecated mma_exec.py)'.
- guide_multi_agent_conductor.md: architecture diagram box
  'Workers call mma_exec.py' -> 'Workers run via the internal
  subprocess template (run_worker_lifecycle; NOT the deprecated
  meta-tooling mma_exec.py)'; the mma_exec.py box relabeled to
  run_worker_lifecycle; See Also 'scripts/mma_exec.py — sub-agent
  entry point' -> 'src/multi_agent_conductor.py:run_worker_lifecycle
  (NOT the deprecated meta-tooling mma_exec.py)'.
2026-07-02 19:39:43 -04:00
ed 3423cc35a0 docs(workflow): fix file sizes, provider count, mma_exec deprecation in TDD section
workflow.md Architecture Fallback section still claimed gui_2 260KB,
ai_client 116KB/5 providers, mcp_client 81KB, app_controller 166KB,
multi_agent_conductor 28KB+10KB, and 'src/models.py (132KB)
centralized registry'. All updated to current sizes + the shim
reality + 8-provider count + mcp_tool_specs split + the run_worker_
lifecycle subprocess template (not mma_exec.py).

Standard Task Workflow steps 4-5 (Delegate Test Creation / Delegate
Implementation) still used the deprecated 'python scripts/mma_exec.py
--role tier3-worker' invocation as the primary example. Updated to
'OpenCode Task tool with subagent_type: tier3-worker' as the canonical
mechanism, with mma_exec noted as DEPRECATED. Same fix applied to the
Phase Completion Verification step (Tier 4 QA Agent).
2026-07-02 19:35:27 -04:00
ed 8b7b8b96c7 docs(styleguides+harness-plan): fix stale models.py refs + harness plan line drift
Styleguides:
- agent_memory_dimensions.md: FileItem/ContextPreset line refs
  (src/models.py:510-559 / 909-937) -> src/project_files.py +
  src/context_presets.py (moved per module_taxonomy_refactor_20260627).
- config_state_owner.md: 'file I/O primitives in src/models.py' ->
  src/project_manager.py (the config-load/save helpers moved).
- python.md §10 NOT-exempt list: '12 per-aggregate types' -> ~19;
  'dataclass types in src/models.py' -> per-system files enumeration
  (mma.py, project_files.py, mcp_tool_specs.py, result_types.py,
  personas.py, workspace_manager.py, mcp_client.py).
- type_aliases.md: type-registry lookup examples updated to the
  per-system files (src_mma.md, src_project_files.md, etc.); the
  'src/models.py: 48 dataclass field types' worked-example line is
  flagged as historical (pre-refactor state).

Harness plan (directive_hotswap_harness_20260627/plan.md):
- §17 line refs corrected: 17.1 220-237 -> 247-264; 17.2 239-250 ->
  266-277; 17.3 252-272 -> 279-299; 17.4 274-299 -> 301-326; 17.5
  301-311 -> 328-338; 17.6 313-323 -> 340-350; 17.7 325-327 ->
  352-354; 17.9 336-409 -> 364-443.
- §17 master range 216-409 -> 243-473.
- §12 175-184 -> 202-211; §13 185-199 -> 212-224; §15 205-215 ->
  234-241.
- error_handling.md: hard rules 212-242 -> 212-264; boundary types
  274-311 -> 284-365.
- type_aliases.md: 40-81 -> 13-87 + 89-160 + 284-365 (the alias
  table + decision pattern 2.5 + boundary/anti-pattern sections).
2026-07-02 19:31:32 -04:00
ed 46f0ec152a docs(guides): fix stale src/models.py refs + line-number drift across 11 guides
Sweep of the per-source-file guides + Readme.md for stale references
to src/models.py as the data model home. models.py is now a ~1.5KB
re-export shim per module_taxonomy_refactor_20260627; dataclasses
moved to src/mma.py, src/project_files.py, src/type_aliases.py,
src/mcp_tool_specs.py, src/result_types.py, src/context_presets.py,
src/workspace_manager.py, src/personas.py.

- guide_gui_2.md: _gui_func line 754->1062; render_main_interface
  line 1259->1898.
- python.md §10 exemption table: App gui_2.py:307->314; AppController
  795->801; RAGEngine 123->125; HookServer 856->941;
  HookServerInstance 130->171; HookHandler 155->208;
  WebSocketServer 908->993.
- guide_multi_agent_conductor.md: Ticket now in src/mma.py (not
  models.py); ConductorEngine 116+->112+; WorkerPool 50-114->52-110.
- guide_agent_memory_dimensions.md: FileItem ref models.py:510-559
  -> src/project_files.py.
- guide_context_aggregation.md: FileItem + ContextPreset refs ->
  src/project_files.py + src/context_presets.py; ai_client _send_*
  count 5->8.
- guide_discussions.md: parse_history_entries now in src/mma.py;
  ContextPreset/FileItem cross-refs updated.
- guide_personas.md: Persona now in src/personas.py; import example
  updated.
- guide_rag.md: RAGConfig now in src/mcp_client.py.
- guide_workspace_profiles.md: WorkspaceProfile now in
  src/workspace_manager.py.
- guide_mma.md: Data Structures section notes src/mma.py as the live
  location.
- Readme.md: MMA Engine + Data Models rows updated for the
  models.py shim reality + 8 providers + mma_exec deprecation.
- guide_ai_client.md: '5 provider SDKs'->8 (added note); PROVIDERS
  line 56->62; __getattr__ re-export line 261->31; provider
  switching examples use real registered model names (claude-sonnet-
  4-5, MiniMax-M2, gemini-2.5-flash, qwen-plus, grok-2, llama-3.1);
  _provider union lists all 8 providers.
2026-07-02 19:25:33 -04:00
ed e80d2952bc docs(harness+conductor): cover 5 missing styleguides in harvest; cosmetic cleanups
directive_hotswap_harness_20260627/spec.md: added the 5 missing
conductor/code_styleguides/*.md files to the harvest source list
(config_state_owner, workspace_paths, test_sandbox, chroma_cache,
code_path_audit) — the original list named 9 of 14. Added note that
the harvester should verify each contains a harvestable directive
before creating v1.md. Also verified tier2-autonomous.md path
exists (17,940 bytes) for plan Step 2.7.

tech-stack.md: removed duplicate src/paths.py entry (the full
description at line 37 already covers it; the line-73 stub was
redundant).

edit_workflow.md: Key Files section now points to src/project_files.py
for FileItem (moved out of src/models.py per the module taxonomy
refactor) and notes the line ~2748 ref may drift.
2026-07-02 19:05:43 -04:00
ed b9228f3ca4 artifacts 2026-07-02 19:03:49 -04:00
ed e5a8a84381 docs(conductor): type_aliases count, tracks.md stale rows, index.md guide count
- product-guidelines.md: '10 aliases' -> core + extended per-aggregate
  dataclasses (Metadata, CommsLogEntry, ... PathInfo, FileItemsDiff,
  JsonPrimitive/JsonValue) reflecting the actual ~19 types in
  src/type_aliases.py.
- tracks.md: marked rows 2 (qwen_llama_grok), 3 (data_oriented_error
  handling), 17 (code_path_audit) as Completed per chronology.md
  (they were stale 'in progress' / 'ready to start'). Added cleanup
  note. data_structure_strengthening already dropped.
- index.md: '27 deep-dive guides' -> 41 (actual count in docs/);
  refreshed last doc-refresh date to 2026-07-02 with the drift-fix
  summary.
2026-07-02 19:02:35 -04:00
ed 84372a9ae0 docs(conductor): update provider count (8 not 5), file sizes, mcp_tool_specs split
product.md and tech-stack.md claimed 5 providers (missing qwen, grok,
llama), wrong file sizes (gui_2 260KB vs 437KB, ai_client 116KB vs
166KB, mcp_client 81KB vs 92KB, app_controller 166KB vs 240KB), and
described models.py as 132KB centralized registry. Updated to 8
providers, current sizes, and the mcp_tool_specs.py extraction.
Centralized Registry Management -> Per-System Registry Management
reflects the post-refactor reality (PROVIDERS in ai_client.py,
tool registry in mcp_tool_specs.py, models.py is a shim).
2026-07-02 18:59:54 -04:00
ed 9d1fef738a docs(styleguide): fix audit_optional_returns.py — script does not exist
python.md §17.8 and §17.10 listed audit_optional_returns.py as the
 implemented successor to audit_optional_in_3_files.py. Verified
via py_find_usages: audit_optional_returns does not exist in scripts/.
The live script is audit_optional_in_3_files.py (covers 4 baseline
files). Corrected the enforcement table + pre-commit workflow blocks.
2026-07-02 18:56:16 -04:00
ed 3ff759ad66 docs(api): fix backwards send_result claim — send() is canonical
product-guidelines.md claimed send_result() was canonical and send()
was removed. The reverse is true: send() is the live public API
returning Result[str, ErrorInfo]; send_result exists only as a local
var in _send_gemini_cli. Corrected the section and added a
reconciliation note. Also fixed guide_ai_client.md send() signature
to show -> Result[str] instead of -> str.
2026-07-02 18:55:20 -04:00
ed f463edf93d docs(guide_models): rewrite for src/models.py shim reality
The guide described models.py as a 132KB centralized registry with
Provider/ModelInfo enums, Ticket/Track classes, AGENT_TOOL_NAMES, and
parse_plan_md. None of that is in models.py anymore — it's a ~1.5KB
re-export shim (Metadata=TrackMetadata alias + PROVIDERS lazy
__getattr__). Dataclasses moved to per-system files (mma.py,
project_files.py, type_aliases.py, mcp_tool_specs.py, result_types.py).
VendorCapabilities moved from the deleted vendor_capabilities.py into
ai_client.py #region. Rewrote the guide to reflect the current
where-each-model-lives table.
2026-07-02 18:53:56 -04:00
ed 2b4c6c7a56 conductor(chronology_v2): user sign-off recorded — track complete 2026-07-02 12:10:34 -04:00
ed fde60ce864 Merge remote-tracking branch 'tier2-clone/tier2/result_migration_polish_20260630' 2026-07-02 11:59:00 -04:00
ed c7db143688 docs(reports): mark test_visual_sim_mma_v2 as fixed in commit 9cfbb980
test_visual_sim_mma_v2 was the user's explicit complaint about pre-
existing flakes. The fix (9cfbb980) addresses three root causes:
dict metadata normalization, App-side state sync in load_track, and
btn_reset pollution cleanup. Update the report to reflect this.
2026-07-02 10:23:45 -04:00
ed 9cfbb980bd fix(mma_lifecycle): load + start + active_tickets sync for batched tests
The test_visual_sim_mma_v2 failure in tier-3 batch context was caused
by state pollution from prior live_gui tests sharing the subprocess:

1. track state file missing for leftover track:
   `_cb_load_track_result` accessed `state.metadata.id` and
   `state.metadata.name`, but `EMPTY_TRACK_STATE` (returned when no
   state.toml exists for a track_id) had `metadata={}` (a dict, not a
   TrackMetadata object). That raised `'dict' object has no attribute
   'id'`. Fixed by normalizing metadata: dict -> TrackMetadata.from_dict,
   TrackMetadata stays, anything else -> TrackMetadata(id=track_id,
   name=track_id).

2. active_track and active_tickets never reached the App:
   `_cb_load_track_result` set `self.active_track` (controller) and
   `self.active_tickets = []` (via `_load_active_tickets`) but never
   mirrored to `self._app.active_track` / `self._app.active_tickets`.
   The /api/gui/mma_status endpoint reads `app.X` first via
   `_get_app_attr`, so it returned None / [] and the test's poll
   `at_id == track_id and bool(s.get('active_tickets'))` failed. Fixed
   by mirroring active_track / active_tickets / active_tier to the App.

3. leftover tracks list in batched run:
   Without btn_reset, `app.tracks` (App-side) accumulates tracks from
   earlier tests in the session. `_get_app_attr(app, 'tracks', [])`
   then returns stale leftovers, and the test's
   `target_track = next((... if 'hello_world' in t.get('title') else
   tracks_list[0]))` picks a leftover with no on-disk state file.
   Fixed in TWO places:
   (a) btn_reset now also clears `app.tracks = []` so each test starts
       clean if it calls btn_reset.
   (b) tests/test_visual_sim_mma_v2.py now calls `client.click('btn_reset')`
       at the start. The test was the only one in its tier that did NOT
       reset; with the live_gui subprocess shared across batched tests,
       that's the source of the state pollution.

Also reverted the `TrackState.metadata` default change (dict -> TrackMetadata)
because it broke TrackState() construction (TrackMetadata requires `id`/`name`).
The metadata normalization in `_cb_load_track_result` is sufficient and
preserves backward compatibility with on-disk state.toml files.

Verified: tests/test_visual_sim_mma_v2.py passes in isolation (58.36s)
and tier-3 batch passes when run standalone. Other tests in tier-3
have pre-existing render-loop contention flakes in batched xdist mode
that are unrelated to this fix.
2026-07-02 10:23:14 -04:00
ed 4d0bd47bbf docs(report): final quality report — 200 Completed / 10 Abandoned (manually verified via git history) 2026-07-02 09:33:28 -04:00
ed a9b9cf3960 fix(chronology): final classifier — work commits OR 'complete' in messages = Completed; 0 evidence = Abandoned (200 Completed / 10 Abandoned / 0 Needs Review) 2026-07-02 09:33:02 -04:00
ed b803f56d58 fix(chronology): honest classifier — archive tracks without completion evidence are Needs Review (not completed, not abandoned) 2026-07-02 09:13:20 -04:00
ed b6adb15666 docs(report): final update — archive = completed (git mv is the completion signal) 2026-07-02 08:59:02 -04:00
ed 864100b4a7 fix(chronology): archive = completed (the git mv IS the completion signal; don't guess Abandoned) 2026-07-02 08:56:15 -04:00
ed 2d5ce12c7b docs(report): update quality + completion reports with honest Needs Review status for 43 ambiguous archive tracks 2026-07-02 08:25:12 -04:00
ed 792dd7d430 fix(chronology): mark genuinely-ambiguous archive tracks as Needs Review instead of guessing Abandoned (work may be in src/ not track folder) 2026-07-02 08:24:25 -04:00
ed f0eba0c5eb docs(report): update TRACK_COMPLETION with honest manual-review notes 2026-07-02 08:19:03 -04:00
ed 03b0403a35 docs(report): update CHRONOLOGY_QUALITY_20260701 with corrected status distribution (167 Completed / 43 Abandoned) + manual review notes 2026-07-02 08:18:40 -04:00
ed cc23a0586d fix(chronology): add 'mark as completed' + archive-move heuristic for old tracks without state.toml 2026-07-02 08:17:54 -04:00
ed ebd4704324 fix(chronology): respect state.toml status as override + plan-progression heuristic for old archive tracks 2026-07-02 08:14:30 -04:00
ed 9f268fd3e2 docs(report): add TRACK_COMPLETION_chronology_v2_20260701 2026-07-01 23:56:44 -04:00