Files
manual_slop/conductor/archive/chronology_20260619/plan.md
T

62 KiB
Raw Blame History

Conductor Chronology v2 Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Rewrite the chronology track to use a git-history-based classifier as the primary evidence (per the handover §"Rewrite _classify_status to use git history as primary evidence"). Move the broken v1 conductor/chronology.mdconductor/chronology.md.broken-v1; generate v2 with the new classifier; add a per-row evidence line + confidence level + "Needs Review" section; enforce a 30% quality gate (FR7) before Tier 1 reviews the "Needs Review" queue; preserve the v1 pruning of tracks.md + the 3-step convention in workflow.md; write a v2 addendum to the migration report + end-of-track report; get user sign-off.

Architecture: Rewrite the helper script (scripts/audit/generate_chronology.py) to use the 5-step git-history algorithm (FR5). Add a new quality gate script (scripts/audit/chronology_quality_gate.py, FR7) that gates the cross-check on the classifier's confidence distribution. The 3-stage cross-check (FR6) is: (1) classifier auto-classifies high-confidence rows, (2) Tier 1 reviews the "Needs Review" queue, (3) user reviews the final v2 with per-row evidence. No CI integration; the file is hand-maintained like tracks.md.

Tech Stack: Python 3.11+ (helper script + quality gate), tomllib, git log (for commit SHAs), pathlib, re. No new production code in src/. No new dependencies. Test runner: uv run python scripts/run_tests_batched.py (per conductor/workflow.md "Tier 2 Autonomous Sandbox" rule #1 — NEVER direct uv run pytest).

Spec reference: conductor/tracks/chronology_20260619/spec.md (276 lines; 7 FRs, 7 NFRs, 14 VCs, 10 Risks, 10 Phases).

Pause points (Tier 2 must stop and report):

  • Phase 5: Quality gate result. If PASS, continue. If ABORT, fall back to manual review (Tier 1 reviews every row) and STOP.
  • Phase 6: Tier 1 review of the "Needs Review" queue. Tier 2 generates the staging file + the per-row evidence log + the queue, then PAUSES for Tier 1.
  • Phase 9: User sign-off. Tier 2 generates the v2 staging file with Tier 1's resolutions, then PAUSES for user.

HARD BAN reminders (per AGENTS.md and conductor/workflow.md):

  • No git restore, git checkout -- <file>, git reset.
  • No day/hour/minute estimates in any artifact.
  • No new src/*.py files.
  • No skipping tests; fix the underlying bug.
  • Per-file atomic commits; no batch commits.

Phase 1: Archive v1 + verify state of carried-forward work

Focus: The v1 work's pruning + 3-step convention + helper script are reused. Move the broken v1 chronology to .broken-v1; reset state.toml to current_phase = 0 so this rewrite starts clean.

  • Task 1.1: Verify tracks.md pruning is intact

    • WHERE: conductor/tracks.md
    • WHAT: Run grep -n "^- \[x\]" conductor/tracks.md | wc -l. Expected: 0 (the only [x] markers should be in the Status legend at the bottom, not in a - [x] line at the start of a row).
    • HOW: git diff --name-only HEAD~1 HEAD -- conductor/tracks.md to see the most recent tracks.md changes (sanity check); then run the grep.
    • SAFETY: If non-zero, STOP and report. The pruning should not have regressed.
    • VERIFY: grep -n "^- \[x\]" conductor/tracks.md returns 0 matches in a - [x] line position.
    • NO COMMIT (verification only).
  • Task 1.2: Verify workflow.md 3-step convention is present

    • WHERE: conductor/workflow.md
    • WHAT: Run grep -n "Archiving a track" conductor/workflow.md. Expected: 1 match.
    • HOW: Single grep; verify the line is the 3-step convention block.
    • SAFETY: If 0 matches or > 1, STOP and report. The 3-step convention should be in workflow.md "Notes > Editing this file" section.
    • VERIFY: The matched line is followed by the 3 numbered steps.
    • NO COMMIT (verification only).
  • Task 1.3: Move chronology.mdchronology.md.broken-v1

    • WHERE: conductor/chronology.mdconductor/chronology.md.broken-v1
    • WHAT: Use git mv to preserve git history (the v1 chronology was committed in commit 8cd92856 and 24 subsequent commits; git mv keeps the rename visible in git log --follow).
    • HOW: git mv conductor/chronology.md conductor/chronology.md.broken-v1. Verify with git status (should show a rename, not a delete+add).
    • SAFETY: Don't delete the file; the rename preserves history. Don't add .broken-v1 to .gitignore — it should be a tracked file (per VC14).
    • VERIFY: git log --follow conductor/chronology.md.broken-v1 shows the v1 history. The file has 218 lines.
    • COMMIT: conductor(chronology): archive v1 chronology to .broken-v1 (VC14)
    • GIT NOTE: "Phase 1.3. v1 chronology moved to .broken-v1 for archival. The new v2 file will be created in Phase 7. v1 status classifier was broken (read stale metadata.json.status); v2 uses git history as primary evidence per FR5."
  • Task 1.4: Reset state.toml to current_phase = 0

    • WHERE: conductor/tracks/chronology_20260619/state.toml
    • WHAT: Set [meta].status = "active", [meta].current_phase = 0, [meta].last_updated = "<today's date YYYY-MM-DD>". Reset all [phases].phase_N.status to "pending", checkpointsha to "". Reset all [tasks].t<N>_<M>.status to "pending", commit_sha to "". Reset all [verification].<flag> to false.
    • HOW: Use manual-slop_edit_file to set the 4 fields (status, current_phase, last_updated, and the rewrite note). For phases/tasks/verification, replace the entire sections with the new content. The new content has phases 1-10 marked pending + tasks marked pending + verification flags false.
    • SAFETY: Don't delete the [meta], [phases], [tasks], [verification], [user_directives_logged] sections; just reset their values. Add a [rewrite_log] section noting the v2 rewrite.
    • VERIFY: uv run python -c "import tomllib; data = tomllib.load(open('conductor/tracks/chronology_20260619/state.toml','rb')); assert data['meta']['current_phase'] == 0; assert data['meta']['status'] == 'active'; print('OK')" exits 0.
    • COMMIT: conductor(plan): reset state.toml to current_phase=0 for v2 rewrite
    • GIT NOTE: "Phase 1.4. state.toml reset for v2 rewrite. The v1 track's 24 commits remain in git history; this rewrite starts at Phase 1 with a git-history-based classifier (per the handover)."
  • Task 1.5: Verify the staging file path is not committed

    • WHERE: conductor/chronology.md.staging (target of Phase 4)
    • WHAT: Confirm the staging file does not exist yet AND that the path is gitignored (so accidental creation doesn't pollute the diff).
    • HOW: Test-Path conductor/chronology.md.staging returns False; grep -n "chronology.md.staging" .gitignore returns a match (or the staging file is added to .gitignore if not already there).
    • SAFETY: If the staging file exists, STOP and report. The v1 chronology should be in .broken-v1 only.
    • VERIFY: Test-Path conductor/chronology.md.staging returns False; the staging path is in .gitignore.
    • NO COMMIT (verification only).
    • CHECKPOINT: conductor(checkpoint): Phase 1 complete — v1 archived, state.toml reset

Phase 2: Rewrite the helper script — git-history classifier (FR5, TDD red+green)

Focus: The script's _classify_status function is rewritten to use the handover's 5-step algorithm. The new signature returns (status, confidence, reason). This is TDD-required per project convention.

  • Task 2.1: Add failing tests for the new _classify_status signature (TDD RED)

    • WHERE: tests/test_generate_chronology.py (modify; add new test class TestClassifyStatus)
    • WHAT: 8 unit tests for the new classifier signature. Each test calls _classify_status(folder_link, init_sha, end_sha, commit_count, first_commit_subject, last_commit_subject, state_phase, metadata_status, last_commit_date) and asserts the returned (status, confidence, reason) tuple.
    • HOW: Add the following tests to the file (1-space indentation per project convention):
      from scripts.audit.generate_chronology import _classify_status
      
      class TestClassifyStatus:
       def test_archive_with_three_plus_work_commits_is_completed_high(self):
        result = _classify_status(
         folder_link="conductor/archive/test_track_20260101",
         init_sha="abc1234", end_sha="def5678", commit_count=5,
         first_commit_subject="feat: add core logic",
         last_commit_subject="fix: edge case",
         state_phase="complete", metadata_status="Completed",
         last_commit_date="2026-01-15",
        )
        assert result == ("Completed", "high", "archive + 5 work commits + state_phase=complete")
      
       def test_archive_with_zero_commits_is_special_low(self):
        result = _classify_status(
         folder_link="conductor/archive/empty_track_20260101",
         init_sha="abc1234", end_sha="abc1234", commit_count=0,
         first_commit_subject="chore: init",
         last_commit_subject="chore: init",
         state_phase="no-state-toml", metadata_status="Completed",
         last_commit_date="2026-01-01",
        )
        status, confidence, _ = result
        assert status == "Special"
        assert confidence == "low"
      
       def test_tracks_with_state_phase_three_plus_is_in_progress_high(self):
        result = _classify_status(
         folder_link="conductor/tracks/active_track_20260101",
         init_sha="abc1234", end_sha="def5678", commit_count=4,
         first_commit_subject="conductor(track): add initial spec",
         last_commit_subject="feat: implement core",
         state_phase="4", metadata_status="active",
         last_commit_date="2026-01-20",
        )
        assert result == ("In Progress", "high", "tracks + 4 work commits + state_phase=4")
      
       def test_tracks_with_zero_to_one_commits_is_active_low(self):
        result = _classify_status(
         folder_link="conductor/tracks/fresh_track_20260101",
         init_sha="abc1234", end_sha="abc1234", commit_count=1,
         first_commit_subject="conductor(track): add initial spec",
         last_commit_subject="conductor(track): add initial spec",
         state_phase="0", metadata_status="spec_written",
         last_commit_date="2026-01-01",
        )
        status, confidence, _ = result
        assert status == "Active"
        assert confidence == "low"
      
       def test_metadata_status_abandoned_is_honoured(self):
        result = _classify_status(
         folder_link="conductor/archive/abandoned_20260101",
         init_sha="abc1234", end_sha="def5678", commit_count=2,
         first_commit_subject="conductor(track): add initial spec",
         last_commit_subject="chore: archive move",
         state_phase="no-state-toml", metadata_status="abandoned",
         last_commit_date="2025-06-01",
        )
        status, confidence, _ = result
        assert status == "Abandoned"
        assert confidence == "high"
      
       def test_metadata_status_superseded_is_mapped_to_special(self):
        result = _classify_status(
         folder_link="conductor/archive/superseded_20260101",
         init_sha="abc1234", end_sha="def5678", commit_count=2,
         first_commit_subject="conductor(track): add initial spec",
         last_commit_subject="chore: archive move",
         state_phase="no-state-toml", metadata_status="superseded",
         last_commit_date="2025-06-01",
        )
        status, confidence, _ = result
        assert status == "Special"
        assert confidence == "high"
      
       def test_metadata_status_contradicts_git_is_low(self):
        result = _classify_status(
         folder_link="conductor/archive/contradict_20260101",
         init_sha="abc1234", end_sha="def5678", commit_count=0,
         first_commit_subject="chore: init",
         last_commit_subject="chore: init",
         state_phase="no-state-toml", metadata_status="Completed",
         last_commit_date="2026-01-01",
        )
        status, confidence, _ = result
        assert status == "Special"
        assert confidence == "low"
      
       def test_tracks_with_three_plus_commits_no_state_toml_is_in_progress(self):
        result = _classify_status(
         folder_link="conductor/tracks/no_state_20260101",
         init_sha="abc1234", end_sha="def5678", commit_count=3,
         first_commit_subject="feat: implement X",
         last_commit_subject="fix: bug in X",
         state_phase="no-state-toml", metadata_status="active",
         last_commit_date="2026-01-15",
        )
        assert result == ("In Progress", "high", "tracks + 3 work commits + no state.toml")
      
       def test_placeholder_in_track_id_keeps_current_status(self):
        result = _classify_status(
         folder_link="conductor/tracks/PLACEHOLDER_test_20260101",
         init_sha="abc1234", end_sha="def5678", commit_count=3,
         first_commit_subject="chore: init",
         last_commit_subject="chore: update",
         state_phase="2", metadata_status="spec_written",
         last_commit_date="2026-01-10",
        )
        assert result == ("spec_written", "high", "PLACEHOLDER track; original status preserved")
      
    • SAFETY: These tests use the _classify_status function with the new signature. The v1 signature was _classify_status(folder_link, current, track_id). The tests will fail with a TypeError or signature mismatch — that's the RED phase.
    • VERIFY: uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py should FAIL with TypeError: _classify_status() got an unexpected keyword argument 'init_sha' (or similar signature mismatch).
    • COMMIT: test(chronology): add 8 failing tests for new _classify_status signature (TDD red)
    • GIT NOTE: "Phase 2.1. TDD red. 8 tests cover the new (status, confidence, reason) tuple signature with the 5-step algorithm. Tests must fail before Task 2.2 rewrites the function."
  • Task 2.2: Rewrite _classify_status to the new signature (TDD GREEN)

    • WHERE: scripts/audit/generate_chronology.py (modify; replace the _classify_status function definition; lines ~163-189)
    • WHAT: Replace the v1 _classify_status(folder_link, current, track_id) function with the new signature _classify_status(folder_link, init_sha, end_sha, commit_count, first_commit_subject, last_commit_subject, state_phase, metadata_status, last_commit_date) that returns (status, confidence, reason).
    • HOW: Use manual-slop_py_update_definition to replace the function body. The new function:
      def _classify_status(
       folder_link: str,
       init_sha: str,
       end_sha: str,
       commit_count: int,
       first_commit_subject: str,
       last_commit_subject: str,
       state_phase: str,
       metadata_status: str,
       last_commit_date: str,
      ) -> tuple[str, str, str]:
       """Classify a track's status using git history as primary evidence.
      
       The 5-step algorithm (per CHRONOLOGY_TRACK_HANDOVER_20260620.md):
       1. Count meaningful commits (commit_count)
       2. Inspect commit messages (first/last subject)
       3. Check state.toml phase progression (state_phase as string from _parse_state_phase)
       4. Default to conservative (low confidence on ambiguous)
       5. Honour explicit metadata (metadata_status)
      
       state_phase values (per _parse_state_phase):
       - "complete" = track shipped
       - "N" (digit string) = current_phase = N
       - "no-state-toml" = no state.toml file
       - "?" = parse error
      
       Returns: (status, confidence, reason) where:
       - status: one of "Active", "In Progress", "Completed", "Abandoned", "Special"
       - confidence: "high" or "low"
       - reason: one-line explanation
       """
       track_id: str = Path(folder_link).name
      
       if "PLACEHOLDER" in track_id:
        return (metadata_status or "unknown", "high", "PLACEHOLDER track; original status preserved")
      
       is_archive: bool = folder_link.startswith("conductor/archive/")
       is_tracks: bool = folder_link.startswith("conductor/tracks/")
      
       work_commit_pattern: re.Pattern[str] = re.compile(r"^(feat|fix|refactor|perf|test)\(")
       is_work_subject: bool = (
        work_commit_pattern.match(first_commit_subject) is not None
        or work_commit_pattern.match(last_commit_subject) is not None
       )
      
       metadata_lower: str = (metadata_status or "").lower()
       state_phase_clean: str = (state_phase or "").strip()
       is_in_flight: bool = state_phase_clean.isdigit() and int(state_phase_clean) >= 3
      
       if metadata_lower in {"abandoned", "superseded"} and not is_work_subject:
        mapped: str = "Abandoned" if metadata_lower == "abandoned" else "Special"
        return (mapped, "high", f"metadata.json.status={metadata_lower}; no work commits; mapped to {mapped}")
      
       if is_archive and commit_count >= 3 and (is_work_subject or state_phase_clean == "complete"):
        return ("Completed", "high", f"archive + {commit_count} work commits + state_phase={state_phase_clean}")
      
       if is_archive and commit_count == 0:
        return ("Special", "low", f"archive folder with 0 work commits; needs review")
      
       if is_tracks and is_in_flight:
        return ("In Progress", "high", f"tracks + {commit_count} work commits + state_phase={state_phase_clean}")
      
       if is_tracks and commit_count >= 3 and state_phase_clean in {"no-state-toml", "?"}:
        return ("In Progress", "high", f"tracks + {commit_count} work commits + no state.toml")
      
       if is_tracks and commit_count <= 1:
        return ("Active", "low", f"tracks + {commit_count} commits; spec-only or early stage")
      
       if metadata_lower in {"completed", "complete", "shipped"} and commit_count == 0:
        return ("Special", "low", f"metadata says {metadata_status} but git shows 0 commits; contradictory")
      
       return ("Special", "low", f"ambiguous: commit_count={commit_count}, state_phase={state_phase_clean}, is_work={is_work_subject}")
      
    • SAFETY: Use 1-space indentation (per project convention). Use tuple[str, str, str] return type (per conductor/code_styleguides/error_handling.md data-oriented pattern; NOT Optional[Tuple]). The function is a pure transformation: inputs → outputs, no side effects, no exceptions. The state_phase is always a string (per _parse_state_phase); the check uses .isdigit() + int() conversion.
    • VERIFY: uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py should now PASS for the 8 new tests in TestClassifyStatus (the 6 v1 tests in TestExtract* may also need updates — see Task 2.3).
    • COMMIT: feat(chronology): rewrite _classify_status to git-history classifier (FR5)
    • GIT NOTE: "Phase 2.2. TDD green. _classify_status now returns (status, confidence, reason) using the 5-step algorithm. The new signature accepts 9 parameters (folder_link, init/end_sha, commit_count, first/last_commit_subject, state_phase, metadata_status, last_commit_date)."
  • Task 2.3: Update walk_track_folders to compute the new per-row evidence fields

    • WHERE: scripts/audit/generate_chronology.py (modify; the walk_track_folders function)
    • WHAT: walk_track_folders currently populates 8 row fields (date, track_id, status, summary, init_sha, end_sha, commit_count, folder_link). Add 6 new fields: first_commit_subject, last_commit_subject, state_phase, metadata_status, last_commit_date, confidence, reason. The init_sha/end_sha/commit_count already exist. Update the call to _classify_status to use the new 9-argument signature.
    • HOW: Use manual-slop_py_update_definition to update the function. The relevant changes:
      # Existing: extract init_sha, end_sha, commit_count via _git_first_line / _git_log
      # Add: extract first_commit_subject + last_commit_subject
      first_commit_subject: str = _git_first_line(folder_relpath, "--reverse", "--format=%s")
      last_commit_subject: str = _git_first_line(folder_relpath, "-1", "--format=%s")
      
      # Add: extract state_phase (string from _parse_state_phase)
      state_phase: str = _parse_state_phase(folder / "state.toml")
      
      # Add: extract metadata_status (read from metadata.json)
      metadata_status: str = ""
      if metadata_path.is_file():
       try:
        data = json.loads(metadata_path.read_text(encoding="utf-8"))
        metadata_status = str(data.get("status", "")).strip()
       except (json.JSONDecodeError, OSError):
        pass
      
      # Add: extract last_commit_date
      last_commit_date: str = _last_commit_date(folder_relpath)
      
      # Update: call _classify_status with the new signature
      status, confidence, reason = _classify_status(
       folder_link=folder_link,
       init_sha=init_sha,
       end_sha=end_sha,
       commit_count=commit_count,
       first_commit_subject=first_commit_subject,
       last_commit_subject=last_commit_subject,
       state_phase=state_phase,
       metadata_status=metadata_status,
       last_commit_date=last_commit_date,
      )
      
      # Update: row dict includes the 6 new fields
      rows.append({
       "date": date,
       "track_id": track_id,
       "status": status,
       "confidence": confidence,
       "summary": summary,
       "init_sha": init_sha,
       "end_sha": end_sha,
       "commit_count": commit_count,
       "folder_link": folder_link,
       "first_commit_subject": first_commit_subject,
       "last_commit_subject": last_commit_subject,
       "state_phase": state_phase,
       "metadata_status": metadata_status,
       "last_commit_date": last_commit_date,
       "reason": reason,
      })
      
    • SAFETY: The existing v1 tests may break because the row dict's keys change. Run the full test file to see what breaks. The 6 v1 tests target extract_slug_date and extract_summary (pure functions), so they should still pass.
    • VERIFY: uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py — the 8 new tests should PASS; the 6 v1 tests should still PASS (no changes to their target functions).
    • COMMIT: refactor(chronology): update walk_track_folders to populate per-row evidence fields
    • GIT NOTE: "Phase 2.3. walk_track_folders now populates 6 new fields per row (first/last_commit_subject, state_phase, metadata_status, last_commit_date, confidence, reason). init_sha/end_sha/commit_count were already present."
  • Task 2.4: Update the v1 tests to match the new row dict shape

    • WHERE: tests/test_generate_chronology.py (modify; the 6 v1 tests in TestExtractSlugDate + TestExtractSummary)
    • WHAT: The 6 v1 tests should still pass because they only test extract_slug_date and extract_summary (pure functions). The new tests for _classify_status and walk_track_folders are in TestClassifyStatus. Run the full test file to confirm.
    • HOW: Re-read the 6 v1 tests. If they fail due to the row dict change, update them. Otherwise no changes needed.
    • SAFETY: Don't delete the 6 v1 tests; they verify extract_slug_date + extract_summary which are still used.
    • VERIFY: uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py — all tests (6 v1 + 8 new) PASS.
    • NO COMMIT (if no changes) or COMMIT: test(chronology): update v1 tests for new row dict shape (FR5)
    • GIT NOTE: "Phase 2.4. v1 tests verified against new row dict shape; no changes needed."
    • CHECKPOINT: conductor(checkpoint): Phase 2 complete — git-history classifier (FR5) green

Phase 3: Update extract_summary + add per-row evidence line to format_markdown (FR1, FR5)

Focus: The summary priority chain needs to reject metadata-field text (the v1 bug). The markdown output needs a per-row evidence line.

  • Task 3.1: Add failing tests for metadata-field text rejection (TDD RED)

    • WHERE: tests/test_generate_chronology.py (modify; add 4 new tests to TestExtractSummary)
    • WHAT: 4 unit tests that the regex ^\*\* rejects metadata-field text. The tests use a fixture folder with spec.md whose first non-empty line is **Priority:** A (foundational; precedes X).
    • HOW: Add the following tests:
      class TestExtractSummaryRejectsMetadata:
       def test_rejects_priority_metadata_line(self, tmp_path: Path):
        (tmp_path / "spec.md").write_text(
         "# Test Track\n\n**Priority:** A (foundational; precedes X)\n\nThis is the actual summary.",
         encoding="utf-8",
        )
        summary: str = extract_summary(tmp_path)
        assert not summary.startswith("**")
        assert "actual summary" in summary
      
       def test_rejects_date_metadata_line(self, tmp_path: Path):
        (tmp_path / "spec.md").write_text(
         "# Test Track\n\n**Date:** 2026-01-01\n\nThe actual summary content.",
         encoding="utf-8",
        )
        summary: str = extract_summary(tmp_path)
        assert not summary.startswith("**")
        assert "actual summary" in summary
      
       def test_rejects_initialized_metadata_line(self, tmp_path: Path):
        (tmp_path / "spec.md").write_text(
         "# Test Track\n\n**Initialized:** 2026-01-01\n\nThe actual summary content.",
         encoding="utf-8",
        )
        summary: str = extract_summary(tmp_path)
        assert not summary.startswith("**")
        assert "actual summary" in summary
      
       def test_rejects_confidence_metadata_line(self, tmp_path: Path):
        (tmp_path / "spec.md").write_text(
         "# Test Track\n\n**Confidence:** high\n\nThe actual summary content.",
         encoding="utf-8",
        )
        summary: str = extract_summary(tmp_path)
        assert not summary.startswith("**")
        assert "actual summary" in summary
      
    • SAFETY: These tests will FAIL with the v1 extract_summary because v1 returns the first non-empty line, which is **Priority:** A.... RED phase.
    • VERIFY: uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py::TestExtractSummaryRejectsMetadata should FAIL.
    • COMMIT: test(chronology): add 4 failing tests for metadata-field text rejection (TDD red)
    • GIT NOTE: "Phase 3.1. TDD red. 4 tests verify that extract_summary skips metadata-field lines (Priority, Date, Initialized, Confidence). Tests must fail before Task 3.2 updates the function."
  • Task 3.2: Update extract_summary to reject metadata-field text (TDD GREEN)

    • WHERE: scripts/audit/generate_chronology.py (modify; the extract_summary function)
    • WHAT: Add a regex check at the start of each candidate line: if the line starts with ** (i.e., matches re.match(r"^\*\*", line)), skip it. Try the next non-empty line.
    • HOW: Use manual-slop_py_update_definition to update the function. The new implementation adds a _METADATA_LINE_RE = re.compile(r"^\*\*") constant and filters candidate lines:
      _METADATA_LINE_RE: re.Pattern[str] = re.compile(r"^\*\*")
      
      def extract_summary(folder_path: Path) -> str:
       """Extract a one-sentence summary from a track folder.
      
       Priority chain (each step rejects metadata-field text starting with '**'):
       1. metadata.json.summary if present and not metadata-field text
       2. First non-empty line of spec.md that does not start with '**'
       3. metadata.json.description if not metadata-field text
       4. First non-empty line of plan.md that does not start with '**'
       5. Generic placeholder based on folder location
       """
       metadata_path: Path = folder_path / "metadata.json"
       if metadata_path.is_file():
        try:
         data: dict = json.loads(metadata_path.read_text(encoding="utf-8"))
        except (json.JSONDecodeError, OSError):
         data = {}
        summary: str | None = data.get("summary")
        if summary and not _METADATA_LINE_RE.match(summary):
         return _truncate_to_25_words(summary)
        description: str | None = data.get("description")
        if description and not _METADATA_LINE_RE.match(description):
         return _truncate_to_25_words(description)
      
       for filename in ("spec.md", "plan.md"):
        file_path: Path = folder_path / filename
        if not file_path.is_file():
         continue
        try:
         text: str = file_path.read_text(encoding="utf-8")
        except OSError:
         continue
        for line in text.splitlines():
         stripped: str = line.strip()
         if not stripped or _METADATA_LINE_RE.match(stripped):
          continue
         if stripped.startswith("#"):
          continue
         if stripped.startswith(">"):
          continue
         return _truncate_to_25_words(_first_sentence(stripped))
      
       if (folder_path.parent.name if folder_path.parent else "") == "archive":
        return "Imported from archive (no spec)"
       return "Track folder (no spec found)"
      
    • SAFETY: Use 1-space indentation. The function remains a pure transformation; no side effects. The json.JSONDecodeError is caught at the SDK boundary (per conductor/code_styleguides/error_handling.md "Exceptions reserved for the SDK boundary") and converted to a recoverable case (empty data dict).
    • VERIFY: uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py — all 18 tests (6 v1 + 8 new classifier + 4 metadata-rejection) PASS.
    • COMMIT: feat(chronology): extract_summary rejects metadata-field text (FR5)
    • GIT NOTE: "Phase 3.2. TDD green. extract_summary now skips lines starting with '**' (Priority, Date, Initialized, Confidence, etc.). Fixes the v1 bug where 167 rows had metadata-field text as their summary."
  • Task 3.3: Add failing tests for per-row evidence line in format_markdown (TDD RED)

    • WHERE: tests/test_generate_chronology.py (modify; add new test class TestFormatMarkdown)
    • WHAT: 3 unit tests for the per-row evidence line. Each test creates a row dict with the new fields and asserts that the output markdown has the evidence sub-line.
    • HOW: Add the following tests:
      class TestFormatMarkdown:
       def test_row_includes_evidence_subline(self):
        rows: list[dict] = [
         {
          "date": "2026-01-01",
          "track_id": "test_track_20260101",
          "status": "Completed",
          "confidence": "high",
          "summary": "Test summary.",
          "folder_link": "conductor/archive/test_track_20260101",
          "init_sha": "abc1234",
          "end_sha": "def5678",
          "commit_count": 5,
          "state_phase": "complete",
          "first_commit_subject": "feat: add core",
          "last_commit_subject": "fix: edge case",
          "reason": "archive + 5 work commits + state_phase=complete",
         },
        ]
        output: str = format_markdown(rows)
        assert "abc1234..def5678" in output
        assert "Evidence:" in output
        assert "feat: add core" in output
        assert "fix: edge case" in output
        assert "confidence=high" in output
      
       def test_row_includes_confidence_in_table(self):
        rows: list[dict] = [
         {
          "date": "2026-01-01",
          "track_id": "test_track_20260101",
          "status": "Special",
          "confidence": "low",
          "summary": "Test summary.",
          "folder_link": "conductor/archive/test_track_20260101",
          "init_sha": "abc1234",
          "end_sha": "abc1234",
          "commit_count": 0,
          "state_phase": None,
          "first_commit_subject": "chore: init",
          "last_commit_subject": "chore: init",
          "reason": "archive folder with 0 work commits",
         },
        ]
        output: str = format_markdown(rows)
        assert "confidence=low" in output
      
       def test_needs_review_section_appended(self):
        rows: list[dict] = [
         {
          "date": "2026-01-01",
          "track_id": "low_conf_20260101",
          "status": "Special",
          "confidence": "low",
          "summary": "Low confidence row.",
          "folder_link": "conductor/archive/low_conf_20260101",
          "init_sha": "abc1234",
          "end_sha": "abc1234",
          "commit_count": 0,
          "state_phase": None,
          "first_commit_subject": "chore: init",
          "last_commit_subject": "chore: init",
          "reason": "archive folder with 0 work commits",
         },
        ]
        output: str = format_markdown(rows)
        assert "## Needs Review" in output
        assert "low_conf_20260101" in output
      
    • SAFETY: These tests will FAIL with the v1 format_markdown because v1 doesn't emit evidence lines or a "Needs Review" section. RED phase.
    • VERIFY: uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py::TestFormatMarkdown should FAIL.
    • COMMIT: test(chronology): add 3 failing tests for per-row evidence line (TDD red)
    • GIT NOTE: "Phase 3.3. TDD red. 3 tests verify the per-row evidence line, confidence in table, and 'Needs Review' section. Tests must fail before Task 3.4 updates format_markdown."
  • Task 3.4: Update format_markdown to emit evidence + "Needs Review" section (TDD GREEN)

    • WHERE: scripts/audit/generate_chronology.py (modify; the format_markdown function)
    • WHAT: Update the table to include the confidence column. Add a sub-line below each row with the evidence. Add a "Needs Review" section at the bottom listing low-confidence rows.
    • HOW: Use manual-slop_py_update_definition to update the function:
      def format_markdown(rows: list[dict]) -> str:
       lines: list[str] = [
        "| Date | ID | Status | Confidence | Summary | Folder | Range |",
        "| --- | --- | --- | --- | --- | --- | --- |",
       ]
       for row in rows:
        range_str: str = f"`{row['init_sha']}..{row['end_sha']}` ({row['commit_count']})"
        confidence_marker: str = f"**{row['confidence']}**" if row["confidence"] == "low" else row["confidence"]
        lines.append(
         f"| {row['date']} | `{row['track_id']}` | {row['status']} | {confidence_marker} | "
         f"{_md_escape(row['summary'])} | `{row['folder_link']}` | {range_str} |"
        )
        evidence_line: str = (
         f"| | | | | | "
         f"Evidence: `{row['init_sha'][:7]}..{row['end_sha'][:7]}` | "
         f"{row['commit_count']} commits | state_phase={row.get('state_phase', 'n/a')} | "
         f"\"{row['first_commit_subject']}\"\"{row['last_commit_subject']}\" | "
         f"reason=\"{row['reason']}\" |"
        )
        lines.append(evidence_line)
      
       low_rows: list[dict] = [r for r in rows if r.get("confidence") == "low"]
       if low_rows:
        lines.append("")
        lines.append("## Needs Review (Tier 1 + User)")
        lines.append("")
        lines.append(f"{len(low_rows)} row(s) had ambiguous git evidence. Resolved by Tier 1; user reviewed in Stage 3.")
        lines.append("")
        for row in low_rows:
         lines.append(
          f"- `{row['track_id']}` (status={row['status']}) — {row['reason']} — resolved by Tier 1"
         )
      
       return "\n".join(lines) + "\n"
      
    • SAFETY: Use 1-space indentation. The function is a pure transformation. The "Needs Review" section only appears if there are low-confidence rows. The evidence line is a sub-line in the table (not a column), keeping the table at 7 columns (was 6; +1 for Confidence).
    • VERIFY: uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py — all 21 tests (6 v1 + 8 classifier + 4 metadata-rejection + 3 format) PASS.
    • COMMIT: feat(chronology): format_markdown emits per-row evidence + Needs Review section (FR1)
    • GIT NOTE: "Phase 3.4. TDD green. format_markdown now emits a 7-column table (added Confidence) with a per-row evidence sub-line and a 'Needs Review' section listing low-confidence rows. v1 had a 6-column table with no evidence."
    • CHECKPOINT: conductor(checkpoint): Phase 3 complete — summary regex + per-row evidence (FR1, FR5)

Phase 4: Add the quality gate script (FR7, TDD)

Focus: A new wrapper script that runs after the classifier's output and gates the cross-check on the confidence distribution.

  • Task 4.1: Add failing tests for the quality gate (TDD RED)

    • WHERE: tests/test_chronology_quality_gate.py (new file)
    • WHAT: 5 unit tests for the quality gate logic. Each test calls check_quality_gate(staging_path, threshold=0.30) and asserts the return code + message.
    • HOW: Create the new file:
      """Tests for scripts/audit/chronology_quality_gate.py."""
      from pathlib import Path
      import pytest
      from scripts.audit.chronology_quality_gate import check_quality_gate
      
      def _write_staging(path: Path, low_count: int, high_count: int) -> Path:
       lines: list[str] = [
        "| Date | ID | Status | Confidence | Summary | Folder | Range |",
        "| --- | --- | --- | --- | --- | --- | --- |",
       ]
       for i in range(high_count):
        lines.append(
         f"| 2026-01-{i+1:02d} | `track_{i}_20260101` | Completed | high | summary | folder | `abc..def` (5) |"
        )
       for i in range(low_count):
        lines.append(
         f"| 2026-01-{i+1:02d} | `low_{i}_20260101` | Special | **low** | summary | folder | `abc..abc` (0) |"
        )
       path.write_text("\n".join(lines), encoding="utf-8")
       return path
      
      def test_zero_percent_low_passes(tmp_path: Path):
       staging: Path = _write_staging(tmp_path / "staging.md", low_count=0, high_count=100)
       passed, message = check_quality_gate(staging, threshold=0.30)
       assert passed is True
       assert "PASS" in message
      
      def test_thirty_percent_low_at_boundary_passes(tmp_path: Path):
       staging: Path = _write_staging(tmp_path / "staging.md", low_count=30, high_count=70)
       passed, _ = check_quality_gate(staging, threshold=0.30)
       assert passed is True
      
      def test_thirty_one_percent_low_fails(tmp_path: Path):
       staging: Path = _write_staging(tmp_path / "staging.md", low_count=31, high_count=69)
       passed, message = check_quality_gate(staging, threshold=0.30)
       assert passed is False
       assert "ABORT" in message
      
      def test_hundred_percent_low_fails(tmp_path: Path):
       staging: Path = _write_staging(tmp_path / "staging.md", low_count=10, high_count=0)
       passed, message = check_quality_gate(staging, threshold=0.30)
       assert passed is False
       assert "ABORT" in message
      
      def test_missing_staging_file_errors(tmp_path: Path):
       with pytest.raises(FileNotFoundError):
        check_quality_gate(tmp_path / "missing.md", threshold=0.30)
      
    • SAFETY: The test file is new; it does not modify any existing file. Use 1-space indentation. Use pytest.raises for the error case. The _write_staging helper creates a minimal staging file with the expected 7-column header.
    • VERIFY: uv run python scripts/run_tests_batched.py tests/test_chronology_quality_gate.py should FAIL with ModuleNotFoundError: No module named 'scripts.audit.chronology_quality_gate'.
    • COMMIT: test(chronology): add 5 failing tests for quality gate (TDD red)
    • GIT NOTE: "Phase 4.1. TDD red. 5 tests cover the 30% threshold logic: 0% passes, 30% at boundary passes, 31% fails, 100% fails, missing file errors. Tests must fail before Task 4.2 writes the script."
  • Task 4.2: Write the quality gate script (TDD GREEN)

    • WHERE: scripts/audit/chronology_quality_gate.py (new file)
    • WHAT: The wrapper script. Exposes check_quality_gate(staging_path, threshold) for programmatic use and a main() for CLI use (--staging PATH --threshold FLOAT).
    • HOW: Create the new file:
      """Quality gate for the chronology classifier (FR7).
      
      Reads the staging file produced by generate_chronology.py and checks the
      confidence distribution. If the ratio of low-confidence rows to total rows
      exceeds the threshold (default 0.30), the gate aborts and the user must
      fall back to manual review (Tier 1 reviews every row).
      """
      from __future__ import annotations
      import argparse
      import re
      import sys
      from pathlib import Path
      
      _LOW_CONFIDENCE_PATTERN: re.Pattern[str] = re.compile(r"\*\*low\*\*|confidence=\"low\"")
      
      def check_quality_gate(
       staging_path: Path,
       threshold: float = 0.30,
      ) -> tuple[bool, str]:
       """Check the confidence distribution in the staging file.
      
       Returns:
        (passed, message) where:
        - passed: True if low_count / total_count <= threshold
        - message: human-readable summary
       """
       if not staging_path.exists():
        raise FileNotFoundError(f"Staging file not found: {staging_path}")
      
       content: str = staging_path.read_text(encoding="utf-8")
       total_count: int = 0
       low_count: int = 0
       for line in content.splitlines():
        if not line.startswith("|"):
         continue
        if "---" in line or "Date" in line:
         continue
        total_count += 1
        if _LOW_CONFIDENCE_PATTERN.search(line):
         low_count += 1
      
       if total_count == 0:
        return (False, "ABORT: staging file has no rows; nothing to review.")
      
       ratio: float = low_count / total_count
       if ratio > threshold:
        return (
         False,
         f"ABORT: classifier is bad; {low_count}/{total_count} ({ratio:.1%}) rows are ambiguous, exceeds {threshold:.0%} threshold. "
         f"Fallback to manual review (Tier 1 reviews every row).",
        )
      
       return (
        True,
        f"PASS: classifier is good; {low_count}/{total_count} ({ratio:.1%}) rows are low confidence (≤ {threshold:.0%} threshold). "
        f"Proceed to Tier 1 review of 'Needs Review' queue.",
       )
      
      def main() -> int:
       parser: argparse.ArgumentParser = argparse.ArgumentParser(
        description="Quality gate for the chronology classifier (FR7)."
       )
       parser.add_argument(
        "--staging",
        type=Path,
        required=True,
        help="Path to the staging file produced by generate_chronology.py",
       )
       parser.add_argument(
        "--threshold",
        type=float,
        default=0.30,
        help="Maximum allowed ratio of low-confidence rows (default: 0.30)",
       )
       args: argparse.Namespace = parser.parse_args()
      
       try:
        passed, message = check_quality_gate(args.staging, threshold=args.threshold)
       except FileNotFoundError as e:
        print(f"ERROR: {e}", file=sys.stderr)
        return 2
      
       print(message)
       return 0 if passed else 1
      
      if __name__ == "__main__":
       sys.exit(main())
      
    • SAFETY: Use 1-space indentation. The function is a pure transformation: file → (passed, message). No side effects. The CLI uses argparse (standard library). The exit code is 0 for PASS, 1 for ABORT, 2 for errors.
    • VERIFY: uv run python scripts/run_tests_batched.py tests/test_chronology_quality_gate.py should now PASS (all 5 tests green).
    • COMMIT: feat(chronology): add quality gate script (FR7)
    • GIT NOTE: "Phase 4.2. TDD green. chronology_quality_gate.py checks the confidence distribution. Default threshold 0.30 (30%). Exit 0 for PASS, 1 for ABORT, 2 for errors. CLI: --staging PATH --threshold FLOAT."
    • CHECKPOINT: conductor(checkpoint): Phase 4 complete — quality gate (FR7) green

Phase 5: Run the new classifier, generate v2 staging (FR6 Stage 1)

Focus: Run the script; verify the staging file has the per-row evidence + confidence + "Needs Review" section.

  • Task 5.1: Run the script, generate the staging file

    • WHERE: scripts/audit/generate_chronology.pyconductor/chronology.md.staging
    • WHAT: Run the script with the existing --draft flag (from v1) to produce a 216-row staging file.
    • HOW: uv run python scripts/audit/generate_chronology.py --draft > conductor/chronology.md.staging. The script's --draft flag emits the markdown table to stdout; the staging file is the captured output.
    • SAFETY: The staging file is gitignored (verified in Task 1.5). The script does not modify any source folder.
    • VERIFY: Test-Path conductor/chronology.md.staging returns True; (Get-Content conductor/chronology.md.staging | Measure-Object -Line).Lines returns a value > 400 (header + 216 rows × ~2 lines per row + Needs Review section).
    • NO COMMIT (staging file is gitignored).
  • Task 5.2: Sanity-check 5 random rows in the staging file

    • WHERE: conductor/chronology.md.staging
    • WHAT: Pick 5 random rows; for each, verify the per-row evidence line + confidence level + summary are correct. If any field is wrong, the script has a bug — fix the script in a follow-up commit BEFORE Phase 6.
    • HOW: For each picked row, run:
      • Get-Content "conductor/<tracks|archive>/<id>/spec.md" | Select-Object -First 1 (verify summary source; should NOT start with **)
      • git log --oneline --reverse -- "conductor/<tracks|archive>/<id>/" (verify init SHA)
      • git log -1 --format='%h' -- "conductor/<tracks|archive>/<id>/" (verify end SHA)
      • Compare the staging file's evidence line to the git output.
    • SAFETY: Don't proceed to Phase 6 if the script is buggy. Fix the script first.
    • NO COMMIT (sanity check, not implementation).
    • CHECKPOINT: conductor(checkpoint): Phase 5 complete — v2 staging generated

Phase 6: Quality gate (FR7) — HARD GATE

Focus: Run the quality gate; if PASS, proceed to Phase 7 (Tier 1 review). If ABORT, fall back to manual review (the user's "B" option).

  • Task 6.1: Run the quality gate

    • WHERE: scripts/audit/chronology_quality_gate.pyconductor/chronology.md.staging
    • WHAT: uv run python scripts/audit/chronology_quality_gate.py --staging conductor/chronology.md.staging.
    • HOW: Single command. Capture the exit code and message.
    • SAFETY: This is a HARD GATE. The exit code determines the next phase.
    • VERIFY: Exit code is 0 (PASS) or 1 (ABORT). The message contains either "PASS" or "ABORT".
  • Task 6.2a: IF PASS — proceed to Phase 7

    • WHERE: continue to Phase 7
    • WHAT: If the gate passed, the classifier is good. Proceed to Phase 7 (Tier 1 review of the "Needs Review" queue).
    • HOW: Note the PASS message + low/total ratio in the per-row evidence log. Continue.
    • NO COMMIT (gate logic, not implementation).
  • Task 6.2b: IF ABORT — fall back to manual review (the user's "B" option)

    • WHERE: same — Tier 1 reviews every row in the staging file
    • WHAT: If the gate aborted (>30% low confidence), the classifier is bad. Fall back to the v1 manual protocol: Tier 1 opens every row in the staging file, verifies the 5 fields (date, ID, status, summary, range), and assigns a status manually.
    • HOW: Skip Phase 7 (no "Needs Review" queue; Tier 1 reviews all 216 rows). Go directly to Phase 8 (promote staging to canonical). Document the abort + the fallback in the migration report v2 addendum (Phase 9).
    • SAFETY: STOP and report to the user. The fallback is a significant scope change; the user should be aware.
    • NO COMMIT (gate logic).
    • CHECKPOINT: conductor(checkpoint): Phase 6 complete — quality gate (PASS or ABORT)

Phase 7: Tier 1 reviews "Needs Review" queue (FR6 Stage 2) — PAUSE FOR TIER 1

Focus: Tier 1 (not the Tier 2 worker) reviews the low-confidence rows. Tier 2 PAUSES here and reports to Tier 1.

  • Task 7.1: Tier 2 generates the "Needs Review" worklist for Tier 1

    • WHERE: conductor/chronology.md.stagingtests/artifacts/chronology_v2_needs_review_worklist.md (new file, gitignored)
    • WHAT: Tier 2 extracts the low-confidence rows from the staging file and writes a worklist for Tier 1. Each worklist entry has: track_id, current status, current confidence, evidence (git SHAs, commit count, first/last commit subject, state_phase), classifier reason, the source folder path, the source spec.md path, and a placeholder for Tier 1's resolution.
    • HOW: Use manual-slop_run_powershell to extract the "Needs Review" section from the staging file. Format the worklist as a markdown table.
    • SAFETY: The worklist is gitignored. Tier 2 does NOT make resolution decisions; that's Tier 1's job.
    • VERIFY: The worklist has one row per low-confidence track.
    • NO COMMIT (gitignored worklist).
  • Task 7.2: Tier 2 generates the per-row evidence log

    • WHERE: tests/artifacts/chronology_v2_evidence_log.md (new file, gitignored)
    • WHAT: A complete log of all 216 rows with: track_id, status, confidence, init_sha, end_sha, commit_count, first_commit_subject, last_commit_subject, state_phase, classifier_reason, tier1_override (placeholder).
    • HOW: Use manual-slop_run_powershell to extract the staging file's evidence sub-lines. Format as a markdown table.
    • SAFETY: The log is gitignored. Tier 2 does NOT fill in tier1_override; that's Tier 1's job.
    • VERIFY: The log has 216 rows.
    • NO COMMIT (gitignored log).
  • Task 7.3: Tier 2 PAUSES and reports to Tier 1

    • WHERE: report to the user
    • WHAT: Tier 2 reports the quality gate result + the worklist + the evidence log to the user. The user decides who reviews the worklist (Tier 1 or the user directly).
    • HOW: Print a summary:
      Phase 6 quality gate: PASS (low/total ratio = X%)
      Phase 7 "Needs Review" queue: N rows need Tier 1 review.
      Worklist: tests/artifacts/chronology_v2_needs_review_worklist.md
      Evidence log: tests/artifacts/chronology_v2_evidence_log.md
      
      [PAUSE] Awaiting Tier 1 review of the "Needs Review" queue.
      
    • SAFETY: STOP. Do not proceed to Phase 8 until Tier 1 has reviewed the worklist.
    • NO COMMIT (gate).
  • Task 7.4: Tier 1 reviews each low-confidence row (OUTSIDE this plan)

    • WHERE: tests/artifacts/chronology_v2_needs_review_worklist.md (filled in by Tier 1)
    • WHAT: Tier 1 (a different agent or a human) opens the worklist, reads each row's evidence, opens the source folder's spec.md and git log, and assigns a status. Tier 1's resolution is written in the worklist's "Tier 1 Resolution" column.
    • HOW: This task is NOT executed by Tier 2. Tier 1 fills the worklist manually.
    • SAFETY: Tier 1's defaults per FR6: in tracks/ + ambiguous → Active; in archive/ + 0 commits → Special; in archive/ + ≥ 3 work commits + state_phase=0 → Completed; truly ambiguous → Special with note.
    • NO COMMIT (Tier 1's work is gitignored).
  • Task 7.5: Tier 2 applies Tier 1's resolutions to the staging file

    • WHERE: conductor/chronology.md.staging (updated by Tier 2)
    • WHAT: Tier 2 reads the filled worklist and updates the staging file: for each low-confidence row, replace the status with Tier 1's resolved status. The "Needs Review" section is replaced with a "Tier 1 Resolutions" section showing each row's resolution note.
    • HOW: Use manual-slop_run_powershell to parse the worklist and update the staging file.
    • SAFETY: Only low-confidence rows are modified; high-confidence rows are preserved.
    • VERIFY: grep -c "Needs Review" conductor/chronology.md.staging returns 0 (the section is renamed). grep -c "Tier 1 Resolutions" conductor/chronology.md.staging returns 1.
    • NO COMMIT (staging file is gitignored).
    • CHECKPOINT: conductor(checkpoint): Phase 7 complete — Tier 1 review applied

Phase 8: Promote v2 staging → canonical (FR1)

Focus: Rename chronology.md.stagingchronology.md; this is the first time the v2 canonical file is committed.

  • Task 8.1: Move staging to canonical, commit
    • WHERE: conductor/chronology.md.stagingconductor/chronology.md
    • WHAT: git mv conductor/chronology.md.staging conductor/chronology.md. Then commit.
    • HOW: git mv preserves git history (the v1 chronology is in .broken-v1; the v2 chronology is a new file). Verify the rename with git status (should show a rename of .staging to canonical + the removal of .gitignore line if applicable).
    • SAFETY: Verify the file content matches the staging file before commit. Don't add the staging path to .gitignore if it would block the rename.
    • VERIFY: git log -1 --stat shows the rename + the new file. conductor/chronology.md has 216 rows.
    • COMMIT: conductor(chronology): add v2 chronology.md with git-history classifier (FR1)
    • GIT NOTE: "Phase 8. v2 chronology.md promoted from staging to canonical. Uses git-history classifier (FR5) + per-row evidence + confidence level + 'Needs Review' section. v1 is preserved as .broken-v1 (VC14). Cross-check passed the 30% quality gate (VC10) and Tier 1 reviewed the 'Needs Review' queue (VC11)."
    • CHECKPOINT: conductor(checkpoint): Phase 8 complete — v2 chronology.md committed

Phase 9: Write v2 addendum to migration report + end-of-track report (FR4, VC9)

Focus: The v1 migration report gets a v2 addendum documenting the rewrite. The end-of-track report gets a v2 addendum.

  • Task 9.1: Write v2 addendum to migration report

    • WHERE: docs/reports/CHRONOLOGY_MIGRATION_20260619.md (append)
    • WHAT: A new section appended to the end of the v1 report: "v2 Rewrite Addendum (2026-06-21)". Contains:
      • Why the rewrite was needed — link to CHRONOLOGY_TRACK_HANDOVER_20260620.md + 1-paragraph summary
      • v1 → v2 status diff — table of all 216 rows showing the v1 status (stale) and v2 status (after the new classifier) + the git evidence per row. The diff is computed from tests/artifacts/chronology_v2_evidence_log.md.
      • Classifier confidence distribution — counts: high / low / total; % of total in Needs Review
      • Tier 1 review log — for each low-confidence row, the resolution note
      • Quality gate result — PASS or ABORT + the ratio
      • User sign-off — a checklist section for the user to fill in during Phase 10
    • HOW: Use manual-slop_run_powershell to extract the evidence log + format the diff. Hand-write the Tier 1 review log + the sign-off checklist.
    • SAFETY: The addendum does not modify the v1 sections; it only appends.
    • VERIFY: The file is now > 200 lines (v1 was 174; addendum is ~50-100 lines).
    • COMMIT: docs(chronology): write v2 rewrite addendum to migration report (FR4)
    • GIT NOTE: "Phase 9.1. v2 addendum appended to CHRONOLOGY_MIGRATION_20260619.md. Documents the rewrite rationale + v1→v2 status diff + classifier confidence + Tier 1 review log + quality gate result + user sign-off checklist."
  • Task 9.2: Write v2 addendum to end-of-track report

    • WHERE: docs/reports/TRACK_COMPLETION_chronology_20260619.md (append)
    • WHAT: A new section: "v2 Rewrite Addendum (2026-06-21)". Contains:
      • Summary (1-2 sentences)
      • v2 final state (5 fields: chronology.md, tracks.md, workflow.md, migration report, end-of-track report)
      • v2 statistics (rows in v2, classifier confidence, Tier 1 review log size, quality gate result)
      • Cross-check summary (per VC10/11/12/13 confirmation)
      • User sign-off (reference to the migration report v2 addendum)
    • HOW: Hand-write the addendum; reference the v1 sections.
    • VERIFY: The file is updated.
    • COMMIT: docs(chronology): write v2 addendum to end-of-track report (VC9)
    • GIT NOTE: "Phase 9.2. v2 addendum to TRACK_COMPLETION. Documents the v2 final state + cross-check summary."
    • CHECKPOINT: conductor(checkpoint): Phase 9 complete — migration report + end-of-track report updated

Phase 10: User sign-off (FR6 Stage 3) — PAUSE FOR USER

Focus: User is the final quality gate. The track is not "done" until the user signs off.

  • Task 10.1: Tier 2 PAUSES and reports to the user

    • WHERE: report to the user
    • WHAT: Tier 2 reports the v2 state to the user:
      Phase 9 complete. v2 chronology.md committed.
      Quality gate: PASS (low/total ratio = X%)
      Tier 1 review: N rows resolved.
      
      Files:
      - conductor/chronology.md (v2; 216 rows)
      - conductor/chronology.md.broken-v1 (v1 archive; 218 lines)
      - docs/reports/CHRONOLOGY_MIGRATION_20260619.md (v1 + v2 addendum)
      - docs/reports/TRACK_COMPLETION_chronology_20260619.md (v1 + v2 addendum)
      
      [PAUSE] Awaiting user sign-off.
      
    • SAFETY: STOP. Do not proceed to Phase 11 until the user signs off.
    • NO COMMIT (gate).
  • Task 10.2: User reviews v2 + evidence log + Tier 1 resolutions (USER TASK, outside this plan)

    • WHERE: conductor/chronology.md, tests/artifacts/chronology_v2_evidence_log.md, docs/reports/CHRONOLOGY_MIGRATION_20260619.md (v2 addendum)
    • WHAT: User confirms: (a) format correct, (b) every row has evidence, (c) Tier 1 resolutions are reasonable, (d) nothing missed.
    • HOW: User fills the sign-off checklist in the v2 addendum of the migration report.
    • SAFETY: This is a user task; Tier 2 does not execute it.
    • NO COMMIT (gate).
  • Task 10.3: Tier 2 records the user sign-off

    • WHERE: docs/reports/CHRONOLOGY_MIGRATION_20260619.md (v2 addendum)
    • WHAT: Tier 2 records the user's sign-off in the v2 addendum (date + confirmation). If the user requested changes, loop back to the relevant phase.
    • HOW: Use manual-slop_edit_file to add the sign-off line.
    • VERIFY: The v2 addendum has a "User Sign-off" section with a date.
    • COMMIT: docs(chronology): record user sign-off in migration report v2 addendum (VC13)
    • GIT NOTE: "Phase 10.3. User sign-off recorded. All VCs satisfied."
    • CHECKPOINT: conductor(checkpoint): Phase 10 complete — user sign-off recorded

Phase 11: Wrap-up

Focus: Mark the track complete in tracks.md, state.toml, and metadata.json.

  • Task 11.1: Mark track complete in conductor/tracks.md

    • WHERE: conductor/tracks.md (the entry for chronology_20260619 in the "Active Tracks" table or the equivalent)
    • WHAT: Flip the status marker from [ ] or [~] to [x]. Add a one-line note: "Completed 2026-06-21 (v2 rewrite after tier-2 failure report)."
    • HOW: Use manual-slop_edit_file to find the entry and update it.
    • SAFETY: Don't touch other entries.
    • VERIFY: grep -n "chronology_20260619" conductor/tracks.md shows the entry with [x].
    • COMMIT: conductor(track): mark chronology_20260619 as complete in tracks.md
    • GIT NOTE: "Phase 11.1. Track marked complete in tracks.md."
  • Task 11.2: Update state.toml to completed

    • WHERE: conductor/tracks/chronology_20260619/state.toml
    • WHAT: Set [meta].status = "completed", [meta].current_phase = "complete", [meta].last_updated = "<today's date>". Mark all 11 phases as completed with their checkpoint SHAs. Mark all tasks as completed with their commit SHAs. Mark all verification flags as true.
    • HOW: Use manual-slop_edit_file to set the fields. The task commit SHAs come from git log --oneline conductor/tracks/chronology_20260619/.
    • SAFETY: Don't change the task descriptions; just flip the status fields.
    • VERIFY: uv run python -c "import tomllib; data = tomllib.load(open('conductor/tracks/chronology_20260619/state.toml','rb')); assert data['meta']['status'] == 'completed'; print('OK')" exits 0.
    • COMMIT: conductor(plan): mark chronology_20260619 v2 as completed in state.toml
    • GIT NOTE: "Phase 11.2. Track complete. All 11 phases satisfied."
  • Task 11.3: Update metadata.json to completed

    • WHERE: conductor/tracks/chronology_20260619/metadata.json
    • WHAT: Set "status": "completed". Update the verification_criteria array with the v2 VCs (replace v1's 12 VCs with v2's 14 VCs). Update the risk_register to include the v2 risks (replace v1's 9 risks with v2's 10 risks). Add a new regressions_and_pre_existing_failures entry: "v1 chronology had 167 rows with stale status; v2 fixed via git-history classifier (FR5)."
    • HOW: Use manual-slop_edit_file to update the JSON.
    • SAFETY: Keep the JSON valid (run python -c "import json; json.load(open('metadata.json'))" after the edit).
    • VERIFY: uv run python -c "import json; data = json.load(open('conductor/tracks/chronology_20260619/metadata.json')); assert data['status'] == 'completed'; assert len(data['verification_criteria']) >= 14; print('OK')" exits 0.
    • COMMIT: conductor(plan): mark chronology_20260619 v2 as completed in metadata.json
    • GIT NOTE: "Phase 11.3. metadata.json updated with v2 VCs and risk register."
  • Task 11.4: Final end-of-track commit

    • WHERE: working tree
    • WHAT: Verify git status is clean except for the v2 file changes. Print a final summary of the 11 phases + their checkpoints.
    • HOW: git log --oneline conductor/tracks/chronology_20260619/ | head -30 (shows the v1 + v2 commit history).
    • SAFETY: Don't commit unrelated working-tree changes.
    • NO COMMIT (verification).
    • CHECKPOINT: conductor(checkpoint): Phase 11 complete — track fully complete

Summary

Plan Phase Scope Pause point? Spec Phase
1 Archive v1 + verify state + reset state.toml No Spec Phase 1
2 Rewrite _classify_status (TDD red+green) No Spec Phase 2 (part 1)
3 Update extract_summary + format_markdown (TDD red+green) No Spec Phase 2 (part 2)
4 Add quality gate script (TDD red+green) No Spec Phase 3
5 Run classifier, generate staging No Spec Phase 4
6 Quality gate (PASS or ABORT) Yes (HARD) Spec Phase 5
7 Tier 1 reviews "Needs Review" queue Yes (PAUSE for Tier 1)
8 Promote staging → canonical No
9 Write v2 addendum to migration + end-of-track reports No
10 User sign-off Yes (PAUSE for user)
11 Wrap-up (mark complete) No

Total: 11 plan phases (10 spec phases; Plan Phase 2 = Spec Phase 2 part 1, Plan Phase 3 = Spec Phase 2 part 2; renumbering shifts from Plan Phase 4 onwards), 4 hard pause points (Plan Phase 6 quality gate, Plan Phase 7 Tier 1 review, Plan Phase 10 user sign-off, plus the Plan Phase 6 ABORT fallback), ~25 commits. The cross-check is the dominant cost; the user's "EVERY SINGLE ENTRY" mandate is preserved at the semantic level via the per-row evidence log + the 3-stage protocol (FR6).

Verification Criteria Recap

All 14 VCs from the spec must be satisfied for the track to be marked complete:

  • VC1: chronology.md v2 with 216 rows, 5-value status enum, per-row evidence + confidence
  • VC2: tracks.md pruning intact (no regression)
  • VC3: workflow.md 3-step convention present (no regression)
  • VC4: migration report v2 addendum
  • VC5: sorted newest first; Folder + Range + Evidence per row
  • VC6: every folder has a row or documented exception
  • VC7: Notable Non-Track Commits section preserved
  • VC8: no new src/*.py files
  • VC9: end-of-track v2 addendum
  • VC10: quality gate PASS (or ABORT with manual fallback)
  • VC11: "Needs Review" queue resolved by Tier 1
  • VC12: per-row evidence log (tests/artifacts/chronology_v2_evidence_log.md)
  • VC13: user sign-off recorded
  • VC14: v1 archive preserved as .broken-v1

Cross-cutting safety

  • No day estimates (per conductor/workflow.md Tier 1 rule).
  • Per-task atomic commits (per conductor/workflow.md Commit Guidelines).
  • Git notes on every commit (per project convention).
  • No git restore / git checkout -- <file> / git reset (per HARD BAN).
  • No new src/*.py files (per AGENTS.md File Size and Naming Convention).
  • No new conductor/code_styleguides/* files (the 3-step convention is in workflow.md, not a new styleguide).
  • TDD for all script changes (Phases 2-4 are red+green cycles).
  • Tier 2 uses uv run python scripts/run_tests_batched.py (never direct uv run pytest; per conductor/workflow.md "Tier 2 Autonomous Sandbox" rule #1).
  • Throw-away scripts go in scripts/tier2/artifacts/chronology_20260619/ (per conductor/workflow.md "Tier 2 Autonomous Sandbox" rule #4).
  • End-of-track report at docs/reports/TRACK_COMPLETION_chronology_20260619.md (already exists; v2 addendum appended in Phase 9.2).
  • Default branch: master (per conductor/workflow.md "Tier 2 Autonomous Sandbox" rule #2).
  • Preserve line endings on edit (per conductor/workflow.md "Tier 2 Autonomous Sandbox" rule #3; this repo has mixed CRLF/LF).