Private
Public Access
1046 lines
62 KiB
Markdown
1046 lines
62 KiB
Markdown
# Conductor Chronology v2 Implementation Plan
|
||
|
||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||
|
||
**Goal:** Rewrite the chronology track to use a **git-history-based classifier** as the primary evidence (per the handover §"Rewrite `_classify_status` to use git history as primary evidence"). Move the broken v1 `conductor/chronology.md` → `conductor/chronology.md.broken-v1`; generate v2 with the new classifier; add a per-row evidence line + confidence level + "Needs Review" section; enforce a 30% quality gate (FR7) before Tier 1 reviews the "Needs Review" queue; preserve the v1 pruning of `tracks.md` + the 3-step convention in `workflow.md`; write a v2 addendum to the migration report + end-of-track report; get user sign-off.
|
||
|
||
**Architecture:** Rewrite the helper script (`scripts/audit/generate_chronology.py`) to use the 5-step git-history algorithm (FR5). Add a new quality gate script (`scripts/audit/chronology_quality_gate.py`, FR7) that gates the cross-check on the classifier's confidence distribution. The 3-stage cross-check (FR6) is: (1) classifier auto-classifies high-confidence rows, (2) Tier 1 reviews the "Needs Review" queue, (3) user reviews the final v2 with per-row evidence. No CI integration; the file is hand-maintained like `tracks.md`.
|
||
|
||
**Tech Stack:** Python 3.11+ (helper script + quality gate), `tomllib`, `git log` (for commit SHAs), `pathlib`, `re`. No new production code in `src/`. No new dependencies. Test runner: `uv run python scripts/run_tests_batched.py` (per `conductor/workflow.md` "Tier 2 Autonomous Sandbox" rule #1 — NEVER direct `uv run pytest`).
|
||
|
||
**Spec reference:** `conductor/tracks/chronology_20260619/spec.md` (276 lines; 7 FRs, 7 NFRs, 14 VCs, 10 Risks, 10 Phases).
|
||
|
||
**Pause points (Tier 2 must stop and report):**
|
||
- **Phase 5:** Quality gate result. If PASS, continue. If ABORT, fall back to manual review (Tier 1 reviews every row) and STOP.
|
||
- **Phase 6:** Tier 1 review of the "Needs Review" queue. Tier 2 generates the staging file + the per-row evidence log + the queue, then PAUSES for Tier 1.
|
||
- **Phase 9:** User sign-off. Tier 2 generates the v2 staging file with Tier 1's resolutions, then PAUSES for user.
|
||
|
||
**HARD BAN reminders** (per `AGENTS.md` and `conductor/workflow.md`):
|
||
- No `git restore`, `git checkout -- <file>`, `git reset`.
|
||
- No day/hour/minute estimates in any artifact.
|
||
- No new `src/*.py` files.
|
||
- No skipping tests; fix the underlying bug.
|
||
- Per-file atomic commits; no batch commits.
|
||
|
||
---
|
||
|
||
## Phase 1: Archive v1 + verify state of carried-forward work
|
||
|
||
Focus: The v1 work's pruning + 3-step convention + helper script are reused. Move the broken v1 chronology to `.broken-v1`; reset `state.toml` to `current_phase = 0` so this rewrite starts clean.
|
||
|
||
- [ ] **Task 1.1: Verify `tracks.md` pruning is intact**
|
||
- WHERE: `conductor/tracks.md`
|
||
- WHAT: Run `grep -n "^- \[x\]" conductor/tracks.md | wc -l`. Expected: `0` (the only `[x]` markers should be in the Status legend at the bottom, not in a `- [x]` line at the start of a row).
|
||
- HOW: `git diff --name-only HEAD~1 HEAD -- conductor/tracks.md` to see the most recent tracks.md changes (sanity check); then run the grep.
|
||
- SAFETY: If non-zero, STOP and report. The pruning should not have regressed.
|
||
- VERIFY: `grep -n "^- \[x\]" conductor/tracks.md` returns 0 matches in a `- [x]` line position.
|
||
- NO COMMIT (verification only).
|
||
|
||
- [ ] **Task 1.2: Verify `workflow.md` 3-step convention is present**
|
||
- WHERE: `conductor/workflow.md`
|
||
- WHAT: Run `grep -n "Archiving a track" conductor/workflow.md`. Expected: 1 match.
|
||
- HOW: Single grep; verify the line is the 3-step convention block.
|
||
- SAFETY: If 0 matches or > 1, STOP and report. The 3-step convention should be in `workflow.md` "Notes > Editing this file" section.
|
||
- VERIFY: The matched line is followed by the 3 numbered steps.
|
||
- NO COMMIT (verification only).
|
||
|
||
- [ ] **Task 1.3: Move `chronology.md` → `chronology.md.broken-v1`**
|
||
- WHERE: `conductor/chronology.md` → `conductor/chronology.md.broken-v1`
|
||
- WHAT: Use `git mv` to preserve git history (the v1 chronology was committed in commit `8cd92856` and 24 subsequent commits; `git mv` keeps the rename visible in `git log --follow`).
|
||
- HOW: `git mv conductor/chronology.md conductor/chronology.md.broken-v1`. Verify with `git status` (should show a rename, not a delete+add).
|
||
- SAFETY: Don't delete the file; the rename preserves history. Don't add `.broken-v1` to `.gitignore` — it should be a tracked file (per VC14).
|
||
- VERIFY: `git log --follow conductor/chronology.md.broken-v1` shows the v1 history. The file has 218 lines.
|
||
- COMMIT: `conductor(chronology): archive v1 chronology to .broken-v1 (VC14)`
|
||
- GIT NOTE: "Phase 1.3. v1 chronology moved to .broken-v1 for archival. The new v2 file will be created in Phase 7. v1 status classifier was broken (read stale metadata.json.status); v2 uses git history as primary evidence per FR5."
|
||
|
||
- [ ] **Task 1.4: Reset `state.toml` to `current_phase = 0`**
|
||
- WHERE: `conductor/tracks/chronology_20260619/state.toml`
|
||
- WHAT: Set `[meta].status = "active"`, `[meta].current_phase = 0`, `[meta].last_updated = "<today's date YYYY-MM-DD>"`. Reset all `[phases].phase_N.status` to `"pending"`, `checkpointsha` to `""`. Reset all `[tasks].t<N>_<M>.status` to `"pending"`, `commit_sha` to `""`. Reset all `[verification].<flag>` to `false`.
|
||
- HOW: Use `manual-slop_edit_file` to set the 4 fields (status, current_phase, last_updated, and the rewrite note). For phases/tasks/verification, replace the entire sections with the new content. The new content has phases 1-10 marked pending + tasks marked pending + verification flags false.
|
||
- SAFETY: Don't delete the `[meta]`, `[phases]`, `[tasks]`, `[verification]`, `[user_directives_logged]` sections; just reset their values. Add a `[rewrite_log]` section noting the v2 rewrite.
|
||
- VERIFY: `uv run python -c "import tomllib; data = tomllib.load(open('conductor/tracks/chronology_20260619/state.toml','rb')); assert data['meta']['current_phase'] == 0; assert data['meta']['status'] == 'active'; print('OK')"` exits 0.
|
||
- COMMIT: `conductor(plan): reset state.toml to current_phase=0 for v2 rewrite`
|
||
- GIT NOTE: "Phase 1.4. state.toml reset for v2 rewrite. The v1 track's 24 commits remain in git history; this rewrite starts at Phase 1 with a git-history-based classifier (per the handover)."
|
||
|
||
- [ ] **Task 1.5: Verify the staging file path is not committed**
|
||
- WHERE: `conductor/chronology.md.staging` (target of Phase 4)
|
||
- WHAT: Confirm the staging file does not exist yet AND that the path is gitignored (so accidental creation doesn't pollute the diff).
|
||
- HOW: `Test-Path conductor/chronology.md.staging` returns False; `grep -n "chronology.md.staging" .gitignore` returns a match (or the staging file is added to `.gitignore` if not already there).
|
||
- SAFETY: If the staging file exists, STOP and report. The v1 chronology should be in `.broken-v1` only.
|
||
- VERIFY: `Test-Path conductor/chronology.md.staging` returns False; the staging path is in `.gitignore`.
|
||
- NO COMMIT (verification only).
|
||
- CHECKPOINT: `conductor(checkpoint): Phase 1 complete — v1 archived, state.toml reset`
|
||
|
||
---
|
||
|
||
## Phase 2: Rewrite the helper script — git-history classifier (FR5, TDD red+green)
|
||
|
||
Focus: The script's `_classify_status` function is rewritten to use the handover's 5-step algorithm. The new signature returns `(status, confidence, reason)`. This is TDD-required per project convention.
|
||
|
||
- [ ] **Task 2.1: Add failing tests for the new `_classify_status` signature (TDD RED)**
|
||
- WHERE: `tests/test_generate_chronology.py` (modify; add new test class `TestClassifyStatus`)
|
||
- WHAT: 8 unit tests for the new classifier signature. Each test calls `_classify_status(folder_link, init_sha, end_sha, commit_count, first_commit_subject, last_commit_subject, state_phase, metadata_status, last_commit_date)` and asserts the returned `(status, confidence, reason)` tuple.
|
||
- HOW: Add the following tests to the file (1-space indentation per project convention):
|
||
```python
|
||
from scripts.audit.generate_chronology import _classify_status
|
||
|
||
class TestClassifyStatus:
|
||
def test_archive_with_three_plus_work_commits_is_completed_high(self):
|
||
result = _classify_status(
|
||
folder_link="conductor/archive/test_track_20260101",
|
||
init_sha="abc1234", end_sha="def5678", commit_count=5,
|
||
first_commit_subject="feat: add core logic",
|
||
last_commit_subject="fix: edge case",
|
||
state_phase="complete", metadata_status="Completed",
|
||
last_commit_date="2026-01-15",
|
||
)
|
||
assert result == ("Completed", "high", "archive + 5 work commits + state_phase=complete")
|
||
|
||
def test_archive_with_zero_commits_is_special_low(self):
|
||
result = _classify_status(
|
||
folder_link="conductor/archive/empty_track_20260101",
|
||
init_sha="abc1234", end_sha="abc1234", commit_count=0,
|
||
first_commit_subject="chore: init",
|
||
last_commit_subject="chore: init",
|
||
state_phase="no-state-toml", metadata_status="Completed",
|
||
last_commit_date="2026-01-01",
|
||
)
|
||
status, confidence, _ = result
|
||
assert status == "Special"
|
||
assert confidence == "low"
|
||
|
||
def test_tracks_with_state_phase_three_plus_is_in_progress_high(self):
|
||
result = _classify_status(
|
||
folder_link="conductor/tracks/active_track_20260101",
|
||
init_sha="abc1234", end_sha="def5678", commit_count=4,
|
||
first_commit_subject="conductor(track): add initial spec",
|
||
last_commit_subject="feat: implement core",
|
||
state_phase="4", metadata_status="active",
|
||
last_commit_date="2026-01-20",
|
||
)
|
||
assert result == ("In Progress", "high", "tracks + 4 work commits + state_phase=4")
|
||
|
||
def test_tracks_with_zero_to_one_commits_is_active_low(self):
|
||
result = _classify_status(
|
||
folder_link="conductor/tracks/fresh_track_20260101",
|
||
init_sha="abc1234", end_sha="abc1234", commit_count=1,
|
||
first_commit_subject="conductor(track): add initial spec",
|
||
last_commit_subject="conductor(track): add initial spec",
|
||
state_phase="0", metadata_status="spec_written",
|
||
last_commit_date="2026-01-01",
|
||
)
|
||
status, confidence, _ = result
|
||
assert status == "Active"
|
||
assert confidence == "low"
|
||
|
||
def test_metadata_status_abandoned_is_honoured(self):
|
||
result = _classify_status(
|
||
folder_link="conductor/archive/abandoned_20260101",
|
||
init_sha="abc1234", end_sha="def5678", commit_count=2,
|
||
first_commit_subject="conductor(track): add initial spec",
|
||
last_commit_subject="chore: archive move",
|
||
state_phase="no-state-toml", metadata_status="abandoned",
|
||
last_commit_date="2025-06-01",
|
||
)
|
||
status, confidence, _ = result
|
||
assert status == "Abandoned"
|
||
assert confidence == "high"
|
||
|
||
def test_metadata_status_superseded_is_mapped_to_special(self):
|
||
result = _classify_status(
|
||
folder_link="conductor/archive/superseded_20260101",
|
||
init_sha="abc1234", end_sha="def5678", commit_count=2,
|
||
first_commit_subject="conductor(track): add initial spec",
|
||
last_commit_subject="chore: archive move",
|
||
state_phase="no-state-toml", metadata_status="superseded",
|
||
last_commit_date="2025-06-01",
|
||
)
|
||
status, confidence, _ = result
|
||
assert status == "Special"
|
||
assert confidence == "high"
|
||
|
||
def test_metadata_status_contradicts_git_is_low(self):
|
||
result = _classify_status(
|
||
folder_link="conductor/archive/contradict_20260101",
|
||
init_sha="abc1234", end_sha="def5678", commit_count=0,
|
||
first_commit_subject="chore: init",
|
||
last_commit_subject="chore: init",
|
||
state_phase="no-state-toml", metadata_status="Completed",
|
||
last_commit_date="2026-01-01",
|
||
)
|
||
status, confidence, _ = result
|
||
assert status == "Special"
|
||
assert confidence == "low"
|
||
|
||
def test_tracks_with_three_plus_commits_no_state_toml_is_in_progress(self):
|
||
result = _classify_status(
|
||
folder_link="conductor/tracks/no_state_20260101",
|
||
init_sha="abc1234", end_sha="def5678", commit_count=3,
|
||
first_commit_subject="feat: implement X",
|
||
last_commit_subject="fix: bug in X",
|
||
state_phase="no-state-toml", metadata_status="active",
|
||
last_commit_date="2026-01-15",
|
||
)
|
||
assert result == ("In Progress", "high", "tracks + 3 work commits + no state.toml")
|
||
|
||
def test_placeholder_in_track_id_keeps_current_status(self):
|
||
result = _classify_status(
|
||
folder_link="conductor/tracks/PLACEHOLDER_test_20260101",
|
||
init_sha="abc1234", end_sha="def5678", commit_count=3,
|
||
first_commit_subject="chore: init",
|
||
last_commit_subject="chore: update",
|
||
state_phase="2", metadata_status="spec_written",
|
||
last_commit_date="2026-01-10",
|
||
)
|
||
assert result == ("spec_written", "high", "PLACEHOLDER track; original status preserved")
|
||
```
|
||
- SAFETY: These tests use the `_classify_status` function with the new signature. The v1 signature was `_classify_status(folder_link, current, track_id)`. The tests will fail with a TypeError or signature mismatch — that's the RED phase.
|
||
- VERIFY: `uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py` should FAIL with `TypeError: _classify_status() got an unexpected keyword argument 'init_sha'` (or similar signature mismatch).
|
||
- COMMIT: `test(chronology): add 8 failing tests for new _classify_status signature (TDD red)`
|
||
- GIT NOTE: "Phase 2.1. TDD red. 8 tests cover the new (status, confidence, reason) tuple signature with the 5-step algorithm. Tests must fail before Task 2.2 rewrites the function."
|
||
|
||
- [ ] **Task 2.2: Rewrite `_classify_status` to the new signature (TDD GREEN)**
|
||
- WHERE: `scripts/audit/generate_chronology.py` (modify; replace the `_classify_status` function definition; lines ~163-189)
|
||
- WHAT: Replace the v1 `_classify_status(folder_link, current, track_id)` function with the new signature `_classify_status(folder_link, init_sha, end_sha, commit_count, first_commit_subject, last_commit_subject, state_phase, metadata_status, last_commit_date)` that returns `(status, confidence, reason)`.
|
||
- HOW: Use `manual-slop_py_update_definition` to replace the function body. The new function:
|
||
```python
|
||
def _classify_status(
|
||
folder_link: str,
|
||
init_sha: str,
|
||
end_sha: str,
|
||
commit_count: int,
|
||
first_commit_subject: str,
|
||
last_commit_subject: str,
|
||
state_phase: str,
|
||
metadata_status: str,
|
||
last_commit_date: str,
|
||
) -> tuple[str, str, str]:
|
||
"""Classify a track's status using git history as primary evidence.
|
||
|
||
The 5-step algorithm (per CHRONOLOGY_TRACK_HANDOVER_20260620.md):
|
||
1. Count meaningful commits (commit_count)
|
||
2. Inspect commit messages (first/last subject)
|
||
3. Check state.toml phase progression (state_phase as string from _parse_state_phase)
|
||
4. Default to conservative (low confidence on ambiguous)
|
||
5. Honour explicit metadata (metadata_status)
|
||
|
||
state_phase values (per _parse_state_phase):
|
||
- "complete" = track shipped
|
||
- "N" (digit string) = current_phase = N
|
||
- "no-state-toml" = no state.toml file
|
||
- "?" = parse error
|
||
|
||
Returns: (status, confidence, reason) where:
|
||
- status: one of "Active", "In Progress", "Completed", "Abandoned", "Special"
|
||
- confidence: "high" or "low"
|
||
- reason: one-line explanation
|
||
"""
|
||
track_id: str = Path(folder_link).name
|
||
|
||
if "PLACEHOLDER" in track_id:
|
||
return (metadata_status or "unknown", "high", "PLACEHOLDER track; original status preserved")
|
||
|
||
is_archive: bool = folder_link.startswith("conductor/archive/")
|
||
is_tracks: bool = folder_link.startswith("conductor/tracks/")
|
||
|
||
work_commit_pattern: re.Pattern[str] = re.compile(r"^(feat|fix|refactor|perf|test)\(")
|
||
is_work_subject: bool = (
|
||
work_commit_pattern.match(first_commit_subject) is not None
|
||
or work_commit_pattern.match(last_commit_subject) is not None
|
||
)
|
||
|
||
metadata_lower: str = (metadata_status or "").lower()
|
||
state_phase_clean: str = (state_phase or "").strip()
|
||
is_in_flight: bool = state_phase_clean.isdigit() and int(state_phase_clean) >= 3
|
||
|
||
if metadata_lower in {"abandoned", "superseded"} and not is_work_subject:
|
||
mapped: str = "Abandoned" if metadata_lower == "abandoned" else "Special"
|
||
return (mapped, "high", f"metadata.json.status={metadata_lower}; no work commits; mapped to {mapped}")
|
||
|
||
if is_archive and commit_count >= 3 and (is_work_subject or state_phase_clean == "complete"):
|
||
return ("Completed", "high", f"archive + {commit_count} work commits + state_phase={state_phase_clean}")
|
||
|
||
if is_archive and commit_count == 0:
|
||
return ("Special", "low", f"archive folder with 0 work commits; needs review")
|
||
|
||
if is_tracks and is_in_flight:
|
||
return ("In Progress", "high", f"tracks + {commit_count} work commits + state_phase={state_phase_clean}")
|
||
|
||
if is_tracks and commit_count >= 3 and state_phase_clean in {"no-state-toml", "?"}:
|
||
return ("In Progress", "high", f"tracks + {commit_count} work commits + no state.toml")
|
||
|
||
if is_tracks and commit_count <= 1:
|
||
return ("Active", "low", f"tracks + {commit_count} commits; spec-only or early stage")
|
||
|
||
if metadata_lower in {"completed", "complete", "shipped"} and commit_count == 0:
|
||
return ("Special", "low", f"metadata says {metadata_status} but git shows 0 commits; contradictory")
|
||
|
||
return ("Special", "low", f"ambiguous: commit_count={commit_count}, state_phase={state_phase_clean}, is_work={is_work_subject}")
|
||
```
|
||
- SAFETY: Use 1-space indentation (per project convention). Use `tuple[str, str, str]` return type (per `conductor/code_styleguides/error_handling.md` data-oriented pattern; NOT `Optional[Tuple]`). The function is a pure transformation: inputs → outputs, no side effects, no exceptions. The `state_phase` is always a string (per `_parse_state_phase`); the check uses `.isdigit()` + `int()` conversion.
|
||
- VERIFY: `uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py` should now PASS for the 8 new tests in `TestClassifyStatus` (the 6 v1 tests in `TestExtract*` may also need updates — see Task 2.3).
|
||
- COMMIT: `feat(chronology): rewrite _classify_status to git-history classifier (FR5)`
|
||
- GIT NOTE: "Phase 2.2. TDD green. _classify_status now returns (status, confidence, reason) using the 5-step algorithm. The new signature accepts 9 parameters (folder_link, init/end_sha, commit_count, first/last_commit_subject, state_phase, metadata_status, last_commit_date)."
|
||
|
||
- [ ] **Task 2.3: Update `walk_track_folders` to compute the new per-row evidence fields**
|
||
- WHERE: `scripts/audit/generate_chronology.py` (modify; the `walk_track_folders` function)
|
||
- WHAT: `walk_track_folders` currently populates 8 row fields (date, track_id, status, summary, init_sha, end_sha, commit_count, folder_link). Add 6 new fields: `first_commit_subject`, `last_commit_subject`, `state_phase`, `metadata_status`, `last_commit_date`, `confidence`, `reason`. The `init_sha`/`end_sha`/`commit_count` already exist. Update the call to `_classify_status` to use the new 9-argument signature.
|
||
- HOW: Use `manual-slop_py_update_definition` to update the function. The relevant changes:
|
||
```python
|
||
# Existing: extract init_sha, end_sha, commit_count via _git_first_line / _git_log
|
||
# Add: extract first_commit_subject + last_commit_subject
|
||
first_commit_subject: str = _git_first_line(folder_relpath, "--reverse", "--format=%s")
|
||
last_commit_subject: str = _git_first_line(folder_relpath, "-1", "--format=%s")
|
||
|
||
# Add: extract state_phase (string from _parse_state_phase)
|
||
state_phase: str = _parse_state_phase(folder / "state.toml")
|
||
|
||
# Add: extract metadata_status (read from metadata.json)
|
||
metadata_status: str = ""
|
||
if metadata_path.is_file():
|
||
try:
|
||
data = json.loads(metadata_path.read_text(encoding="utf-8"))
|
||
metadata_status = str(data.get("status", "")).strip()
|
||
except (json.JSONDecodeError, OSError):
|
||
pass
|
||
|
||
# Add: extract last_commit_date
|
||
last_commit_date: str = _last_commit_date(folder_relpath)
|
||
|
||
# Update: call _classify_status with the new signature
|
||
status, confidence, reason = _classify_status(
|
||
folder_link=folder_link,
|
||
init_sha=init_sha,
|
||
end_sha=end_sha,
|
||
commit_count=commit_count,
|
||
first_commit_subject=first_commit_subject,
|
||
last_commit_subject=last_commit_subject,
|
||
state_phase=state_phase,
|
||
metadata_status=metadata_status,
|
||
last_commit_date=last_commit_date,
|
||
)
|
||
|
||
# Update: row dict includes the 6 new fields
|
||
rows.append({
|
||
"date": date,
|
||
"track_id": track_id,
|
||
"status": status,
|
||
"confidence": confidence,
|
||
"summary": summary,
|
||
"init_sha": init_sha,
|
||
"end_sha": end_sha,
|
||
"commit_count": commit_count,
|
||
"folder_link": folder_link,
|
||
"first_commit_subject": first_commit_subject,
|
||
"last_commit_subject": last_commit_subject,
|
||
"state_phase": state_phase,
|
||
"metadata_status": metadata_status,
|
||
"last_commit_date": last_commit_date,
|
||
"reason": reason,
|
||
})
|
||
```
|
||
- SAFETY: The existing v1 tests may break because the row dict's keys change. Run the full test file to see what breaks. The 6 v1 tests target `extract_slug_date` and `extract_summary` (pure functions), so they should still pass.
|
||
- VERIFY: `uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py` — the 8 new tests should PASS; the 6 v1 tests should still PASS (no changes to their target functions).
|
||
- COMMIT: `refactor(chronology): update walk_track_folders to populate per-row evidence fields`
|
||
- GIT NOTE: "Phase 2.3. walk_track_folders now populates 6 new fields per row (first/last_commit_subject, state_phase, metadata_status, last_commit_date, confidence, reason). init_sha/end_sha/commit_count were already present."
|
||
|
||
- [ ] **Task 2.4: Update the v1 tests to match the new row dict shape**
|
||
- WHERE: `tests/test_generate_chronology.py` (modify; the 6 v1 tests in `TestExtractSlugDate` + `TestExtractSummary`)
|
||
- WHAT: The 6 v1 tests should still pass because they only test `extract_slug_date` and `extract_summary` (pure functions). The new tests for `_classify_status` and `walk_track_folders` are in `TestClassifyStatus`. Run the full test file to confirm.
|
||
- HOW: Re-read the 6 v1 tests. If they fail due to the row dict change, update them. Otherwise no changes needed.
|
||
- SAFETY: Don't delete the 6 v1 tests; they verify `extract_slug_date` + `extract_summary` which are still used.
|
||
- VERIFY: `uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py` — all tests (6 v1 + 8 new) PASS.
|
||
- NO COMMIT (if no changes) or COMMIT: `test(chronology): update v1 tests for new row dict shape (FR5)`
|
||
- GIT NOTE: "Phase 2.4. v1 tests verified against new row dict shape; no changes needed."
|
||
- CHECKPOINT: `conductor(checkpoint): Phase 2 complete — git-history classifier (FR5) green`
|
||
|
||
---
|
||
|
||
## Phase 3: Update `extract_summary` + add per-row evidence line to `format_markdown` (FR1, FR5)
|
||
|
||
Focus: The summary priority chain needs to reject metadata-field text (the v1 bug). The markdown output needs a per-row evidence line.
|
||
|
||
- [ ] **Task 3.1: Add failing tests for metadata-field text rejection (TDD RED)**
|
||
- WHERE: `tests/test_generate_chronology.py` (modify; add 4 new tests to `TestExtractSummary`)
|
||
- WHAT: 4 unit tests that the regex `^\*\*` rejects metadata-field text. The tests use a fixture folder with `spec.md` whose first non-empty line is `**Priority:** A (foundational; precedes X)`.
|
||
- HOW: Add the following tests:
|
||
```python
|
||
class TestExtractSummaryRejectsMetadata:
|
||
def test_rejects_priority_metadata_line(self, tmp_path: Path):
|
||
(tmp_path / "spec.md").write_text(
|
||
"# Test Track\n\n**Priority:** A (foundational; precedes X)\n\nThis is the actual summary.",
|
||
encoding="utf-8",
|
||
)
|
||
summary: str = extract_summary(tmp_path)
|
||
assert not summary.startswith("**")
|
||
assert "actual summary" in summary
|
||
|
||
def test_rejects_date_metadata_line(self, tmp_path: Path):
|
||
(tmp_path / "spec.md").write_text(
|
||
"# Test Track\n\n**Date:** 2026-01-01\n\nThe actual summary content.",
|
||
encoding="utf-8",
|
||
)
|
||
summary: str = extract_summary(tmp_path)
|
||
assert not summary.startswith("**")
|
||
assert "actual summary" in summary
|
||
|
||
def test_rejects_initialized_metadata_line(self, tmp_path: Path):
|
||
(tmp_path / "spec.md").write_text(
|
||
"# Test Track\n\n**Initialized:** 2026-01-01\n\nThe actual summary content.",
|
||
encoding="utf-8",
|
||
)
|
||
summary: str = extract_summary(tmp_path)
|
||
assert not summary.startswith("**")
|
||
assert "actual summary" in summary
|
||
|
||
def test_rejects_confidence_metadata_line(self, tmp_path: Path):
|
||
(tmp_path / "spec.md").write_text(
|
||
"# Test Track\n\n**Confidence:** high\n\nThe actual summary content.",
|
||
encoding="utf-8",
|
||
)
|
||
summary: str = extract_summary(tmp_path)
|
||
assert not summary.startswith("**")
|
||
assert "actual summary" in summary
|
||
```
|
||
- SAFETY: These tests will FAIL with the v1 `extract_summary` because v1 returns the first non-empty line, which is `**Priority:** A...`. RED phase.
|
||
- VERIFY: `uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py::TestExtractSummaryRejectsMetadata` should FAIL.
|
||
- COMMIT: `test(chronology): add 4 failing tests for metadata-field text rejection (TDD red)`
|
||
- GIT NOTE: "Phase 3.1. TDD red. 4 tests verify that extract_summary skips metadata-field lines (Priority, Date, Initialized, Confidence). Tests must fail before Task 3.2 updates the function."
|
||
|
||
- [ ] **Task 3.2: Update `extract_summary` to reject metadata-field text (TDD GREEN)**
|
||
- WHERE: `scripts/audit/generate_chronology.py` (modify; the `extract_summary` function)
|
||
- WHAT: Add a regex check at the start of each candidate line: if the line starts with `**` (i.e., matches `re.match(r"^\*\*", line)`), skip it. Try the next non-empty line.
|
||
- HOW: Use `manual-slop_py_update_definition` to update the function. The new implementation adds a `_METADATA_LINE_RE = re.compile(r"^\*\*")` constant and filters candidate lines:
|
||
```python
|
||
_METADATA_LINE_RE: re.Pattern[str] = re.compile(r"^\*\*")
|
||
|
||
def extract_summary(folder_path: Path) -> str:
|
||
"""Extract a one-sentence summary from a track folder.
|
||
|
||
Priority chain (each step rejects metadata-field text starting with '**'):
|
||
1. metadata.json.summary if present and not metadata-field text
|
||
2. First non-empty line of spec.md that does not start with '**'
|
||
3. metadata.json.description if not metadata-field text
|
||
4. First non-empty line of plan.md that does not start with '**'
|
||
5. Generic placeholder based on folder location
|
||
"""
|
||
metadata_path: Path = folder_path / "metadata.json"
|
||
if metadata_path.is_file():
|
||
try:
|
||
data: dict = json.loads(metadata_path.read_text(encoding="utf-8"))
|
||
except (json.JSONDecodeError, OSError):
|
||
data = {}
|
||
summary: str | None = data.get("summary")
|
||
if summary and not _METADATA_LINE_RE.match(summary):
|
||
return _truncate_to_25_words(summary)
|
||
description: str | None = data.get("description")
|
||
if description and not _METADATA_LINE_RE.match(description):
|
||
return _truncate_to_25_words(description)
|
||
|
||
for filename in ("spec.md", "plan.md"):
|
||
file_path: Path = folder_path / filename
|
||
if not file_path.is_file():
|
||
continue
|
||
try:
|
||
text: str = file_path.read_text(encoding="utf-8")
|
||
except OSError:
|
||
continue
|
||
for line in text.splitlines():
|
||
stripped: str = line.strip()
|
||
if not stripped or _METADATA_LINE_RE.match(stripped):
|
||
continue
|
||
if stripped.startswith("#"):
|
||
continue
|
||
if stripped.startswith(">"):
|
||
continue
|
||
return _truncate_to_25_words(_first_sentence(stripped))
|
||
|
||
if (folder_path.parent.name if folder_path.parent else "") == "archive":
|
||
return "Imported from archive (no spec)"
|
||
return "Track folder (no spec found)"
|
||
```
|
||
- SAFETY: Use 1-space indentation. The function remains a pure transformation; no side effects. The `json.JSONDecodeError` is caught at the SDK boundary (per `conductor/code_styleguides/error_handling.md` "Exceptions reserved for the SDK boundary") and converted to a recoverable case (empty `data` dict).
|
||
- VERIFY: `uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py` — all 18 tests (6 v1 + 8 new classifier + 4 metadata-rejection) PASS.
|
||
- COMMIT: `feat(chronology): extract_summary rejects metadata-field text (FR5)`
|
||
- GIT NOTE: "Phase 3.2. TDD green. extract_summary now skips lines starting with '**' (Priority, Date, Initialized, Confidence, etc.). Fixes the v1 bug where 167 rows had metadata-field text as their summary."
|
||
|
||
- [ ] **Task 3.3: Add failing tests for per-row evidence line in `format_markdown` (TDD RED)**
|
||
- WHERE: `tests/test_generate_chronology.py` (modify; add new test class `TestFormatMarkdown`)
|
||
- WHAT: 3 unit tests for the per-row evidence line. Each test creates a row dict with the new fields and asserts that the output markdown has the evidence sub-line.
|
||
- HOW: Add the following tests:
|
||
```python
|
||
class TestFormatMarkdown:
|
||
def test_row_includes_evidence_subline(self):
|
||
rows: list[dict] = [
|
||
{
|
||
"date": "2026-01-01",
|
||
"track_id": "test_track_20260101",
|
||
"status": "Completed",
|
||
"confidence": "high",
|
||
"summary": "Test summary.",
|
||
"folder_link": "conductor/archive/test_track_20260101",
|
||
"init_sha": "abc1234",
|
||
"end_sha": "def5678",
|
||
"commit_count": 5,
|
||
"state_phase": "complete",
|
||
"first_commit_subject": "feat: add core",
|
||
"last_commit_subject": "fix: edge case",
|
||
"reason": "archive + 5 work commits + state_phase=complete",
|
||
},
|
||
]
|
||
output: str = format_markdown(rows)
|
||
assert "abc1234..def5678" in output
|
||
assert "Evidence:" in output
|
||
assert "feat: add core" in output
|
||
assert "fix: edge case" in output
|
||
assert "confidence=high" in output
|
||
|
||
def test_row_includes_confidence_in_table(self):
|
||
rows: list[dict] = [
|
||
{
|
||
"date": "2026-01-01",
|
||
"track_id": "test_track_20260101",
|
||
"status": "Special",
|
||
"confidence": "low",
|
||
"summary": "Test summary.",
|
||
"folder_link": "conductor/archive/test_track_20260101",
|
||
"init_sha": "abc1234",
|
||
"end_sha": "abc1234",
|
||
"commit_count": 0,
|
||
"state_phase": None,
|
||
"first_commit_subject": "chore: init",
|
||
"last_commit_subject": "chore: init",
|
||
"reason": "archive folder with 0 work commits",
|
||
},
|
||
]
|
||
output: str = format_markdown(rows)
|
||
assert "confidence=low" in output
|
||
|
||
def test_needs_review_section_appended(self):
|
||
rows: list[dict] = [
|
||
{
|
||
"date": "2026-01-01",
|
||
"track_id": "low_conf_20260101",
|
||
"status": "Special",
|
||
"confidence": "low",
|
||
"summary": "Low confidence row.",
|
||
"folder_link": "conductor/archive/low_conf_20260101",
|
||
"init_sha": "abc1234",
|
||
"end_sha": "abc1234",
|
||
"commit_count": 0,
|
||
"state_phase": None,
|
||
"first_commit_subject": "chore: init",
|
||
"last_commit_subject": "chore: init",
|
||
"reason": "archive folder with 0 work commits",
|
||
},
|
||
]
|
||
output: str = format_markdown(rows)
|
||
assert "## Needs Review" in output
|
||
assert "low_conf_20260101" in output
|
||
```
|
||
- SAFETY: These tests will FAIL with the v1 `format_markdown` because v1 doesn't emit evidence lines or a "Needs Review" section. RED phase.
|
||
- VERIFY: `uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py::TestFormatMarkdown` should FAIL.
|
||
- COMMIT: `test(chronology): add 3 failing tests for per-row evidence line (TDD red)`
|
||
- GIT NOTE: "Phase 3.3. TDD red. 3 tests verify the per-row evidence line, confidence in table, and 'Needs Review' section. Tests must fail before Task 3.4 updates format_markdown."
|
||
|
||
- [ ] **Task 3.4: Update `format_markdown` to emit evidence + "Needs Review" section (TDD GREEN)**
|
||
- WHERE: `scripts/audit/generate_chronology.py` (modify; the `format_markdown` function)
|
||
- WHAT: Update the table to include the confidence column. Add a sub-line below each row with the evidence. Add a "Needs Review" section at the bottom listing `low`-confidence rows.
|
||
- HOW: Use `manual-slop_py_update_definition` to update the function:
|
||
```python
|
||
def format_markdown(rows: list[dict]) -> str:
|
||
lines: list[str] = [
|
||
"| Date | ID | Status | Confidence | Summary | Folder | Range |",
|
||
"| --- | --- | --- | --- | --- | --- | --- |",
|
||
]
|
||
for row in rows:
|
||
range_str: str = f"`{row['init_sha']}..{row['end_sha']}` ({row['commit_count']})"
|
||
confidence_marker: str = f"**{row['confidence']}**" if row["confidence"] == "low" else row["confidence"]
|
||
lines.append(
|
||
f"| {row['date']} | `{row['track_id']}` | {row['status']} | {confidence_marker} | "
|
||
f"{_md_escape(row['summary'])} | `{row['folder_link']}` | {range_str} |"
|
||
)
|
||
evidence_line: str = (
|
||
f"| | | | | | "
|
||
f"Evidence: `{row['init_sha'][:7]}..{row['end_sha'][:7]}` | "
|
||
f"{row['commit_count']} commits | state_phase={row.get('state_phase', 'n/a')} | "
|
||
f"\"{row['first_commit_subject']}\" → \"{row['last_commit_subject']}\" | "
|
||
f"reason=\"{row['reason']}\" |"
|
||
)
|
||
lines.append(evidence_line)
|
||
|
||
low_rows: list[dict] = [r for r in rows if r.get("confidence") == "low"]
|
||
if low_rows:
|
||
lines.append("")
|
||
lines.append("## Needs Review (Tier 1 + User)")
|
||
lines.append("")
|
||
lines.append(f"{len(low_rows)} row(s) had ambiguous git evidence. Resolved by Tier 1; user reviewed in Stage 3.")
|
||
lines.append("")
|
||
for row in low_rows:
|
||
lines.append(
|
||
f"- `{row['track_id']}` (status={row['status']}) — {row['reason']} — resolved by Tier 1"
|
||
)
|
||
|
||
return "\n".join(lines) + "\n"
|
||
```
|
||
- SAFETY: Use 1-space indentation. The function is a pure transformation. The "Needs Review" section only appears if there are `low`-confidence rows. The evidence line is a sub-line in the table (not a column), keeping the table at 7 columns (was 6; +1 for Confidence).
|
||
- VERIFY: `uv run python scripts/run_tests_batched.py tests/test_generate_chronology.py` — all 21 tests (6 v1 + 8 classifier + 4 metadata-rejection + 3 format) PASS.
|
||
- COMMIT: `feat(chronology): format_markdown emits per-row evidence + Needs Review section (FR1)`
|
||
- GIT NOTE: "Phase 3.4. TDD green. format_markdown now emits a 7-column table (added Confidence) with a per-row evidence sub-line and a 'Needs Review' section listing low-confidence rows. v1 had a 6-column table with no evidence."
|
||
- CHECKPOINT: `conductor(checkpoint): Phase 3 complete — summary regex + per-row evidence (FR1, FR5)`
|
||
|
||
---
|
||
|
||
## Phase 4: Add the quality gate script (FR7, TDD)
|
||
|
||
Focus: A new wrapper script that runs after the classifier's output and gates the cross-check on the confidence distribution.
|
||
|
||
- [ ] **Task 4.1: Add failing tests for the quality gate (TDD RED)**
|
||
- WHERE: `tests/test_chronology_quality_gate.py` (new file)
|
||
- WHAT: 5 unit tests for the quality gate logic. Each test calls `check_quality_gate(staging_path, threshold=0.30)` and asserts the return code + message.
|
||
- HOW: Create the new file:
|
||
```python
|
||
"""Tests for scripts/audit/chronology_quality_gate.py."""
|
||
from pathlib import Path
|
||
import pytest
|
||
from scripts.audit.chronology_quality_gate import check_quality_gate
|
||
|
||
def _write_staging(path: Path, low_count: int, high_count: int) -> Path:
|
||
lines: list[str] = [
|
||
"| Date | ID | Status | Confidence | Summary | Folder | Range |",
|
||
"| --- | --- | --- | --- | --- | --- | --- |",
|
||
]
|
||
for i in range(high_count):
|
||
lines.append(
|
||
f"| 2026-01-{i+1:02d} | `track_{i}_20260101` | Completed | high | summary | folder | `abc..def` (5) |"
|
||
)
|
||
for i in range(low_count):
|
||
lines.append(
|
||
f"| 2026-01-{i+1:02d} | `low_{i}_20260101` | Special | **low** | summary | folder | `abc..abc` (0) |"
|
||
)
|
||
path.write_text("\n".join(lines), encoding="utf-8")
|
||
return path
|
||
|
||
def test_zero_percent_low_passes(tmp_path: Path):
|
||
staging: Path = _write_staging(tmp_path / "staging.md", low_count=0, high_count=100)
|
||
passed, message = check_quality_gate(staging, threshold=0.30)
|
||
assert passed is True
|
||
assert "PASS" in message
|
||
|
||
def test_thirty_percent_low_at_boundary_passes(tmp_path: Path):
|
||
staging: Path = _write_staging(tmp_path / "staging.md", low_count=30, high_count=70)
|
||
passed, _ = check_quality_gate(staging, threshold=0.30)
|
||
assert passed is True
|
||
|
||
def test_thirty_one_percent_low_fails(tmp_path: Path):
|
||
staging: Path = _write_staging(tmp_path / "staging.md", low_count=31, high_count=69)
|
||
passed, message = check_quality_gate(staging, threshold=0.30)
|
||
assert passed is False
|
||
assert "ABORT" in message
|
||
|
||
def test_hundred_percent_low_fails(tmp_path: Path):
|
||
staging: Path = _write_staging(tmp_path / "staging.md", low_count=10, high_count=0)
|
||
passed, message = check_quality_gate(staging, threshold=0.30)
|
||
assert passed is False
|
||
assert "ABORT" in message
|
||
|
||
def test_missing_staging_file_errors(tmp_path: Path):
|
||
with pytest.raises(FileNotFoundError):
|
||
check_quality_gate(tmp_path / "missing.md", threshold=0.30)
|
||
```
|
||
- SAFETY: The test file is new; it does not modify any existing file. Use 1-space indentation. Use `pytest.raises` for the error case. The `_write_staging` helper creates a minimal staging file with the expected 7-column header.
|
||
- VERIFY: `uv run python scripts/run_tests_batched.py tests/test_chronology_quality_gate.py` should FAIL with `ModuleNotFoundError: No module named 'scripts.audit.chronology_quality_gate'`.
|
||
- COMMIT: `test(chronology): add 5 failing tests for quality gate (TDD red)`
|
||
- GIT NOTE: "Phase 4.1. TDD red. 5 tests cover the 30% threshold logic: 0% passes, 30% at boundary passes, 31% fails, 100% fails, missing file errors. Tests must fail before Task 4.2 writes the script."
|
||
|
||
- [ ] **Task 4.2: Write the quality gate script (TDD GREEN)**
|
||
- WHERE: `scripts/audit/chronology_quality_gate.py` (new file)
|
||
- WHAT: The wrapper script. Exposes `check_quality_gate(staging_path, threshold)` for programmatic use and a `main()` for CLI use (`--staging PATH --threshold FLOAT`).
|
||
- HOW: Create the new file:
|
||
```python
|
||
"""Quality gate for the chronology classifier (FR7).
|
||
|
||
Reads the staging file produced by generate_chronology.py and checks the
|
||
confidence distribution. If the ratio of low-confidence rows to total rows
|
||
exceeds the threshold (default 0.30), the gate aborts and the user must
|
||
fall back to manual review (Tier 1 reviews every row).
|
||
"""
|
||
from __future__ import annotations
|
||
import argparse
|
||
import re
|
||
import sys
|
||
from pathlib import Path
|
||
|
||
_LOW_CONFIDENCE_PATTERN: re.Pattern[str] = re.compile(r"\*\*low\*\*|confidence=\"low\"")
|
||
|
||
def check_quality_gate(
|
||
staging_path: Path,
|
||
threshold: float = 0.30,
|
||
) -> tuple[bool, str]:
|
||
"""Check the confidence distribution in the staging file.
|
||
|
||
Returns:
|
||
(passed, message) where:
|
||
- passed: True if low_count / total_count <= threshold
|
||
- message: human-readable summary
|
||
"""
|
||
if not staging_path.exists():
|
||
raise FileNotFoundError(f"Staging file not found: {staging_path}")
|
||
|
||
content: str = staging_path.read_text(encoding="utf-8")
|
||
total_count: int = 0
|
||
low_count: int = 0
|
||
for line in content.splitlines():
|
||
if not line.startswith("|"):
|
||
continue
|
||
if "---" in line or "Date" in line:
|
||
continue
|
||
total_count += 1
|
||
if _LOW_CONFIDENCE_PATTERN.search(line):
|
||
low_count += 1
|
||
|
||
if total_count == 0:
|
||
return (False, "ABORT: staging file has no rows; nothing to review.")
|
||
|
||
ratio: float = low_count / total_count
|
||
if ratio > threshold:
|
||
return (
|
||
False,
|
||
f"ABORT: classifier is bad; {low_count}/{total_count} ({ratio:.1%}) rows are ambiguous, exceeds {threshold:.0%} threshold. "
|
||
f"Fallback to manual review (Tier 1 reviews every row).",
|
||
)
|
||
|
||
return (
|
||
True,
|
||
f"PASS: classifier is good; {low_count}/{total_count} ({ratio:.1%}) rows are low confidence (≤ {threshold:.0%} threshold). "
|
||
f"Proceed to Tier 1 review of 'Needs Review' queue.",
|
||
)
|
||
|
||
def main() -> int:
|
||
parser: argparse.ArgumentParser = argparse.ArgumentParser(
|
||
description="Quality gate for the chronology classifier (FR7)."
|
||
)
|
||
parser.add_argument(
|
||
"--staging",
|
||
type=Path,
|
||
required=True,
|
||
help="Path to the staging file produced by generate_chronology.py",
|
||
)
|
||
parser.add_argument(
|
||
"--threshold",
|
||
type=float,
|
||
default=0.30,
|
||
help="Maximum allowed ratio of low-confidence rows (default: 0.30)",
|
||
)
|
||
args: argparse.Namespace = parser.parse_args()
|
||
|
||
try:
|
||
passed, message = check_quality_gate(args.staging, threshold=args.threshold)
|
||
except FileNotFoundError as e:
|
||
print(f"ERROR: {e}", file=sys.stderr)
|
||
return 2
|
||
|
||
print(message)
|
||
return 0 if passed else 1
|
||
|
||
if __name__ == "__main__":
|
||
sys.exit(main())
|
||
```
|
||
- SAFETY: Use 1-space indentation. The function is a pure transformation: file → (passed, message). No side effects. The CLI uses `argparse` (standard library). The exit code is 0 for PASS, 1 for ABORT, 2 for errors.
|
||
- VERIFY: `uv run python scripts/run_tests_batched.py tests/test_chronology_quality_gate.py` should now PASS (all 5 tests green).
|
||
- COMMIT: `feat(chronology): add quality gate script (FR7)`
|
||
- GIT NOTE: "Phase 4.2. TDD green. chronology_quality_gate.py checks the confidence distribution. Default threshold 0.30 (30%). Exit 0 for PASS, 1 for ABORT, 2 for errors. CLI: --staging PATH --threshold FLOAT."
|
||
- CHECKPOINT: `conductor(checkpoint): Phase 4 complete — quality gate (FR7) green`
|
||
|
||
---
|
||
|
||
## Phase 5: Run the new classifier, generate v2 staging (FR6 Stage 1)
|
||
|
||
Focus: Run the script; verify the staging file has the per-row evidence + confidence + "Needs Review" section.
|
||
|
||
- [ ] **Task 5.1: Run the script, generate the staging file**
|
||
- WHERE: `scripts/audit/generate_chronology.py` → `conductor/chronology.md.staging`
|
||
- WHAT: Run the script with the existing `--draft` flag (from v1) to produce a 216-row staging file.
|
||
- HOW: `uv run python scripts/audit/generate_chronology.py --draft > conductor/chronology.md.staging`. The script's `--draft` flag emits the markdown table to stdout; the staging file is the captured output.
|
||
- SAFETY: The staging file is gitignored (verified in Task 1.5). The script does not modify any source folder.
|
||
- VERIFY: `Test-Path conductor/chronology.md.staging` returns True; `(Get-Content conductor/chronology.md.staging | Measure-Object -Line).Lines` returns a value > 400 (header + 216 rows × ~2 lines per row + Needs Review section).
|
||
- NO COMMIT (staging file is gitignored).
|
||
|
||
- [ ] **Task 5.2: Sanity-check 5 random rows in the staging file**
|
||
- WHERE: `conductor/chronology.md.staging`
|
||
- WHAT: Pick 5 random rows; for each, verify the per-row evidence line + confidence level + summary are correct. If any field is wrong, the script has a bug — fix the script in a follow-up commit BEFORE Phase 6.
|
||
- HOW: For each picked row, run:
|
||
- `Get-Content "conductor/<tracks|archive>/<id>/spec.md" | Select-Object -First 1` (verify summary source; should NOT start with `**`)
|
||
- `git log --oneline --reverse -- "conductor/<tracks|archive>/<id>/"` (verify init SHA)
|
||
- `git log -1 --format='%h' -- "conductor/<tracks|archive>/<id>/"` (verify end SHA)
|
||
- Compare the staging file's evidence line to the git output.
|
||
- SAFETY: Don't proceed to Phase 6 if the script is buggy. Fix the script first.
|
||
- NO COMMIT (sanity check, not implementation).
|
||
- CHECKPOINT: `conductor(checkpoint): Phase 5 complete — v2 staging generated`
|
||
|
||
---
|
||
|
||
## Phase 6: Quality gate (FR7) — HARD GATE
|
||
|
||
Focus: Run the quality gate; if PASS, proceed to Phase 7 (Tier 1 review). If ABORT, fall back to manual review (the user's "B" option).
|
||
|
||
- [ ] **Task 6.1: Run the quality gate**
|
||
- WHERE: `scripts/audit/chronology_quality_gate.py` ← `conductor/chronology.md.staging`
|
||
- WHAT: `uv run python scripts/audit/chronology_quality_gate.py --staging conductor/chronology.md.staging`.
|
||
- HOW: Single command. Capture the exit code and message.
|
||
- SAFETY: This is a HARD GATE. The exit code determines the next phase.
|
||
- VERIFY: Exit code is 0 (PASS) or 1 (ABORT). The message contains either "PASS" or "ABORT".
|
||
|
||
- [ ] **Task 6.2a: IF PASS — proceed to Phase 7**
|
||
- WHERE: continue to Phase 7
|
||
- WHAT: If the gate passed, the classifier is good. Proceed to Phase 7 (Tier 1 review of the "Needs Review" queue).
|
||
- HOW: Note the PASS message + low/total ratio in the per-row evidence log. Continue.
|
||
- NO COMMIT (gate logic, not implementation).
|
||
|
||
- [ ] **Task 6.2b: IF ABORT — fall back to manual review (the user's "B" option)**
|
||
- WHERE: same — Tier 1 reviews every row in the staging file
|
||
- WHAT: If the gate aborted (>30% low confidence), the classifier is bad. Fall back to the v1 manual protocol: Tier 1 opens every row in the staging file, verifies the 5 fields (date, ID, status, summary, range), and assigns a status manually.
|
||
- HOW: Skip Phase 7 (no "Needs Review" queue; Tier 1 reviews all 216 rows). Go directly to Phase 8 (promote staging to canonical). Document the abort + the fallback in the migration report v2 addendum (Phase 9).
|
||
- SAFETY: STOP and report to the user. The fallback is a significant scope change; the user should be aware.
|
||
- NO COMMIT (gate logic).
|
||
- CHECKPOINT: `conductor(checkpoint): Phase 6 complete — quality gate (PASS or ABORT)`
|
||
|
||
---
|
||
|
||
## Phase 7: Tier 1 reviews "Needs Review" queue (FR6 Stage 2) — PAUSE FOR TIER 1
|
||
|
||
Focus: Tier 1 (not the Tier 2 worker) reviews the `low`-confidence rows. **Tier 2 PAUSES here and reports to Tier 1.**
|
||
|
||
- [ ] **Task 7.1: Tier 2 generates the "Needs Review" worklist for Tier 1**
|
||
- WHERE: `conductor/chronology.md.staging` → `tests/artifacts/chronology_v2_needs_review_worklist.md` (new file, gitignored)
|
||
- WHAT: Tier 2 extracts the `low`-confidence rows from the staging file and writes a worklist for Tier 1. Each worklist entry has: track_id, current status, current confidence, evidence (git SHAs, commit count, first/last commit subject, state_phase), classifier reason, the source folder path, the source spec.md path, and a placeholder for Tier 1's resolution.
|
||
- HOW: Use `manual-slop_run_powershell` to extract the "Needs Review" section from the staging file. Format the worklist as a markdown table.
|
||
- SAFETY: The worklist is gitignored. Tier 2 does NOT make resolution decisions; that's Tier 1's job.
|
||
- VERIFY: The worklist has one row per `low`-confidence track.
|
||
- NO COMMIT (gitignored worklist).
|
||
|
||
- [ ] **Task 7.2: Tier 2 generates the per-row evidence log**
|
||
- WHERE: `tests/artifacts/chronology_v2_evidence_log.md` (new file, gitignored)
|
||
- WHAT: A complete log of all 216 rows with: track_id, status, confidence, init_sha, end_sha, commit_count, first_commit_subject, last_commit_subject, state_phase, classifier_reason, tier1_override (placeholder).
|
||
- HOW: Use `manual-slop_run_powershell` to extract the staging file's evidence sub-lines. Format as a markdown table.
|
||
- SAFETY: The log is gitignored. Tier 2 does NOT fill in tier1_override; that's Tier 1's job.
|
||
- VERIFY: The log has 216 rows.
|
||
- NO COMMIT (gitignored log).
|
||
|
||
- [ ] **Task 7.3: Tier 2 PAUSES and reports to Tier 1**
|
||
- WHERE: report to the user
|
||
- WHAT: Tier 2 reports the quality gate result + the worklist + the evidence log to the user. **The user decides who reviews the worklist** (Tier 1 or the user directly).
|
||
- HOW: Print a summary:
|
||
```
|
||
Phase 6 quality gate: PASS (low/total ratio = X%)
|
||
Phase 7 "Needs Review" queue: N rows need Tier 1 review.
|
||
Worklist: tests/artifacts/chronology_v2_needs_review_worklist.md
|
||
Evidence log: tests/artifacts/chronology_v2_evidence_log.md
|
||
|
||
[PAUSE] Awaiting Tier 1 review of the "Needs Review" queue.
|
||
```
|
||
- SAFETY: STOP. Do not proceed to Phase 8 until Tier 1 has reviewed the worklist.
|
||
- NO COMMIT (gate).
|
||
|
||
- [ ] **Task 7.4: Tier 1 reviews each `low`-confidence row (OUTSIDE this plan)**
|
||
- WHERE: `tests/artifacts/chronology_v2_needs_review_worklist.md` (filled in by Tier 1)
|
||
- WHAT: Tier 1 (a different agent or a human) opens the worklist, reads each row's evidence, opens the source folder's `spec.md` and `git log`, and assigns a status. Tier 1's resolution is written in the worklist's "Tier 1 Resolution" column.
|
||
- HOW: This task is NOT executed by Tier 2. Tier 1 fills the worklist manually.
|
||
- SAFETY: Tier 1's defaults per FR6: in `tracks/` + ambiguous → `Active`; in `archive/` + 0 commits → `Special`; in `archive/` + ≥ 3 work commits + state_phase=0 → `Completed`; truly ambiguous → `Special` with note.
|
||
- NO COMMIT (Tier 1's work is gitignored).
|
||
|
||
- [ ] **Task 7.5: Tier 2 applies Tier 1's resolutions to the staging file**
|
||
- WHERE: `conductor/chronology.md.staging` (updated by Tier 2)
|
||
- WHAT: Tier 2 reads the filled worklist and updates the staging file: for each `low`-confidence row, replace the status with Tier 1's resolved status. The "Needs Review" section is replaced with a "Tier 1 Resolutions" section showing each row's resolution note.
|
||
- HOW: Use `manual-slop_run_powershell` to parse the worklist and update the staging file.
|
||
- SAFETY: Only `low`-confidence rows are modified; `high`-confidence rows are preserved.
|
||
- VERIFY: `grep -c "Needs Review" conductor/chronology.md.staging` returns 0 (the section is renamed). `grep -c "Tier 1 Resolutions" conductor/chronology.md.staging` returns 1.
|
||
- NO COMMIT (staging file is gitignored).
|
||
- CHECKPOINT: `conductor(checkpoint): Phase 7 complete — Tier 1 review applied`
|
||
|
||
---
|
||
|
||
## Phase 8: Promote v2 staging → canonical (FR1)
|
||
|
||
Focus: Rename `chronology.md.staging` → `chronology.md`; this is the first time the v2 canonical file is committed.
|
||
|
||
- [ ] **Task 8.1: Move staging to canonical, commit**
|
||
- WHERE: `conductor/chronology.md.staging` → `conductor/chronology.md`
|
||
- WHAT: `git mv conductor/chronology.md.staging conductor/chronology.md`. Then commit.
|
||
- HOW: `git mv` preserves git history (the v1 chronology is in `.broken-v1`; the v2 chronology is a new file). Verify the rename with `git status` (should show a rename of `.staging` to canonical + the removal of `.gitignore` line if applicable).
|
||
- SAFETY: Verify the file content matches the staging file before commit. Don't add the staging path to `.gitignore` if it would block the rename.
|
||
- VERIFY: `git log -1 --stat` shows the rename + the new file. `conductor/chronology.md` has 216 rows.
|
||
- COMMIT: `conductor(chronology): add v2 chronology.md with git-history classifier (FR1)`
|
||
- GIT NOTE: "Phase 8. v2 chronology.md promoted from staging to canonical. Uses git-history classifier (FR5) + per-row evidence + confidence level + 'Needs Review' section. v1 is preserved as .broken-v1 (VC14). Cross-check passed the 30% quality gate (VC10) and Tier 1 reviewed the 'Needs Review' queue (VC11)."
|
||
- CHECKPOINT: `conductor(checkpoint): Phase 8 complete — v2 chronology.md committed`
|
||
|
||
---
|
||
|
||
## Phase 9: Write v2 addendum to migration report + end-of-track report (FR4, VC9)
|
||
|
||
Focus: The v1 migration report gets a v2 addendum documenting the rewrite. The end-of-track report gets a v2 addendum.
|
||
|
||
- [ ] **Task 9.1: Write v2 addendum to migration report**
|
||
- WHERE: `docs/reports/CHRONOLOGY_MIGRATION_20260619.md` (append)
|
||
- WHAT: A new section appended to the end of the v1 report: "v2 Rewrite Addendum (2026-06-21)". Contains:
|
||
- **Why the rewrite was needed** — link to `CHRONOLOGY_TRACK_HANDOVER_20260620.md` + 1-paragraph summary
|
||
- **v1 → v2 status diff** — table of all 216 rows showing the v1 status (stale) and v2 status (after the new classifier) + the git evidence per row. The diff is computed from `tests/artifacts/chronology_v2_evidence_log.md`.
|
||
- **Classifier confidence distribution** — counts: `high` / `low` / total; % of total in `Needs Review`
|
||
- **Tier 1 review log** — for each `low`-confidence row, the resolution note
|
||
- **Quality gate result** — PASS or ABORT + the ratio
|
||
- **User sign-off** — a checklist section for the user to fill in during Phase 10
|
||
- HOW: Use `manual-slop_run_powershell` to extract the evidence log + format the diff. Hand-write the Tier 1 review log + the sign-off checklist.
|
||
- SAFETY: The addendum does not modify the v1 sections; it only appends.
|
||
- VERIFY: The file is now > 200 lines (v1 was 174; addendum is ~50-100 lines).
|
||
- COMMIT: `docs(chronology): write v2 rewrite addendum to migration report (FR4)`
|
||
- GIT NOTE: "Phase 9.1. v2 addendum appended to CHRONOLOGY_MIGRATION_20260619.md. Documents the rewrite rationale + v1→v2 status diff + classifier confidence + Tier 1 review log + quality gate result + user sign-off checklist."
|
||
|
||
- [ ] **Task 9.2: Write v2 addendum to end-of-track report**
|
||
- WHERE: `docs/reports/TRACK_COMPLETION_chronology_20260619.md` (append)
|
||
- WHAT: A new section: "v2 Rewrite Addendum (2026-06-21)". Contains:
|
||
- Summary (1-2 sentences)
|
||
- v2 final state (5 fields: chronology.md, tracks.md, workflow.md, migration report, end-of-track report)
|
||
- v2 statistics (rows in v2, classifier confidence, Tier 1 review log size, quality gate result)
|
||
- Cross-check summary (per VC10/11/12/13 confirmation)
|
||
- User sign-off (reference to the migration report v2 addendum)
|
||
- HOW: Hand-write the addendum; reference the v1 sections.
|
||
- VERIFY: The file is updated.
|
||
- COMMIT: `docs(chronology): write v2 addendum to end-of-track report (VC9)`
|
||
- GIT NOTE: "Phase 9.2. v2 addendum to TRACK_COMPLETION. Documents the v2 final state + cross-check summary."
|
||
- CHECKPOINT: `conductor(checkpoint): Phase 9 complete — migration report + end-of-track report updated`
|
||
|
||
---
|
||
|
||
## Phase 10: User sign-off (FR6 Stage 3) — PAUSE FOR USER
|
||
|
||
Focus: User is the final quality gate. The track is not "done" until the user signs off.
|
||
|
||
- [ ] **Task 10.1: Tier 2 PAUSES and reports to the user**
|
||
- WHERE: report to the user
|
||
- WHAT: Tier 2 reports the v2 state to the user:
|
||
```
|
||
Phase 9 complete. v2 chronology.md committed.
|
||
Quality gate: PASS (low/total ratio = X%)
|
||
Tier 1 review: N rows resolved.
|
||
|
||
Files:
|
||
- conductor/chronology.md (v2; 216 rows)
|
||
- conductor/chronology.md.broken-v1 (v1 archive; 218 lines)
|
||
- docs/reports/CHRONOLOGY_MIGRATION_20260619.md (v1 + v2 addendum)
|
||
- docs/reports/TRACK_COMPLETION_chronology_20260619.md (v1 + v2 addendum)
|
||
|
||
[PAUSE] Awaiting user sign-off.
|
||
```
|
||
- SAFETY: STOP. Do not proceed to Phase 11 until the user signs off.
|
||
- NO COMMIT (gate).
|
||
|
||
- [ ] **Task 10.2: User reviews v2 + evidence log + Tier 1 resolutions (USER TASK, outside this plan)**
|
||
- WHERE: `conductor/chronology.md`, `tests/artifacts/chronology_v2_evidence_log.md`, `docs/reports/CHRONOLOGY_MIGRATION_20260619.md` (v2 addendum)
|
||
- WHAT: User confirms: (a) format correct, (b) every row has evidence, (c) Tier 1 resolutions are reasonable, (d) nothing missed.
|
||
- HOW: User fills the sign-off checklist in the v2 addendum of the migration report.
|
||
- SAFETY: This is a user task; Tier 2 does not execute it.
|
||
- NO COMMIT (gate).
|
||
|
||
- [ ] **Task 10.3: Tier 2 records the user sign-off**
|
||
- WHERE: `docs/reports/CHRONOLOGY_MIGRATION_20260619.md` (v2 addendum)
|
||
- WHAT: Tier 2 records the user's sign-off in the v2 addendum (date + confirmation). If the user requested changes, loop back to the relevant phase.
|
||
- HOW: Use `manual-slop_edit_file` to add the sign-off line.
|
||
- VERIFY: The v2 addendum has a "User Sign-off" section with a date.
|
||
- COMMIT: `docs(chronology): record user sign-off in migration report v2 addendum (VC13)`
|
||
- GIT NOTE: "Phase 10.3. User sign-off recorded. All VCs satisfied."
|
||
- CHECKPOINT: `conductor(checkpoint): Phase 10 complete — user sign-off recorded`
|
||
|
||
---
|
||
|
||
## Phase 11: Wrap-up
|
||
|
||
Focus: Mark the track complete in `tracks.md`, `state.toml`, and `metadata.json`.
|
||
|
||
- [ ] **Task 11.1: Mark track complete in `conductor/tracks.md`**
|
||
- WHERE: `conductor/tracks.md` (the entry for `chronology_20260619` in the "Active Tracks" table or the equivalent)
|
||
- WHAT: Flip the status marker from `[ ]` or `[~]` to `[x]`. Add a one-line note: "Completed 2026-06-21 (v2 rewrite after tier-2 failure report)."
|
||
- HOW: Use `manual-slop_edit_file` to find the entry and update it.
|
||
- SAFETY: Don't touch other entries.
|
||
- VERIFY: `grep -n "chronology_20260619" conductor/tracks.md` shows the entry with `[x]`.
|
||
- COMMIT: `conductor(track): mark chronology_20260619 as complete in tracks.md`
|
||
- GIT NOTE: "Phase 11.1. Track marked complete in tracks.md."
|
||
|
||
- [ ] **Task 11.2: Update `state.toml` to completed**
|
||
- WHERE: `conductor/tracks/chronology_20260619/state.toml`
|
||
- WHAT: Set `[meta].status = "completed"`, `[meta].current_phase = "complete"`, `[meta].last_updated = "<today's date>"`. Mark all 11 phases as `completed` with their checkpoint SHAs. Mark all tasks as `completed` with their commit SHAs. Mark all verification flags as `true`.
|
||
- HOW: Use `manual-slop_edit_file` to set the fields. The task commit SHAs come from `git log --oneline conductor/tracks/chronology_20260619/`.
|
||
- SAFETY: Don't change the task descriptions; just flip the status fields.
|
||
- VERIFY: `uv run python -c "import tomllib; data = tomllib.load(open('conductor/tracks/chronology_20260619/state.toml','rb')); assert data['meta']['status'] == 'completed'; print('OK')"` exits 0.
|
||
- COMMIT: `conductor(plan): mark chronology_20260619 v2 as completed in state.toml`
|
||
- GIT NOTE: "Phase 11.2. Track complete. All 11 phases satisfied."
|
||
|
||
- [ ] **Task 11.3: Update `metadata.json` to completed**
|
||
- WHERE: `conductor/tracks/chronology_20260619/metadata.json`
|
||
- WHAT: Set `"status": "completed"`. Update the `verification_criteria` array with the v2 VCs (replace v1's 12 VCs with v2's 14 VCs). Update the `risk_register` to include the v2 risks (replace v1's 9 risks with v2's 10 risks). Add a new `regressions_and_pre_existing_failures` entry: "v1 chronology had 167 rows with stale status; v2 fixed via git-history classifier (FR5)."
|
||
- HOW: Use `manual-slop_edit_file` to update the JSON.
|
||
- SAFETY: Keep the JSON valid (run `python -c "import json; json.load(open('metadata.json'))"` after the edit).
|
||
- VERIFY: `uv run python -c "import json; data = json.load(open('conductor/tracks/chronology_20260619/metadata.json')); assert data['status'] == 'completed'; assert len(data['verification_criteria']) >= 14; print('OK')"` exits 0.
|
||
- COMMIT: `conductor(plan): mark chronology_20260619 v2 as completed in metadata.json`
|
||
- GIT NOTE: "Phase 11.3. metadata.json updated with v2 VCs and risk register."
|
||
|
||
- [ ] **Task 11.4: Final end-of-track commit**
|
||
- WHERE: working tree
|
||
- WHAT: Verify `git status` is clean except for the v2 file changes. Print a final summary of the 11 phases + their checkpoints.
|
||
- HOW: `git log --oneline conductor/tracks/chronology_20260619/ | head -30` (shows the v1 + v2 commit history).
|
||
- SAFETY: Don't commit unrelated working-tree changes.
|
||
- NO COMMIT (verification).
|
||
- CHECKPOINT: `conductor(checkpoint): Phase 11 complete — track fully complete`
|
||
|
||
---
|
||
|
||
## Summary
|
||
|
||
| Plan Phase | Scope | Pause point? | Spec Phase |
|
||
|---|---|---|---|
|
||
| 1 | Archive v1 + verify state + reset state.toml | No | Spec Phase 1 |
|
||
| 2 | Rewrite `_classify_status` (TDD red+green) | No | Spec Phase 2 (part 1) |
|
||
| 3 | Update `extract_summary` + `format_markdown` (TDD red+green) | No | Spec Phase 2 (part 2) |
|
||
| 4 | Add quality gate script (TDD red+green) | No | Spec Phase 3 |
|
||
| 5 | Run classifier, generate staging | No | Spec Phase 4 |
|
||
| 6 | Quality gate (PASS or ABORT) | Yes (HARD) | Spec Phase 5 |
|
||
| 7 | Tier 1 reviews "Needs Review" queue | Yes (PAUSE for Tier 1) |
|
||
| 8 | Promote staging → canonical | No |
|
||
| 9 | Write v2 addendum to migration + end-of-track reports | No |
|
||
| 10 | User sign-off | Yes (PAUSE for user) |
|
||
| 11 | Wrap-up (mark complete) | No |
|
||
|
||
**Total: 11 plan phases (10 spec phases; Plan Phase 2 = Spec Phase 2 part 1, Plan Phase 3 = Spec Phase 2 part 2; renumbering shifts from Plan Phase 4 onwards), 4 hard pause points (Plan Phase 6 quality gate, Plan Phase 7 Tier 1 review, Plan Phase 10 user sign-off, plus the Plan Phase 6 ABORT fallback), ~25 commits.** The cross-check is the dominant cost; the user's "EVERY SINGLE ENTRY" mandate is preserved at the semantic level via the per-row evidence log + the 3-stage protocol (FR6).
|
||
|
||
## Verification Criteria Recap
|
||
|
||
All 14 VCs from the spec must be satisfied for the track to be marked complete:
|
||
- VC1: `chronology.md` v2 with 216 rows, 5-value status enum, per-row evidence + confidence
|
||
- VC2: `tracks.md` pruning intact (no regression)
|
||
- VC3: `workflow.md` 3-step convention present (no regression)
|
||
- VC4: migration report v2 addendum
|
||
- VC5: sorted newest first; Folder + Range + Evidence per row
|
||
- VC6: every folder has a row or documented exception
|
||
- VC7: Notable Non-Track Commits section preserved
|
||
- VC8: no new `src/*.py` files
|
||
- VC9: end-of-track v2 addendum
|
||
- VC10: quality gate PASS (or ABORT with manual fallback)
|
||
- VC11: "Needs Review" queue resolved by Tier 1
|
||
- VC12: per-row evidence log (`tests/artifacts/chronology_v2_evidence_log.md`)
|
||
- VC13: user sign-off recorded
|
||
- VC14: v1 archive preserved as `.broken-v1`
|
||
|
||
## Cross-cutting safety
|
||
|
||
- **No day estimates** (per `conductor/workflow.md` Tier 1 rule).
|
||
- **Per-task atomic commits** (per `conductor/workflow.md` Commit Guidelines).
|
||
- **Git notes on every commit** (per project convention).
|
||
- **No `git restore` / `git checkout -- <file>` / `git reset`** (per HARD BAN).
|
||
- **No new `src/*.py` files** (per `AGENTS.md` File Size and Naming Convention).
|
||
- **No new `conductor/code_styleguides/*` files** (the 3-step convention is in `workflow.md`, not a new styleguide).
|
||
- **TDD for all script changes** (Phases 2-4 are red+green cycles).
|
||
- **Tier 2 uses `uv run python scripts/run_tests_batched.py`** (never direct `uv run pytest`; per `conductor/workflow.md` "Tier 2 Autonomous Sandbox" rule #1).
|
||
- **Throw-away scripts go in `scripts/tier2/artifacts/chronology_20260619/`** (per `conductor/workflow.md` "Tier 2 Autonomous Sandbox" rule #4).
|
||
- **End-of-track report at `docs/reports/TRACK_COMPLETION_chronology_20260619.md`** (already exists; v2 addendum appended in Phase 9.2).
|
||
- **Default branch: `master`** (per `conductor/workflow.md` "Tier 2 Autonomous Sandbox" rule #2).
|
||
- **Preserve line endings on edit** (per `conductor/workflow.md` "Tier 2 Autonomous Sandbox" rule #3; this repo has mixed CRLF/LF).
|