Files

32 KiB

Track Specification: Conductor Chronology v2 (2026-06-21 rewrite)

Overview

This is the v2 rewrite of chronology_20260619. The first run (Phases 1-9, 24 commits, 2026-06-19 to 2026-06-20) shipped conductor/chronology.md with a broken status classifier that read stale metadata.json.status fields. The user mandate — "EVERY SINGLE ENTRY MUST BE CROSS CHECKED" — was satisfied at a structural level (folder set == row set) but the semantic level (status correctness, summary quality) was not. Two classifier iterations followed (commits 4109a667 and 271e6895); both used heuristic-based fallbacks and neither used git history as the explicit evidence source the user wants.

This rewrite replaces the spec/plan/state.toml; the 24 prior commits + the broken v1 chronology remain in git history as the foundation. The substantive changes are:

  1. FR1 (chronology structure): rewritten — new status enum (5 values), per-row evidence line, per-row confidence level, "Needs Review" section.
  2. FR5 (helper script): rewritten — git-history classifier with confidence assignment.
  3. FR6 (cross-check): rewritten — 3-stage protocol (classifier auto + Tier 1 reviews "Needs Review" queue + user reviews final).
  4. FR7 (new): classifier quality gate — if > 30% of rows are ambiguous, abort to manual review (the user's "B" fallback).

Phases that produced the existing tracks.md pruning + workflow.md 3-step convention + the v1 migration report are reused. This rewrite adds a v2 addendum to the migration report.

Current State Audit (as of 2026-06-21, commit 3aea92f1)

Already Implemented (carried forward, NO REWORK)

  1. conductor/tracks.md "Phase 9: Chore Tracks" section — pruned to one-line stub pointing to chronology.md (commit be38dd5).
  2. conductor/tracks.md "Active Research Tracks" [x] entries — pruned (commit cca4767).
  3. conductor/tracks.md "Follow-up" [shipped] entries — pruned (commit b3a9c45).
  4. conductor/workflow.md "Notes > Editing this file" section — has the 3-step archiving convention (commit b697cd8).
  5. scripts/audit/generate_chronology.py — exists (338 lines). Functions: extract_slug_date, extract_summary, walk_track_folders, format_markdown, _classify_status, _parse_state_phase, _last_commit_date. The broken function is _classify_status (lines ~163-189) which reads the current parameter (originally from metadata.json.status) and uses folder-location + state_phase heuristics. This function is the target of FR5's rewrite.
  6. tests/test_generate_chronology.py — 6 unit tests, all passing against the current (broken) classifier. Need extension per FR5.
  7. conductor/chronology.md — 218 lines, 216 rows, v1 with broken status classifier. Statuses include active, spec_written, spec_approved, planning (stale metadata.json.status values). 41 Completed, 0 Abandoned, 167 rows with stale status per the handover report (line 14-16). Target of Phase 1's move-to-broken-v1.
  8. docs/reports/CHRONOLOGY_MIGRATION_20260619.md — v1 migration report; needs v2 addendum (FR4).
  9. docs/reports/CHRONOLOGY_TRACK_HANDOVER_20260620.md — tier-2's hand-off; documents the failure + the recommended fix (the 5-step git-history algorithm).
  10. docs/reports/TRACK_COMPLETION_chronology_20260619.md — v1 end-of-track report; needs v2 addendum.

Gaps to Fill (This Track's Scope)

# Gap Where Resolution
G1 v1 chronology.md has 167/216 rows with wrong status (stale metadata.json.status values) conductor/chronology.md Move v1 to conductor/chronology.md.broken-v1 (Phase 1); generate v2 with git-history classifier (Phase 4)
G2 v1 chronology.md has summaries that are metadata-field text (**Priority:** A..., **Date:** 2026-06-20) not the actual track summary Same as G1 v2's priority chain (FR5 §"Summary extraction") rejects metadata-field text via regex
G3 _classify_status reads stale metadata.json.status scripts/audit/generate_chronology.py:~163-189 Rewrite to use the 5-step git-history algorithm (handover §"Root cause of failure")
G4 No "Needs Review" queue mechanism n/a (new) Add per-row confidence (FR5) + "Needs Review" section in chronology.md (FR1)
G5 No quality gate to detect a bad classifier n/a (new) Add scripts/audit/chronology_quality_gate.py (FR7)
G6 v1 cross-check was bulk-verified (structural check, not per-row semantic check) n/a (process change) v2 cross-check is 3-stage (FR6): classifier auto + Tier 1 reviews "Needs Review" + user reviews final with per-row evidence log
G7 v1 per-row evidence is missing n/a (new) Add per-row evidence line to chronology.md (FR1) + standalone evidence log file (FR6 §"per-row evidence log")
G8 state.toml is at current_phase = 10 with a false "complete" state conductor/tracks/chronology_20260619/state.toml Reset to current_phase = 0; this rewrite starts fresh
G9 v1 migration report has 167 stale-status rows in the per-row log docs/reports/CHRONOLOGY_MIGRATION_20260619.md v2 addendum shows the diff (v1 status → v2 status) with the git evidence per row
G10 No fallback path if the classifier is bad n/a (new) FR7 quality gate; if > 30% ambiguous → abort to manual review (the user's "B" fallback per chat 2026-06-21)

Goals

  1. One canonical index. conductor/chronology.md is the only file consulted to see "what has this project done." No more scanning 3 sections of tracks.md. (Carried from v1; unchanged.)
  2. No info loss. Every track that has a folder in conductor/tracks/ or conductor/archive/ has a row in chronology.md (or a documented exception). (Carried from v1; unchanged.)
  3. Forward-compatible. When a new track ships, the convention is clear: move folder to archive/, remove [x] from tracks.md, add a row to chronology.md with the new format. (Carried from v1; unchanged.)
  4. Git history is the explicit evidence. Each row's status is derived from git log -- <folder> (commit count + commit messages). metadata.json.status is informational only — the classifier does not trust it for the final status.
  5. "EVERY SINGLE ENTRY" mandate preserved at the semantic level. Every row has: (a) a status decision, (b) the git evidence that supports the decision, (c) a per-row confidence level, (d) a "Needs Review" flag if confidence is low. The "cross-check" is the row's evidence trail, not a separate audit pass.
  6. Conservative classifier + hard quality gate. The classifier auto-classifies only when evidence is clear; ambiguous rows are flagged for human review. If > 30% of rows are ambiguous, the classifier is bad → abort to manual review (the user's "B" fallback per chat 2026-06-21).
  7. No day estimates. Per conductor/workflow.md Tier 1 Track Initialization Rules (added 2026-06-16). Scope measured in files/sites.

Functional Requirements

FR1. conductor/chronology.md v2 structure (REWRITTEN)

WHERE: conductor/chronology.md (replaces v1).

WHAT: Same overall structure as v1 (table format, newest first, "Notable Non-Track Commits" section at the bottom), with these changes:

Status enum (5 values, replaces v1's 6-value enum):

  • Active — folder in tracks/ + work has started (≥ 1 feat/fix/refactor commit) but state.toml.current_phase < 3
  • In Progress — folder in tracks/ + state.toml.current_phase ≥ 3 (or no state.toml + ≥ 3 work commits)
  • Completed — folder in archive/ + ≥ 3 work commits (or state.toml.current_phase == "complete")
  • Abandoned — folder in tracks/ or archive/ + 0-1 work commits + last commit > 14 days ago + no feat/fix/refactor in commit history
  • Special — explicit human-decision; e.g., research note, scratch dir, archived by mistake, deleted

Notably ABSENT from the v2 enum (present in v1): Shipped, Superseded, planning, spec_written, spec_approved, active (lowercase). The v2 enum is the canonical set; v1's status values are stale metadata leaks.

Per-row confidence level (NEW):

  • high — auto-classified by the script; git evidence + folder location + state.toml (if present) all point to the same status
  • low — in the "Needs Review" queue; needs Tier 1 + user review

Per-row evidence line (NEW): Each row gets a sub-line in the format:

Evidence: <7-char-init-sha>..<7-char-end-sha> | N commits | state_phase=<N or "n/a" or "complete"> | "<first-commit-subject>" → "<last-commit-subject>" | confidence=<high|low>

"Needs Review" section (NEW): At the bottom of chronology.md, a section listing all low-confidence rows with a one-line reason each. Format:

## Needs Review (Tier 1 + User)

These rows had ambiguous git evidence. Resolved by Tier 1; user reviewed in Stage 3.

- `<track_id>` (status=<resolved>) — <one-line reason> — resolved by Tier 1

Other v1 fields preserved unchanged: Date, Track ID, Summary (≤ 25 words), Folder, Range (<init-sha>..<end-sha> with commit count), Notable Non-Track Commits section.

Worked example (new format):

| 2026-06-19 | `chronology_20260619` | In Progress | **Confidence:** low | v2 rewrite of the chronology track after tier-2's failure report identified the broken status classifier. | `conductor/tracks/chronology_20260619` | `87923c93..3aea92f1` (12) |
| | | | | | Evidence: `87923c9..3aea92f` | 12 commits | state_phase=n/a (this rewrite) | "conductor(track): add initial spec for chronology_20260619" → "botched the chronology, going to rewrite the track." | confidence=low |

FR2. conductor/tracks.md pruning (CARRIED FORWARD; no changes)

Already complete in v1 (commits be38dd5, cca4767, b3a9c45). This rewrite verifies the pruning is intact and re-commits nothing.

Verification step: Phase 1 of the v2 plan runs grep -n "^- \[x\]" conductor/tracks.md and confirms 0 matches (other than the Status legend at the bottom of the file).

FR3. conductor/workflow.md 3-step convention (CARRIED FORWARD; no changes)

Already complete in v1 (commit b697cd8). This rewrite verifies the 3-step block is present and re-commits nothing.

Verification step: Phase 1 of the v2 plan runs grep -n "Archiving a track" conductor/workflow.md and confirms 1 match.

FR4. Migration report v2 addendum (UPDATED)

WHERE: docs/reports/CHRONOLOGY_MIGRATION_20260619.md (extends existing report).

WHAT: A new section appended to the end of the v1 report: "v2 Rewrite Addendum (2026-06-21)". Contains:

  • Why the rewrite was needed — link to CHRONOLOGY_TRACK_HANDOVER_20260620.md + summary of the root cause
  • v1 → v2 status diff — table of all 216 rows showing the v1 status (stale) and v2 status (after the new classifier) + the git evidence per row
  • Classifier confidence distribution — counts: high / low / total; % of total in Needs Review
  • Tier 1 review log — for each low-confidence row, the resolution note (assigned status + reason + override if any)
  • Quality gate result — was the 30% threshold hit? If so, the abort-to-B was triggered.
  • Outstanding issues — any rows the user flagged for follow-up

FR5. Helper script rewrite — git-history classifier (REWRITTEN)

WHERE: scripts/audit/generate_chronology.py (rewritten) + tests/test_generate_chronology.py (extended).

WHAT: The script's _classify_status function is rewritten to use the handover's 5-step algorithm. The new signature is:

def _classify_status(
 folder_link: str,
 init_sha: str,
 end_sha: str,
 commit_count: int,
 first_commit_subject: str,
 last_commit_subject: str,
 state_phase: str | None,
 metadata_status: str | None,
 last_commit_date: str,
) -> tuple[str, str, str]:
 """Classify a track's status using git history as primary evidence.

 Returns:
  (status, confidence, reason) where:
  - status: one of "Active", "In Progress", "Completed", "Abandoned", "Special"
  - confidence: "high" or "low"
  - reason: one-line explanation of the classification
 """

The 5-step algorithm (per the handover §"Rewrite _classify_status to use git history as primary evidence"):

  1. Count meaningful commits. commit_count (already computed by the script via git log --oneline -- <folder> | wc -l). 1-2 commits (just spec/plan creation) is a strong signal for Active (in tracks/) or Abandoned (in archive/). ≥ 3 work commits is a strong signal for Completed (in archive/) or In Progress (in tracks/).

  2. Inspect commit messages. first_commit_subject and last_commit_subject (already extracted by the script). Classify each commit as work (matches ^(feat|fix|refactor|perf|test)\() or meta (matches ^(chore|docs|conductor)\() or other (everything else).

  3. Check state.toml phase progression. state_phase is parsed from state.toml.current_phase if the file exists; else None. The thresholds:

    • state_phase == "complete"Completed (high confidence if corroborated by git)
    • state_phase >= 3In Progress (high confidence if corroborated by git)
    • state_phase in (0, 1, 2)Active (high confidence if corroborated by git)
    • state_phase is None → no signal from state.toml; classifier relies on git + folder
  4. Default to conservative. When git history is ambiguous (1-3 commits with no clear work pattern), flag as low confidence → "Needs Review". The classifier NEVER auto-marks Abandoned — that's a Special decision reserved for Tier 1 + user.

  5. Honour explicit metadata. If metadata_status is abandoned or superseded (or Special), and git evidence is not contradictory, trust the metadata. If git evidence contradicts metadata (e.g., archive/ + 0 commits + metadata_status = "Completed"), the classifier flags low confidence and the user resolves in Stage 3.

Per-row confidence assignment:

  • high — git evidence + folder location + state.toml (if present) all point to the same status. Default for unambiguous cases.
  • low — any of: (a) < 3 commits total, (b) conflicting signals (e.g., archive/ + 0 commits + state_phase 0), (c) no state.toml + ambiguous git history, (d) metadata_status contradicts git.

Summary extraction (REWRITTEN priority chain): The v1 priority chain is replaced with a regex-aware version:

  1. metadata.json.summary if present and does not start with ** (regex: ^\*\*)
  2. First non-empty line of spec.md that does not start with **
  3. metadata.json.description if not starting with **
  4. First non-empty line of plan.md that does not start with **
  5. Generic placeholder: "Imported from archive (no spec)" for archive rows, "Track folder (no spec found)" for tracks/ rows

The regex ^\*\* rejects metadata-field text like **Priority:** A..., **Date:** 2026-06-20, **Created:** 2026-06-19, **Initialized:** 2026-06-19, **Parent umbrella:** ..., **Confidence:** ....

New script: scripts/audit/chronology_quality_gate.py (FR7's wrapper).

  • Reads the staging chronology.md.staging file.
  • Counts high and low confidence rows.
  • Computes low_count / total_count.
  • If ratio > 0.30 → exit code 1, prints "ABORT: classifier is bad; >30% of rows are ambiguous. Fall back to manual review (v1 protocol)."
  • If ratio ≤ 0.30 → exit code 0, prints "PASS: classifier is good. Proceed to Tier 1 review of 'Needs Review' queue."

Tests extended: the existing 6 tests stay; add 8-10 new tests covering:

  • _classify_status returns correct status for each (folder, commit_count, state_phase) combination
  • low confidence is assigned for ambiguous cases (1-2 commits, conflicting signals)
  • high confidence is assigned for unambiguous cases
  • Summary priority chain rejects metadata-field text (regression test for the v1 bug)
  • The staging file has per-row evidence + confidence lines
  • The "Needs Review" section is correctly populated
  • The quality gate script exits 1 when > 30% ambiguous, 0 when ≤ 30%
  • The quality gate script prints the correct summary

FR6. Per-row cross-check (REWRITTEN — 3-stage protocol)

WHERE: conductor/chronology.md v2 (after classifier run), then "Needs Review" queue (Tier 1 review), then final v2 (user review).

WHAT: The cross-check is 3-stage (replaces v1's single-stage Tier 1 review of every row):

Stage 1: Classifier auto-classification (script run).

  • The script runs walk_track_folders() over conductor/tracks/ and conductor/archive/.
  • For each folder, the script extracts: date, track_id, init_sha, end_sha, commit_count, first_commit_subject, last_commit_subject, state_phase, metadata_status, last_commit_date, summary.
  • The script's rewritten _classify_status() assigns (status, confidence, reason) for each row.
  • Output: conductor/chronology.md.staging with the per-row evidence line + confidence level + "Needs Review" section.
  • The script is READ-ONLY on the source folders; it writes to chronology.md.staging only.
  • Quality gate (FR7) runs immediately after: if the gate passes, proceed to Stage 2; if the gate fails, the staging file is preserved and the task aborts to manual review (per FR7).

Stage 2: Tier 1 review of the "Needs Review" queue (only if quality gate passes).

  • Tier 1 opens conductor/chronology.md.staging.
  • Tier 1 filters to the "Needs Review" section (rows with confidence=low).
  • For each low-confidence row, Tier 1:
    1. Opens the track's spec.md (or plan.md / metadata.json if no spec).
    2. Runs git log --oneline -- <folder> and reviews the commit history.
    3. Verifies the row's evidence line is accurate.
    4. Assigns a status from the 5-value enum (or flags for user decision).
    5. Writes a one-line resolution note (e.g., "Resolved: Active — work in progress, state_phase=2; classifier flagged low because no spec.md yet").
  • Tier 1's defaults:
    • In tracks/ + ambiguous → Active with a one-line note
    • In archive/ + 0 commits → Special with note "archive folder with no work commits"
    • In archive/ + ≥ 3 work commits + state_phase=0 (missing/incomplete) → Completed with note "archive + N work commits; state.toml is stale"
    • Truly ambiguous → Special with note "needs user decision; flagged in Stage 3"
  • After Tier 1 resolves all low-confidence rows, the staging file is updated: the "Needs Review" section is moved to a "Tier 1 Resolutions" section showing each row's resolution note.

Stage 3: User review of final v2.

  • User opens conductor/chronology.md.staging (now with Stage 2 resolutions).
  • User reviews: (a) the format is correct, (b) every row has evidence + decision, (c) Tier 1's resolutions are reasonable, (d) nothing missed.
  • User either approves (proceed to Phase 7 promotion) or requests changes (loop back to Stage 2 or 1).

The per-row evidence log (NEW FILE).

  • Path: tests/artifacts/chronology_v2_evidence_log.md (gitignored).
  • Format: one row per track with: track_id, status, confidence, init_sha, end_sha, commit_count, first_commit_subject, last_commit_subject, state_phase, classifier_reason, tier1_override (if any).
  • Generated by the script during Stage 1; extended by Tier 1 during Stage 2; reviewed by the user in Stage 3.

FR7. Classifier quality gate (NEW)

WHERE: scripts/audit/chronology_quality_gate.py (new file) + tests/test_chronology_quality_gate.py (new tests).

WHAT: A wrapper script that runs after the classifier's Stage 1 output. The script:

  1. Reads conductor/chronology.md.staging (the script's output).
  2. Parses each row's confidence level.
  3. Counts high and low confidence rows.
  4. Computes low_count / total_count.
  5. If ratio > 0.30 → exit code 1, prints "ABORT: classifier is bad; >30% of rows are ambiguous. Fall back to manual review (v1 protocol). Tier 1 should manually review every row in the staging file."
  6. If ratio ≤ 0.30 → exit code 0, prints "PASS: classifier is good. rows need Tier 1 review; proceed to Stage 2."

The 30% threshold is a hard gate. Tier 1 doesn't start Stage 2 until the gate passes. If the gate fails, the staging file is preserved as chronology.md.staging.aborted and the task falls back to the v1 manual protocol (Tier 1 reviews every row).

Tests for the quality gate:

  • Staging file with 0% low → exit 0
  • Staging file with 30% low (boundary) → exit 0
  • Staging file with 31% low → exit 1
  • Staging file with 100% low → exit 1
  • Staging file with malformed rows → exit 2 (parse error)

Non-Functional Requirements

(Carried from v1, mostly unchanged.)

  • NFR1. Manually maintained. Per user choice (2026-06-19), the ongoing workflow is hand-edited. No auto-generation in CI; no script runs on every commit. The one-shot migration is a single event; the file is then edited like tracks.md.
  • NFR2. Compact. Each row is ≤ 5 lines (the bullet + 3 sub-lines for Folder/Range/Evidence, OR a single condensed line for very old tracks where the folder is the only link). The file is scannable, not a wall of text.
  • NFR3. Re-derivable. A reader can rebuild the chronology from git log + the track folders if needed. The init SHA + end SHA + evidence line in each row is the contract; the summary is the human-friendly gloss.
  • NFR4. No day estimates. Per conductor/workflow.md Tier 1 Track Initialization Rules (added 2026-06-16). All scope is measured in files/sites.
  • NFR5. No TDD required for the chronology itself. This is a documentation/tooling track, not a feature track. The helper script (FR5) gets 8-10 new unit tests for the new classifier (TDD-required per project convention).
  • NFR6. Evidence is auditable (NEW). The per-row evidence log (tests/artifacts/chronology_v2_evidence_log.md) is human-readable; every classification decision is reproducible from the log + git history. A reader can verify any row's status by running git log -- <folder> and comparing to the evidence log.
  • NFR7. Classifier is conservative (NEW). When in doubt, low confidence. The cost of a false low (Tier 1 reviews it) is small; the cost of a false high (wrong status committed without review) is high. The classifier's bias is toward low.

Architecture Reference

  • docs/reports/CHRONOLOGY_TRACK_HANDOVER_20260620.md — the failure report; the source of the new classifier algorithm (5-step algorithm, §"Rewrite _classify_status to use git history as primary evidence", lines 53-68).
  • docs/reports/CHRONOLOGY_MIGRATION_20260619.md — v1 migration report; the v2 addendum (FR4) extends it.
  • conductor/code_styleguides/data_oriented_design.md — applies: the chronology is data (one row per track), the classifier is a transformation (git history → status), the evidence log is a projection (data + decision + provenance).
  • conductor/code_styleguides/error_handling.md — applies to the helper script: the script's _classify_status returns (status, confidence, reason) (a data-oriented "and/or" pattern, not an exception). The "Needs Review" queue is a recoverable case (low confidence), not an error.
  • conductor/tracks.md:459 — the existing "lightweight chronology" reference. v2 formalizes that role.
  • conductor/workflow.md "Notes > Editing this file" — the existing convention for moving tracks to archive/. The 3-step convention (FR3) is appended here.

Out of Scope

(Carried from v1, mostly unchanged.)

  1. Auto-generation on every commit. Per the user's "manual maintenance" choice (2026-06-19), there's no script that updates chronology.md automatically. The file is hand-edited when a track is archived.
  2. Tracking "in-flight" tracks in chronology.md. In-flight tracks ([~] in tracks.md) appear in chronology.md with status Active or In Progress (per v2's enum). The active task list still lives in tracks.md.
  3. Tracking "planned but not specced" backlog items. These stay in tracks.md under "Follow-up" and "Backlog". They aren't tracks until they have a folder.
  4. Restructuring tracks.md beyond [x] removal. The 3 sections that held [x] entries are now stubs (v1 Phase 3); no new structure is imposed.
  5. A separate chronology/ folder for the file. The file lives at the conductor root (conductor/chronology.md), not in a subdirectory.
  6. Reformatting existing spec.md / plan.md files. The migration reads from them; it does not modify them.
  7. A web view of the chronology. It's a markdown file for in-repo reading. No GUI integration is in scope.
  8. A separate chronology.md.draft workflow (NEW for v2). v1 used .draft files; v2 doesn't. The classifier emits directly to a staging file (chronology.md.staging); the staging file is renamed to chronology.md after Stage 2 (Tier 1 review). The .staging suffix is gitignored.

Verification Criteria

For the track to be marked complete, ALL of the following must be true:

  • VC1. conductor/chronology.md v2 exists with 216 rows; all 5 status values are used; per-row evidence line is present; per-row confidence level is present.
  • VC2. conductor/tracks.md pruning is intact (no regression from v1's pruning; grep -n "^- \[x\]" conductor/tracks.md returns 0 matches).
  • VC3. conductor/workflow.md 3-step convention is present (no regression; grep -n "Archiving a track" conductor/workflow.md returns 1 match).
  • VC4. docs/reports/CHRONOLOGY_MIGRATION_20260619.md has the v2 addendum (per FR4).
  • VC5. Sorted newest first; every row has Folder + Range + Evidence lines.
  • VC6. Every folder in conductor/tracks/ and conductor/archive/ has a corresponding row, OR a documented exception in the v2 addendum.
  • VC7. "Notable Non-Track Commits" section is preserved (may be empty if no notable commits found).
  • VC8. No new src/*.py files created (per AGENTS.md File Size and Naming Convention rule).
  • VC9. v2 addendum to docs/reports/TRACK_COMPLETION_chronology_20260619.md (per project convention).
  • VC10. Classifier quality gate (FR7). The scripts/audit/chronology_quality_gate.py ran; result was PASS (low confidence ≤ 30%). If the gate failed, the abort-to-B was triggered and Tier 1 manually reviewed every row.
  • VC11. "Needs Review" queue resolved (FR6 Stage 2). Every low-confidence row in the staging file has a Tier 1 resolution note; the queue is empty in the final chronology.md (Tier 1's resolutions are reflected in the per-row status).
  • VC12. Per-row evidence log (FR6). tests/artifacts/chronology_v2_evidence_log.md has one row per track with status + confidence + evidence + decision (Tier 1 override if any).
  • VC13. User sign-off (FR6 Stage 3). User confirmed: format correct, every row has evidence, Tier 1 resolutions are reasonable, nothing missed. Sign-off recorded in the v2 addendum (FR4).
  • VC14. v1 archive preserved (this rewrite's prerequisite). conductor/chronology.md.broken-v1 exists with the v1 218-line file; git log shows the rewrite is a continuation (commit 3aea92f1 "botched the chronology, going to rewrite the track."), not a re-do.

Risk Assessment

Risk Likelihood Scope impact Mitigation
R1: Classifier is too aggressive (false high confidence) medium Wrong status committed; user catches in Stage 3 FR7 quality gate (30% abort); per-row evidence makes the classifier's reasoning auditable; conservative bias (NFR7)
R2: Classifier is too conservative (>30% low) medium FR7 aborts → fallback to v1 manual protocol (Tier 1 reviews every row) The fallback is the user's "B" option (per chat 2026-06-21); explicitly designed in FR7
R3: Tier 1's resolutions are wrong (Stage 2) low User catches in Stage 3 Per-row resolution notes + evidence log make Tier 1's reasoning auditable; user's Stage 3 review is the final gate
R4: state.toml parsing fails (some folders lack state.toml) low Rows fall to "ambiguous" → low confidence → queued for review Classifier tolerates missing state.toml (FR5 §"3. Check state.toml phase progression"); "ambiguous" is the correct behavior per the conservative bias
R5: v1 archive move loses data low Minimal — git mv is safe Use git mv for the rename; verify with git log --follow after
R6: User disagrees with Tier 1's resolutions low Loops back to Stage 2 The user is the final gate (Stage 3); explicit Stage 3 review
R7: Summary extraction still picks metadata-field text (regression of v1 bug) low Row has bad summary v2's priority chain + regex rejection (^\*\*); tested by extended test suite (FR5 §"Tests extended")
R8: The 30% threshold is wrong (too low or too high) medium If too low: abort too easily. If too high: accept a bad classifier. The 30% value is the user's "A only if classifier is good" trade-off; if the user wants to adjust, FR7's wrapper script accepts --threshold as a CLI flag
R9: Evidence line format is too verbose (clutters the table) low User complains in Stage 3; loops back to FR1 The evidence line is a sub-line (not a column); the table remains 6 columns. If the user wants it more terse, FR1 can be revised.
R10: v1's broken chronology is referenced by other docs low Confusion between v1 and v2 conductor/chronology.md.broken-v1 is clearly labeled; the v2 file is chronology.md; the v1 report is extended with the v2 addendum that explains the rename

Execution Plan (high-level — see plan.md for worker-ready tasks)

  • Phase 1: Archive v1 + verify state of carried-forward work. Move conductor/chronology.mdconductor/chronology.md.broken-v1; reset state.toml to current_phase = 0; verify tracks.md pruning + workflow.md 3-step convention are intact.
  • Phase 2: Rewrite the helper script + extend tests (FR5). Rewrite _classify_status to use the 5-step git-history algorithm; add per-row confidence assignment; rewrite summary priority chain with regex rejection; add 8-10 new unit tests.
  • Phase 3: Add the quality gate script (FR7). New file scripts/audit/chronology_quality_gate.py; 5 new unit tests for the threshold logic.
  • Phase 4: Run the new classifier, generate v2 staging (FR6 Stage 1). Run the script; verify the staging file has per-row evidence + confidence + "Needs Review" section.
  • Phase 5: Quality gate (FR7). Run chronology_quality_gate.py; if PASS, proceed; if ABORT, fallback to manual review protocol.
  • Phase 6: Tier 1 reviews "Needs Review" queue (FR6 Stage 2). Tier 1 resolves each low-confidence row; updates the staging file with Tier 1's resolutions; updates the per-row evidence log.
  • Phase 7: Promote v2 staging → canonical (FR1). Rename chronology.md.stagingchronology.md; commit.
  • Phase 8: Write v2 addendum to migration report + end-of-track report (FR4 + VC9). Add the v2 rewrite section; document the v1 → v2 status diff + Tier 1 review log; write end-of-track v2 addendum.
  • Phase 9: User sign-off (FR6 Stage 3). User reviews v2 + evidence log + Tier 1 resolutions. Records sign-off in the v2 addendum.
  • Phase 10: Wrap-up. Mark track complete in tracks.md + state.toml; set status = "completed" in metadata.json.

See Also

  • docs/reports/CHRONOLOGY_TRACK_HANDOVER_20260620.md — the failure report; the source of the new classifier algorithm.
  • docs/reports/CHRONOLOGY_MIGRATION_20260619.md — v1 migration report; the v2 addendum extends it.
  • conductor/tracks.md:459 — the existing "lightweight chronology" reference that v2 formalizes.
  • conductor/workflow.md "Notes > Editing this file" — the existing archive convention; the 3-step convention (FR3) is appended here.
  • conductor/code_styleguides/feature_flags.md — "delete to turn off" convention; the helper script (FR5) follows it.
  • conductor/code_styleguides/data_oriented_design.md — applies: the chronology is data, the classifier is a transformation, the evidence log is a projection.
  • conductor/code_styleguides/error_handling.md — applies to the helper script: _classify_status returns (status, confidence, reason) (data-oriented "and/or" pattern).
  • docs/reports/TRACK_COMPLETION_tier2_autonomous_sandbox_20260616.md — precedent for one-page end-of-track reports.
  • AGENTS.md "File Size and Naming Convention" — the hard rule against creating new src/<thing>.py files; v2 doesn't touch src/.
  • AGENTS.md "Critical Anti-Patterns" — the no-day-estimates rule; the no-git restore ban; the report-instead-of-fix pattern (the handover IS a fix, not a report).
  • conductor/workflow.md "Tier 1 Track Initialization Rules" — the no-day-estimates rule followed in this spec.
  • conductor/workflow.md "Skip-Marker Policy" — applies: the v1 chronology's broken rows are not "skipped"; they are re-classified in v2.