32 KiB
Track Specification: Conductor Chronology v2 (2026-06-21 rewrite)
Overview
This is the v2 rewrite of chronology_20260619. The first run (Phases 1-9, 24 commits, 2026-06-19 to 2026-06-20) shipped conductor/chronology.md with a broken status classifier that read stale metadata.json.status fields. The user mandate — "EVERY SINGLE ENTRY MUST BE CROSS CHECKED" — was satisfied at a structural level (folder set == row set) but the semantic level (status correctness, summary quality) was not. Two classifier iterations followed (commits 4109a667 and 271e6895); both used heuristic-based fallbacks and neither used git history as the explicit evidence source the user wants.
This rewrite replaces the spec/plan/state.toml; the 24 prior commits + the broken v1 chronology remain in git history as the foundation. The substantive changes are:
- FR1 (chronology structure): rewritten — new status enum (5 values), per-row evidence line, per-row confidence level, "Needs Review" section.
- FR5 (helper script): rewritten — git-history classifier with confidence assignment.
- FR6 (cross-check): rewritten — 3-stage protocol (classifier auto + Tier 1 reviews "Needs Review" queue + user reviews final).
- FR7 (new): classifier quality gate — if > 30% of rows are ambiguous, abort to manual review (the user's "B" fallback).
Phases that produced the existing tracks.md pruning + workflow.md 3-step convention + the v1 migration report are reused. This rewrite adds a v2 addendum to the migration report.
Current State Audit (as of 2026-06-21, commit 3aea92f1)
Already Implemented (carried forward, NO REWORK)
conductor/tracks.md"Phase 9: Chore Tracks" section — pruned to one-line stub pointing tochronology.md(commitbe38dd5).conductor/tracks.md"Active Research Tracks"[x]entries — pruned (commitcca4767).conductor/tracks.md"Follow-up"[shipped]entries — pruned (commitb3a9c45).conductor/workflow.md"Notes > Editing this file" section — has the 3-step archiving convention (commitb697cd8).scripts/audit/generate_chronology.py— exists (338 lines). Functions:extract_slug_date,extract_summary,walk_track_folders,format_markdown,_classify_status,_parse_state_phase,_last_commit_date. The broken function is_classify_status(lines ~163-189) which reads thecurrentparameter (originally frommetadata.json.status) and uses folder-location + state_phase heuristics. This function is the target of FR5's rewrite.tests/test_generate_chronology.py— 6 unit tests, all passing against the current (broken) classifier. Need extension per FR5.conductor/chronology.md— 218 lines, 216 rows, v1 with broken status classifier. Statuses includeactive,spec_written,spec_approved,planning(stale metadata.json.status values). 41Completed, 0Abandoned, 167 rows with stale status per the handover report (line 14-16). Target of Phase 1's move-to-broken-v1.docs/reports/CHRONOLOGY_MIGRATION_20260619.md— v1 migration report; needs v2 addendum (FR4).docs/reports/CHRONOLOGY_TRACK_HANDOVER_20260620.md— tier-2's hand-off; documents the failure + the recommended fix (the 5-step git-history algorithm).docs/reports/TRACK_COMPLETION_chronology_20260619.md— v1 end-of-track report; needs v2 addendum.
Gaps to Fill (This Track's Scope)
| # | Gap | Where | Resolution |
|---|---|---|---|
| G1 | v1 chronology.md has 167/216 rows with wrong status (stale metadata.json.status values) |
conductor/chronology.md |
Move v1 to conductor/chronology.md.broken-v1 (Phase 1); generate v2 with git-history classifier (Phase 4) |
| G2 | v1 chronology.md has summaries that are metadata-field text (**Priority:** A..., **Date:** 2026-06-20) not the actual track summary |
Same as G1 | v2's priority chain (FR5 §"Summary extraction") rejects metadata-field text via regex |
| G3 | _classify_status reads stale metadata.json.status |
scripts/audit/generate_chronology.py:~163-189 |
Rewrite to use the 5-step git-history algorithm (handover §"Root cause of failure") |
| G4 | No "Needs Review" queue mechanism | n/a (new) | Add per-row confidence (FR5) + "Needs Review" section in chronology.md (FR1) |
| G5 | No quality gate to detect a bad classifier | n/a (new) | Add scripts/audit/chronology_quality_gate.py (FR7) |
| G6 | v1 cross-check was bulk-verified (structural check, not per-row semantic check) | n/a (process change) | v2 cross-check is 3-stage (FR6): classifier auto + Tier 1 reviews "Needs Review" + user reviews final with per-row evidence log |
| G7 | v1 per-row evidence is missing | n/a (new) | Add per-row evidence line to chronology.md (FR1) + standalone evidence log file (FR6 §"per-row evidence log") |
| G8 | state.toml is at current_phase = 10 with a false "complete" state |
conductor/tracks/chronology_20260619/state.toml |
Reset to current_phase = 0; this rewrite starts fresh |
| G9 | v1 migration report has 167 stale-status rows in the per-row log | docs/reports/CHRONOLOGY_MIGRATION_20260619.md |
v2 addendum shows the diff (v1 status → v2 status) with the git evidence per row |
| G10 | No fallback path if the classifier is bad | n/a (new) | FR7 quality gate; if > 30% ambiguous → abort to manual review (the user's "B" fallback per chat 2026-06-21) |
Goals
- One canonical index.
conductor/chronology.mdis the only file consulted to see "what has this project done." No more scanning 3 sections oftracks.md. (Carried from v1; unchanged.) - No info loss. Every track that has a folder in
conductor/tracks/orconductor/archive/has a row inchronology.md(or a documented exception). (Carried from v1; unchanged.) - Forward-compatible. When a new track ships, the convention is clear: move folder to
archive/, remove[x]fromtracks.md, add a row tochronology.mdwith the new format. (Carried from v1; unchanged.) - Git history is the explicit evidence. Each row's status is derived from
git log -- <folder>(commit count + commit messages).metadata.json.statusis informational only — the classifier does not trust it for the final status. - "EVERY SINGLE ENTRY" mandate preserved at the semantic level. Every row has: (a) a status decision, (b) the git evidence that supports the decision, (c) a per-row confidence level, (d) a "Needs Review" flag if confidence is low. The "cross-check" is the row's evidence trail, not a separate audit pass.
- Conservative classifier + hard quality gate. The classifier auto-classifies only when evidence is clear; ambiguous rows are flagged for human review. If > 30% of rows are ambiguous, the classifier is bad → abort to manual review (the user's "B" fallback per chat 2026-06-21).
- No day estimates. Per
conductor/workflow.mdTier 1 Track Initialization Rules (added 2026-06-16). Scope measured in files/sites.
Functional Requirements
FR1. conductor/chronology.md v2 structure (REWRITTEN)
WHERE: conductor/chronology.md (replaces v1).
WHAT: Same overall structure as v1 (table format, newest first, "Notable Non-Track Commits" section at the bottom), with these changes:
Status enum (5 values, replaces v1's 6-value enum):
Active— folder intracks/+ work has started (≥ 1feat/fix/refactorcommit) butstate.toml.current_phase< 3In Progress— folder intracks/+state.toml.current_phase≥ 3 (or nostate.toml+ ≥ 3 work commits)Completed— folder inarchive/+ ≥ 3 work commits (orstate.toml.current_phase == "complete")Abandoned— folder intracks/orarchive/+ 0-1 work commits + last commit > 14 days ago + nofeat/fix/refactorin commit historySpecial— explicit human-decision; e.g., research note, scratch dir, archived by mistake, deleted
Notably ABSENT from the v2 enum (present in v1): Shipped, Superseded, planning, spec_written, spec_approved, active (lowercase). The v2 enum is the canonical set; v1's status values are stale metadata leaks.
Per-row confidence level (NEW):
high— auto-classified by the script; git evidence + folder location + state.toml (if present) all point to the same statuslow— in the "Needs Review" queue; needs Tier 1 + user review
Per-row evidence line (NEW): Each row gets a sub-line in the format:
Evidence: <7-char-init-sha>..<7-char-end-sha> | N commits | state_phase=<N or "n/a" or "complete"> | "<first-commit-subject>" → "<last-commit-subject>" | confidence=<high|low>
"Needs Review" section (NEW):
At the bottom of chronology.md, a section listing all low-confidence rows with a one-line reason each. Format:
## Needs Review (Tier 1 + User)
These rows had ambiguous git evidence. Resolved by Tier 1; user reviewed in Stage 3.
- `<track_id>` (status=<resolved>) — <one-line reason> — resolved by Tier 1
Other v1 fields preserved unchanged: Date, Track ID, Summary (≤ 25 words), Folder, Range (<init-sha>..<end-sha> with commit count), Notable Non-Track Commits section.
Worked example (new format):
| 2026-06-19 | `chronology_20260619` | In Progress | **Confidence:** low | v2 rewrite of the chronology track after tier-2's failure report identified the broken status classifier. | `conductor/tracks/chronology_20260619` | `87923c93..3aea92f1` (12) |
| | | | | | Evidence: `87923c9..3aea92f` | 12 commits | state_phase=n/a (this rewrite) | "conductor(track): add initial spec for chronology_20260619" → "botched the chronology, going to rewrite the track." | confidence=low |
FR2. conductor/tracks.md pruning (CARRIED FORWARD; no changes)
Already complete in v1 (commits be38dd5, cca4767, b3a9c45). This rewrite verifies the pruning is intact and re-commits nothing.
Verification step: Phase 1 of the v2 plan runs grep -n "^- \[x\]" conductor/tracks.md and confirms 0 matches (other than the Status legend at the bottom of the file).
FR3. conductor/workflow.md 3-step convention (CARRIED FORWARD; no changes)
Already complete in v1 (commit b697cd8). This rewrite verifies the 3-step block is present and re-commits nothing.
Verification step: Phase 1 of the v2 plan runs grep -n "Archiving a track" conductor/workflow.md and confirms 1 match.
FR4. Migration report v2 addendum (UPDATED)
WHERE: docs/reports/CHRONOLOGY_MIGRATION_20260619.md (extends existing report).
WHAT: A new section appended to the end of the v1 report: "v2 Rewrite Addendum (2026-06-21)". Contains:
- Why the rewrite was needed — link to
CHRONOLOGY_TRACK_HANDOVER_20260620.md+ summary of the root cause - v1 → v2 status diff — table of all 216 rows showing the v1 status (stale) and v2 status (after the new classifier) + the git evidence per row
- Classifier confidence distribution — counts:
high/low/ total; % of total inNeeds Review - Tier 1 review log — for each
low-confidence row, the resolution note (assigned status + reason + override if any) - Quality gate result — was the 30% threshold hit? If so, the abort-to-B was triggered.
- Outstanding issues — any rows the user flagged for follow-up
FR5. Helper script rewrite — git-history classifier (REWRITTEN)
WHERE: scripts/audit/generate_chronology.py (rewritten) + tests/test_generate_chronology.py (extended).
WHAT: The script's _classify_status function is rewritten to use the handover's 5-step algorithm. The new signature is:
def _classify_status(
folder_link: str,
init_sha: str,
end_sha: str,
commit_count: int,
first_commit_subject: str,
last_commit_subject: str,
state_phase: str | None,
metadata_status: str | None,
last_commit_date: str,
) -> tuple[str, str, str]:
"""Classify a track's status using git history as primary evidence.
Returns:
(status, confidence, reason) where:
- status: one of "Active", "In Progress", "Completed", "Abandoned", "Special"
- confidence: "high" or "low"
- reason: one-line explanation of the classification
"""
The 5-step algorithm (per the handover §"Rewrite _classify_status to use git history as primary evidence"):
-
Count meaningful commits.
commit_count(already computed by the script viagit log --oneline -- <folder> | wc -l). 1-2 commits (just spec/plan creation) is a strong signal forActive(intracks/) orAbandoned(inarchive/). ≥ 3 work commits is a strong signal forCompleted(inarchive/) orIn Progress(intracks/). -
Inspect commit messages.
first_commit_subjectandlast_commit_subject(already extracted by the script). Classify each commit aswork(matches^(feat|fix|refactor|perf|test)\() ormeta(matches^(chore|docs|conductor)\() orother(everything else). -
Check
state.tomlphase progression.state_phaseis parsed fromstate.toml.current_phaseif the file exists; elseNone. The thresholds:state_phase == "complete"→Completed(high confidence if corroborated by git)state_phase >= 3→In Progress(high confidence if corroborated by git)state_phase in (0, 1, 2)→Active(high confidence if corroborated by git)state_phase is None→ no signal from state.toml; classifier relies on git + folder
-
Default to conservative. When git history is ambiguous (1-3 commits with no clear
workpattern), flag aslowconfidence → "Needs Review". The classifier NEVER auto-marksAbandoned— that's aSpecialdecision reserved for Tier 1 + user. -
Honour explicit metadata. If
metadata_statusisabandonedorsuperseded(orSpecial), and git evidence is not contradictory, trust the metadata. If git evidence contradicts metadata (e.g.,archive/+ 0 commits +metadata_status = "Completed"), the classifier flagslowconfidence and the user resolves in Stage 3.
Per-row confidence assignment:
high— git evidence + folder location + state.toml (if present) all point to the same status. Default for unambiguous cases.low— any of: (a) < 3 commits total, (b) conflicting signals (e.g.,archive/+ 0 commits + state_phase 0), (c) nostate.toml+ ambiguous git history, (d)metadata_statuscontradicts git.
Summary extraction (REWRITTEN priority chain): The v1 priority chain is replaced with a regex-aware version:
metadata.json.summaryif present and does not start with**(regex:^\*\*)- First non-empty line of
spec.mdthat does not start with** metadata.json.descriptionif not starting with**- First non-empty line of
plan.mdthat does not start with** - Generic placeholder:
"Imported from archive (no spec)"for archive rows,"Track folder (no spec found)"for tracks/ rows
The regex ^\*\* rejects metadata-field text like **Priority:** A..., **Date:** 2026-06-20, **Created:** 2026-06-19, **Initialized:** 2026-06-19, **Parent umbrella:** ..., **Confidence:** ....
New script: scripts/audit/chronology_quality_gate.py (FR7's wrapper).
- Reads the staging
chronology.md.stagingfile. - Counts
highandlowconfidence rows. - Computes
low_count / total_count. - If ratio > 0.30 → exit code 1, prints "ABORT: classifier is bad; >30% of rows are ambiguous. Fall back to manual review (v1 protocol)."
- If ratio ≤ 0.30 → exit code 0, prints "PASS: classifier is good. Proceed to Tier 1 review of 'Needs Review' queue."
Tests extended: the existing 6 tests stay; add 8-10 new tests covering:
_classify_statusreturns correct status for each (folder, commit_count, state_phase) combinationlowconfidence is assigned for ambiguous cases (1-2 commits, conflicting signals)highconfidence is assigned for unambiguous cases- Summary priority chain rejects metadata-field text (regression test for the v1 bug)
- The staging file has per-row evidence + confidence lines
- The "Needs Review" section is correctly populated
- The quality gate script exits 1 when > 30% ambiguous, 0 when ≤ 30%
- The quality gate script prints the correct summary
FR6. Per-row cross-check (REWRITTEN — 3-stage protocol)
WHERE: conductor/chronology.md v2 (after classifier run), then "Needs Review" queue (Tier 1 review), then final v2 (user review).
WHAT: The cross-check is 3-stage (replaces v1's single-stage Tier 1 review of every row):
Stage 1: Classifier auto-classification (script run).
- The script runs
walk_track_folders()overconductor/tracks/andconductor/archive/. - For each folder, the script extracts: date, track_id, init_sha, end_sha, commit_count, first_commit_subject, last_commit_subject, state_phase, metadata_status, last_commit_date, summary.
- The script's rewritten
_classify_status()assigns (status, confidence, reason) for each row. - Output:
conductor/chronology.md.stagingwith the per-row evidence line + confidence level + "Needs Review" section. - The script is READ-ONLY on the source folders; it writes to
chronology.md.stagingonly. - Quality gate (FR7) runs immediately after: if the gate passes, proceed to Stage 2; if the gate fails, the staging file is preserved and the task aborts to manual review (per FR7).
Stage 2: Tier 1 review of the "Needs Review" queue (only if quality gate passes).
- Tier 1 opens
conductor/chronology.md.staging. - Tier 1 filters to the "Needs Review" section (rows with
confidence=low). - For each
low-confidence row, Tier 1:- Opens the track's
spec.md(orplan.md/metadata.jsonif no spec). - Runs
git log --oneline -- <folder>and reviews the commit history. - Verifies the row's evidence line is accurate.
- Assigns a status from the 5-value enum (or flags for user decision).
- Writes a one-line resolution note (e.g., "Resolved: Active — work in progress, state_phase=2; classifier flagged low because no spec.md yet").
- Opens the track's
- Tier 1's defaults:
- In
tracks/+ ambiguous →Activewith a one-line note - In
archive/+ 0 commits →Specialwith note "archive folder with no work commits" - In
archive/+ ≥ 3 work commits + state_phase=0 (missing/incomplete) →Completedwith note "archive + N work commits; state.toml is stale" - Truly ambiguous →
Specialwith note "needs user decision; flagged in Stage 3"
- In
- After Tier 1 resolves all
low-confidence rows, the staging file is updated: the "Needs Review" section is moved to a "Tier 1 Resolutions" section showing each row's resolution note.
Stage 3: User review of final v2.
- User opens
conductor/chronology.md.staging(now with Stage 2 resolutions). - User reviews: (a) the format is correct, (b) every row has evidence + decision, (c) Tier 1's resolutions are reasonable, (d) nothing missed.
- User either approves (proceed to Phase 7 promotion) or requests changes (loop back to Stage 2 or 1).
The per-row evidence log (NEW FILE).
- Path:
tests/artifacts/chronology_v2_evidence_log.md(gitignored). - Format: one row per track with: track_id, status, confidence, init_sha, end_sha, commit_count, first_commit_subject, last_commit_subject, state_phase, classifier_reason, tier1_override (if any).
- Generated by the script during Stage 1; extended by Tier 1 during Stage 2; reviewed by the user in Stage 3.
FR7. Classifier quality gate (NEW)
WHERE: scripts/audit/chronology_quality_gate.py (new file) + tests/test_chronology_quality_gate.py (new tests).
WHAT: A wrapper script that runs after the classifier's Stage 1 output. The script:
- Reads
conductor/chronology.md.staging(the script's output). - Parses each row's confidence level.
- Counts
highandlowconfidence rows. - Computes
low_count / total_count. - If ratio > 0.30 → exit code 1, prints "ABORT: classifier is bad; >30% of rows are ambiguous. Fall back to manual review (v1 protocol). Tier 1 should manually review every row in the staging file."
- If ratio ≤ 0.30 → exit code 0, prints "PASS: classifier is good. rows need Tier 1 review; proceed to Stage 2."
The 30% threshold is a hard gate. Tier 1 doesn't start Stage 2 until the gate passes. If the gate fails, the staging file is preserved as chronology.md.staging.aborted and the task falls back to the v1 manual protocol (Tier 1 reviews every row).
Tests for the quality gate:
- Staging file with 0% low → exit 0
- Staging file with 30% low (boundary) → exit 0
- Staging file with 31% low → exit 1
- Staging file with 100% low → exit 1
- Staging file with malformed rows → exit 2 (parse error)
Non-Functional Requirements
(Carried from v1, mostly unchanged.)
- NFR1. Manually maintained. Per user choice (2026-06-19), the ongoing workflow is hand-edited. No auto-generation in CI; no script runs on every commit. The one-shot migration is a single event; the file is then edited like
tracks.md. - NFR2. Compact. Each row is ≤ 5 lines (the bullet + 3 sub-lines for Folder/Range/Evidence, OR a single condensed line for very old tracks where the folder is the only link). The file is scannable, not a wall of text.
- NFR3. Re-derivable. A reader can rebuild the chronology from
git log+ the track folders if needed. The init SHA + end SHA + evidence line in each row is the contract; the summary is the human-friendly gloss. - NFR4. No day estimates. Per
conductor/workflow.mdTier 1 Track Initialization Rules (added 2026-06-16). All scope is measured in files/sites. - NFR5. No TDD required for the chronology itself. This is a documentation/tooling track, not a feature track. The helper script (FR5) gets 8-10 new unit tests for the new classifier (TDD-required per project convention).
- NFR6. Evidence is auditable (NEW). The per-row evidence log (
tests/artifacts/chronology_v2_evidence_log.md) is human-readable; every classification decision is reproducible from the log + git history. A reader can verify any row's status by runninggit log -- <folder>and comparing to the evidence log. - NFR7. Classifier is conservative (NEW). When in doubt,
lowconfidence. The cost of a falselow(Tier 1 reviews it) is small; the cost of a falsehigh(wrong status committed without review) is high. The classifier's bias is towardlow.
Architecture Reference
docs/reports/CHRONOLOGY_TRACK_HANDOVER_20260620.md— the failure report; the source of the new classifier algorithm (5-step algorithm, §"Rewrite_classify_statusto use git history as primary evidence", lines 53-68).docs/reports/CHRONOLOGY_MIGRATION_20260619.md— v1 migration report; the v2 addendum (FR4) extends it.conductor/code_styleguides/data_oriented_design.md— applies: the chronology is data (one row per track), the classifier is a transformation (git history → status), the evidence log is a projection (data + decision + provenance).conductor/code_styleguides/error_handling.md— applies to the helper script: the script's_classify_statusreturns(status, confidence, reason)(a data-oriented "and/or" pattern, not an exception). The "Needs Review" queue is a recoverable case (low confidence), not an error.conductor/tracks.md:459— the existing "lightweight chronology" reference. v2 formalizes that role.conductor/workflow.md"Notes > Editing this file" — the existing convention for moving tracks toarchive/. The 3-step convention (FR3) is appended here.
Out of Scope
(Carried from v1, mostly unchanged.)
- Auto-generation on every commit. Per the user's "manual maintenance" choice (2026-06-19), there's no script that updates
chronology.mdautomatically. The file is hand-edited when a track is archived. - Tracking "in-flight" tracks in
chronology.md. In-flight tracks ([~]intracks.md) appear inchronology.mdwith statusActiveorIn Progress(per v2's enum). The active task list still lives intracks.md. - Tracking "planned but not specced" backlog items. These stay in
tracks.mdunder "Follow-up" and "Backlog". They aren't tracks until they have a folder. - Restructuring
tracks.mdbeyond[x]removal. The 3 sections that held[x]entries are now stubs (v1 Phase 3); no new structure is imposed. - A separate
chronology/folder for the file. The file lives at the conductor root (conductor/chronology.md), not in a subdirectory. - Reformatting existing
spec.md/plan.mdfiles. The migration reads from them; it does not modify them. - A web view of the chronology. It's a markdown file for in-repo reading. No GUI integration is in scope.
- A separate
chronology.md.draftworkflow (NEW for v2). v1 used.draftfiles; v2 doesn't. The classifier emits directly to a staging file (chronology.md.staging); the staging file is renamed tochronology.mdafter Stage 2 (Tier 1 review). The.stagingsuffix is gitignored.
Verification Criteria
For the track to be marked complete, ALL of the following must be true:
- VC1.
conductor/chronology.mdv2 exists with 216 rows; all 5 status values are used; per-row evidence line is present; per-row confidence level is present. - VC2.
conductor/tracks.mdpruning is intact (no regression from v1's pruning;grep -n "^- \[x\]" conductor/tracks.mdreturns 0 matches). - VC3.
conductor/workflow.md3-step convention is present (no regression;grep -n "Archiving a track" conductor/workflow.mdreturns 1 match). - VC4.
docs/reports/CHRONOLOGY_MIGRATION_20260619.mdhas the v2 addendum (per FR4). - VC5. Sorted newest first; every row has Folder + Range + Evidence lines.
- VC6. Every folder in
conductor/tracks/andconductor/archive/has a corresponding row, OR a documented exception in the v2 addendum. - VC7. "Notable Non-Track Commits" section is preserved (may be empty if no notable commits found).
- VC8. No new
src/*.pyfiles created (perAGENTS.mdFile Size and Naming Convention rule). - VC9. v2 addendum to
docs/reports/TRACK_COMPLETION_chronology_20260619.md(per project convention). - VC10. Classifier quality gate (FR7). The
scripts/audit/chronology_quality_gate.pyran; result was PASS (low confidence ≤ 30%). If the gate failed, the abort-to-B was triggered and Tier 1 manually reviewed every row. - VC11. "Needs Review" queue resolved (FR6 Stage 2). Every
low-confidence row in the staging file has a Tier 1 resolution note; the queue is empty in the finalchronology.md(Tier 1's resolutions are reflected in the per-row status). - VC12. Per-row evidence log (FR6).
tests/artifacts/chronology_v2_evidence_log.mdhas one row per track with status + confidence + evidence + decision (Tier 1 override if any). - VC13. User sign-off (FR6 Stage 3). User confirmed: format correct, every row has evidence, Tier 1 resolutions are reasonable, nothing missed. Sign-off recorded in the v2 addendum (FR4).
- VC14. v1 archive preserved (this rewrite's prerequisite).
conductor/chronology.md.broken-v1exists with the v1 218-line file;git logshows the rewrite is a continuation (commit3aea92f1"botched the chronology, going to rewrite the track."), not a re-do.
Risk Assessment
| Risk | Likelihood | Scope impact | Mitigation |
|---|---|---|---|
R1: Classifier is too aggressive (false high confidence) |
medium | Wrong status committed; user catches in Stage 3 | FR7 quality gate (30% abort); per-row evidence makes the classifier's reasoning auditable; conservative bias (NFR7) |
R2: Classifier is too conservative (>30% low) |
medium | FR7 aborts → fallback to v1 manual protocol (Tier 1 reviews every row) | The fallback is the user's "B" option (per chat 2026-06-21); explicitly designed in FR7 |
| R3: Tier 1's resolutions are wrong (Stage 2) | low | User catches in Stage 3 | Per-row resolution notes + evidence log make Tier 1's reasoning auditable; user's Stage 3 review is the final gate |
R4: state.toml parsing fails (some folders lack state.toml) |
low | Rows fall to "ambiguous" → low confidence → queued for review |
Classifier tolerates missing state.toml (FR5 §"3. Check state.toml phase progression"); "ambiguous" is the correct behavior per the conservative bias |
| R5: v1 archive move loses data | low | Minimal — git mv is safe |
Use git mv for the rename; verify with git log --follow after |
| R6: User disagrees with Tier 1's resolutions | low | Loops back to Stage 2 | The user is the final gate (Stage 3); explicit Stage 3 review |
| R7: Summary extraction still picks metadata-field text (regression of v1 bug) | low | Row has bad summary | v2's priority chain + regex rejection (^\*\*); tested by extended test suite (FR5 §"Tests extended") |
| R8: The 30% threshold is wrong (too low or too high) | medium | If too low: abort too easily. If too high: accept a bad classifier. | The 30% value is the user's "A only if classifier is good" trade-off; if the user wants to adjust, FR7's wrapper script accepts --threshold as a CLI flag |
| R9: Evidence line format is too verbose (clutters the table) | low | User complains in Stage 3; loops back to FR1 | The evidence line is a sub-line (not a column); the table remains 6 columns. If the user wants it more terse, FR1 can be revised. |
| R10: v1's broken chronology is referenced by other docs | low | Confusion between v1 and v2 | conductor/chronology.md.broken-v1 is clearly labeled; the v2 file is chronology.md; the v1 report is extended with the v2 addendum that explains the rename |
Execution Plan (high-level — see plan.md for worker-ready tasks)
- Phase 1: Archive v1 + verify state of carried-forward work. Move
conductor/chronology.md→conductor/chronology.md.broken-v1; resetstate.tomltocurrent_phase = 0; verifytracks.mdpruning +workflow.md3-step convention are intact. - Phase 2: Rewrite the helper script + extend tests (FR5). Rewrite
_classify_statusto use the 5-step git-history algorithm; add per-row confidence assignment; rewrite summary priority chain with regex rejection; add 8-10 new unit tests. - Phase 3: Add the quality gate script (FR7). New file
scripts/audit/chronology_quality_gate.py; 5 new unit tests for the threshold logic. - Phase 4: Run the new classifier, generate v2 staging (FR6 Stage 1). Run the script; verify the staging file has per-row evidence + confidence + "Needs Review" section.
- Phase 5: Quality gate (FR7). Run
chronology_quality_gate.py; if PASS, proceed; if ABORT, fallback to manual review protocol. - Phase 6: Tier 1 reviews "Needs Review" queue (FR6 Stage 2). Tier 1 resolves each
low-confidence row; updates the staging file with Tier 1's resolutions; updates the per-row evidence log. - Phase 7: Promote v2 staging → canonical (FR1). Rename
chronology.md.staging→chronology.md; commit. - Phase 8: Write v2 addendum to migration report + end-of-track report (FR4 + VC9). Add the v2 rewrite section; document the v1 → v2 status diff + Tier 1 review log; write end-of-track v2 addendum.
- Phase 9: User sign-off (FR6 Stage 3). User reviews v2 + evidence log + Tier 1 resolutions. Records sign-off in the v2 addendum.
- Phase 10: Wrap-up. Mark track complete in
tracks.md+state.toml; set status = "completed" inmetadata.json.
See Also
docs/reports/CHRONOLOGY_TRACK_HANDOVER_20260620.md— the failure report; the source of the new classifier algorithm.docs/reports/CHRONOLOGY_MIGRATION_20260619.md— v1 migration report; the v2 addendum extends it.conductor/tracks.md:459— the existing "lightweight chronology" reference that v2 formalizes.conductor/workflow.md"Notes > Editing this file" — the existing archive convention; the 3-step convention (FR3) is appended here.conductor/code_styleguides/feature_flags.md— "delete to turn off" convention; the helper script (FR5) follows it.conductor/code_styleguides/data_oriented_design.md— applies: the chronology is data, the classifier is a transformation, the evidence log is a projection.conductor/code_styleguides/error_handling.md— applies to the helper script:_classify_statusreturns(status, confidence, reason)(data-oriented "and/or" pattern).docs/reports/TRACK_COMPLETION_tier2_autonomous_sandbox_20260616.md— precedent for one-page end-of-track reports.AGENTS.md"File Size and Naming Convention" — the hard rule against creating newsrc/<thing>.pyfiles; v2 doesn't touchsrc/.AGENTS.md"Critical Anti-Patterns" — the no-day-estimates rule; the no-git restoreban; the report-instead-of-fix pattern (the handover IS a fix, not a report).conductor/workflow.md"Tier 1 Track Initialization Rules" — the no-day-estimates rule followed in this spec.conductor/workflow.md"Skip-Marker Policy" — applies: the v1 chronology's broken rows are not "skipped"; they are re-classified in v2.