manual_slop

Private

Public Access

Author	SHA1	Message	Date
conductor-tier2	9cc51ca9af	conductor(track): nagent review - deep-dive + 6 pitfalls + 10 actionable takeaways Reference/analysis track. Produces 0 code changes. Artifacts (conductor/tracks/nagent_review_20260608/): - spec.md (240 lines) - track wrapper with Application/Meta-Tooling framing - report.md (571 lines) - 14-section deep-dive; primary deliverable - comparison_table.md (79 lines) - flat side-by-side reference - decisions.md (286 lines) - 10 future-track candidates with priority matrix - nagent_takeaways_20260608.md (363 lines) - 10 actionable patterns grounded in code (file:line refs into nagent source and Manual Slop source) - metadata.json (132 lines) - structured metadata + verification criteria - state.toml (113 lines) - per-task tracking + user-corrections log (7 entries) 14 nagent principles covered in report.md (durable work, text-in/text-out, editable state, visible protocol, the loop, per-file memory, repo history, neighborhoods, sub-conversations, controlled writes, large files, tool discovery, framework differences, build your own). 6 pitfalls (revised from 8 after user-corrections): 1. No structured output protocol in Application AI (opaque function calling) 2. Provider-specific history in process globals (ai_client._anthropic_history + _deepseek_history + _minimax_history) 3. RAG is not 'history as data' (fuzzy, not auditable) 4. AI client is a stateful singleton (2,685-line ai_client.py) 5. No non-MMA disposable sub-conversations (1:1 gap; user-flagged want) 6. Hard-coded tool discovery (45-tool if/elif in mcp_client.py) User-corrections applied (3 rounds, 7 total corrections recorded): - Editable discussions: PARTIAL -> PARITY (DIFFERENT FOCUS) with full A1-A7 per-entry + B1-B11 discussion-level + C1-C5 undo/redo operation matrix - Per-file memory: DOMAIN MISMATCH -> MANUAL SLOP IS STRONGER IN CURATION DIMENSION (FileItem + ContextPreset vs nagent's inode-keyed conversation log; complementary, not equivalent) - Sub-conversations: MMA has it; 1:1 does not -> 'PARITY for MMA; GAP for 1:1 discussions' (user wants this) - RAG: opt-in, not gap; user wants pre-staging via sub-conversation - Personas: config bundling (can opt out via AI settings) - Tool discovery: deferred (user has 'intent based DSL' idea but 'no where near that ideation yet') 10 actionable takeaways (separate from the 6 pitfalls - those are diagnosis, these are prescription): 1. State visibility (UI inspector for in-process state) 2. Readable conversation log (text-greppable, not just JSON-L) 3. Sub-agents for 1:1 (HIGH priority - user-flagged) 4. File-identity over file-path (st_dev:st_ino rename-safe) 5. One loop shape visible in diagnostics 6. Visible retry on protocol failure 7. Meta-Tooling DSL (intent-based, deferred) 8. Self-describing tools (subsumed by mcp_architecture_refactor_20260606) 9. Single source of truth for disc_entries + provider history 10. Sub-agent return type constraint (bake into candidate #1 spec) Domain classification: every recommendation tagged Application / Meta-Tooling / Both per docs/guide_meta_boundary.md. nagent lives in the Meta-Tooling domain; Manual Slop's Application AI is a different kind of thing. No code modified by this track (reference/analysis only). All 7 files parse cleanly (JSON, TOML, Markdown). All internal cross-links resolve. Track is 'active' awaiting human review; future-track candidates live in decisions.md and nagent_takeaways_20260608.md.	2026-06-08 18:44:35 -04:00
ed	c9a991bbb8	test(live_workflow): bump project switch wait timeout 30s -> 120s The 30s wait_for_project_switch timeout was an excessive constraint. In batch context, prior sims' AI discussion turn workers saturate the 8-worker io_pool, queueing this switch for tens of seconds. The other defensive waits in the test (warmup 60s, prior switch 60s) already use 60s+, so 30s was the inconsistent outlier. User confirmed: 'I think not completing in 30s is an excessive constraint if thats whats going on.' Verification: - test_full_live_workflow isolation: 11.69s PASS - 7-test batch (test_full_live_workflow + 4 extended sims + 2 markdown): 85.83s PASS	2026-06-08 18:14:18 -04:00
ed	87d7c5bff2	test(io_pool): update assertion for 8-worker pool size	2026-06-08 17:51:39 -04:00
ed	4a33848620	fix(io_pool): increase worker count from 4 to 8 to prevent test hangs Root cause: test_full_live_workflow in batch context (with prior sims running AI discussion turns) would queue its _do_project_switch behind the auto-pruner's scan of tests/logs/ (154MB, 6519 files). The 4-worker pool was saturated, so the switch would never run within 30s. Fix: bump IO_POOL_MAX_WORKERS from 4 to 8. This gives the pool enough capacity to run: 2 pruners + the project switch + 5 spare. Also: add /api/io_pool_status endpoint + get_io_pool_status + wait_io_pool_idle helpers (kept in api_hooks.py and api_hook_client.py for the test_api_hook_client_io_pool.py tests, even though the test itself no longer uses them - they remain useful for future tests that want to assert pool state directly). Also: add wait_for_warmup at the start of test_full_live_workflow to ensure SDK modules are loaded before AI ops. Test verification: - test_full_live_workflow in isolation: 11.83s PASS - test_full_live_workflow in batch (with 4 prior sims): 83.46s PASS - 30/30 related unit tests PASS	2026-06-08 17:49:34 -04:00
ed	9afc93bce2	fix(app_controller): clear project-switch state in _handle_reset_session When a prior test in the tier-3-live_gui batch leaves a _do_project_switch background thread running, the next test's btn_project_new_automated click sees _project_switch_in_progress=True (from the prior thread) and queues the new path via _project_switch_pending_path. The queued switch is never actually submitted to the io_pool, so is_project_stale() stays True and AI ops (_handle_generate_send) bail with 'project switch in progress; AI ops disabled'. Fix: _handle_reset_session now also clears _project_switch_in_progress, _project_switch_pending_path, and _project_switch_error (under the existing _project_switch_lock). This way, even if the prior background thread is still running, the controller reports an idle state and the new switch can be submitted normally. Also: - src/api_hook_client.py: reverted wait_for_project_switch to require in_progress=False (was relaxed to return on queued path, which misled the caller into thinking the switch was done) - tests/test_handle_reset_session_clears_project.py: new test test_handle_reset_session_clears_project_switch_state asserts is_project_stale() returns False after reset - tests/test_api_hook_client_wait_for_project_switch.py: updated test_wait_for_project_switch_does_not_return_on_queued (in_progress + matching path should keep waiting, not return early) - tests/test_live_workflow.py: added pre-wait for any in-flight switch before doing btn_reset (so the test waits up to 60s for the prior switch to complete if needed) - conductor/todos/TODO_test_full_live_workflow.md: updated Task 4 with the deeper hang analysis and recommended fix Known follow-up: test_full_live_workflow still hangs in tier-3 batch even with this fix, because the new _do_project_switch itself is hung in the io_pool (likely saturation from prior sims' AI discussion turn workers). Deeper investigation required.	2026-06-08 15:19:30 -04:00
ed	5087ee988d	chore: move TODO_test_full_live_workflow.md to conductor/todos/ Following the conductor convention of organizing track-related artifacts under conductor/. The TODO tracks the test_full_live_workflow race condition fix and its follow-up items (Tasks 3, 7 still pending; known batch hang documented). Tasks 1, 2 (with regression fix), 4, 5, 6 are SHIPPED in prior commits.	2026-06-08 14:05:40 -04:00
ed	3391e18f64	chore(pyproject): register pytest.mark.live marker Silences the PytestUnknownMarkWarning emitted by test_visual_mma.py and test_visual_sim_gui_ux.py (3 instances). The @pytest.mark.live mark already exists in the test files; pyproject.toml just didn't know about it. - pyproject.toml: added 'live: marks tests as live visualization tests (not in CI by default)' to [tool.pytest.ini_options].markers	2026-06-08 13:59:18 -04:00
ed	d09f70ea44	docs(todo): mark Tasks 4+5 as SHIPPED; note known batch hang issue	2026-06-08 13:37:13 -04:00
ed	b6972c31de	test(live_workflow): use wait_for_project_switch + defensive file check Replaces the 10x1s blind poll of derived state with a condition-based wait on /api/project_switch_status. Also adds a defensive file existence check that fails fast (within 5s) if the click was dropped or the project creation handler crashed. The new wait surfaces a clear error message ('Project switch did not complete in 30s. Last status: ...') instead of the generic 'Project failed to activate', and exposes _project_switch_error if the controller reported one. - tests/test_live_workflow.py: replaced poll loop (lines 57-65) with wait_for_project_switch + os.path.exists defensive check	2026-06-08 13:26:54 -04:00
ed	a6605d9889	feat(api_hook_client): add wait_for_project_switch for deterministic test waits Adds a polling helper that blocks until the project switch completes, errors out, or times out. Replaces the fragile 10x1s blind poll in test_full_live_workflow with a condition-based wait on the /api/project_switch_status endpoint. Features: - Polls /api/project_switch_status every 200ms (configurable) - Returns immediately on error (with the error in the result) - Path matching: exact match OR basename match (handles absolute vs relative) - Times out with a clear 'timeout' flag instead of a generic assertion - Optional expected_path: if None, returns on any in_progress=False - src/api_hook_client.py: new wait_for_project_switch method (37 lines) - tests/test_api_hook_client_wait_for_project_switch.py: 6 unit tests with mocked _make_request covering all paths	2026-06-08 13:04:12 -04:00
ed	54e46ee815	docs(todo): note regression discovered and fixed in test_context_sim_live	2026-06-08 12:35:24 -04:00
ed	4548726a2b	conductor(tracks): restructure - chronological by phase + status groupings + active queue table	2026-06-08 12:26:56 -04:00
ed	e0a3eb8c05	fix(app_controller): regression in test_context_sim_live from clearing active_project_path Task 2 (_handle_reset_session reset) introduced a regression: setting self.active_project_path to empty caused an infinite re-switch loop in _do_project_switch because _flush_to_project writes to active_project_path (raises OSError on empty path), and the finally block re-submitted the failed switch on every iteration. Result: test_context_sim_live saw switching-to status for 5+ seconds and MD-only generation was blocked. Fix: keep self.active_project_path as-is in _handle_reset_session. Only reset self.project (to a fresh default_project dict) and self.project_paths (to empty list). The stale project state issue is solved by replacing the project dict; the active_project_path stays valid for _flush_to_project. - src/app_controller.py: refined _handle_reset_session project reset - tests/test_handle_reset_session_clears_project.py: updated contract test to assert active_project_path is preserved	2026-06-08 12:24:10 -04:00
ed	40d61bf3d8	docs(todo): mark Tasks 1+2 as SHIPPED for test_full_live_workflow fix	2026-06-08 10:15:54 -04:00
ed	6ecb31ea0a	feat(app_controller): reset project state in _handle_reset_session Stale project state from prior live_gui tests (shared session-scoped subprocess) was leaking into subsequent tests, causing the test_full_live_workflow race condition: 'Project not switched' errors when self.project still claimed to be a different project. The fix: _handle_reset_session now mirrors the default-project branch of __init__ (lines 1743-1745), creating a fresh default project dict, clearing active_project_path and project_paths, and reinitializing the workspace manager. - src/app_controller.py: 6 new lines in _handle_reset_session - tests/test_handle_reset_session_clears_project.py: 3 tests (active_project_path, project_paths, self.project)	2026-06-08 10:13:07 -04:00
ed	abb3856525	feat(api_hooks): add /api/project_switch_status endpoint for deterministic test signaling Adds a new endpoint that exposes the project-switch state machine so tests can poll for completion instead of guessing with timeouts. - AppController: track _project_switch_error on failure paths - src/api_hooks.py: GET /api/project_switch_status returns {in_progress, pending_path, active_path, error} - src/api_hook_client.py: get_project_switch_status() helper - tests/test_api_hooks_project_switch.py: 3 unit tests for client + endpoint shape, 1 live_gui test for the default-idle case	2026-06-08 09:55:36 -04:00
ed	c531cebe03	conductor(plan): review pass — fix cross-references, add NOT_READY + with_errors + Lottes/Valigo, split §3.4 into 8 sub-tasks	2026-06-08 09:38:27 -04:00
ed	8248a49f1e	docs(todo): simple todo list for fixing test_full_live_workflow race	2026-06-08 09:25:18 -04:00
ed	08ee7547be	docs(reports): root cause report for test_full_live_workflow race condition	2026-06-08 09:24:14 -04:00
ed	64823493c0	conductor(closeout): ship test_batching_refactor_20260606 with CLOSEOUT.md and follow-up recommendation	2026-06-08 08:36:22 -04:00
ed	488ae04459	fix(run_tests_batched): detect batch failure from output when proc.returncode is wrong	2026-06-08 02:03:50 -04:00
ed	5c6eb620a1	fix(run_tests_batched): colorize non-xdist format (tests/... STATUS), filter 'Error during log pruning' noise	2026-06-08 01:54:56 -04:00
ed	272b7841ae	fix(run_tests_batched): filter xdist scheduling queue output (test paths without status prefix)	2026-06-08 01:51:07 -04:00
ed	a2d16541d0	fix(run_tests_batched): keep pytest's full -v output, only filter LogPruner/win errors, colorize per-test status	2026-06-08 01:49:39 -04:00
ed	21cb57b31d	fix(run_tests_batched): graceful xdist fallback, live progress streaming, ANSI colors, absolute default paths	2026-06-08 01:28:53 -04:00
ed	fb6b4bd3eb	conductor(tracks): mark test_batching_refactor_20260606 as completed	2026-06-08 01:18:20 -04:00
ed	50bd894f8d	conductor(archive): ship test_batching_refactor_20260606 to archive	2026-06-08 01:16:58 -04:00
ed	50f26f0d5c	chore: delete legacy run_tests_batched.py (was preserved for one cycle)	2026-06-08 01:15:12 -04:00
ed	ac7e638b23	chore: gitignore tests/.test_durations.json (developer-local cache)	2026-06-08 01:14:51 -04:00
ed	9eac02ddcb	feat(tests): populate test_categories.toml with cross-cutting entries	2026-06-08 01:14:12 -04:00
ed	796eec0058	conductor(plan): mark Phases 2,3 complete in test_batching_refactor_20260606	2026-06-08 01:09:02 -04:00
ed	5252b6d782	docs(testing): document new run_tests_batched.py in Running Tests section	2026-06-08 01:00:50 -04:00
ed	e6ad2ecda2	chore: preserve old run_tests_batched.py as .legacy for one cycle	2026-06-08 00:59:49 -04:00
ed	2c3a0512f2	feat(run_tests_batched): full CLI with --tiers, --durations, actual pytest execution	2026-06-08 00:58:53 -04:00
ed	7610c9c1dc	conductor(plan): mark Phase 1 complete in test_batching_refactor_20260606	2026-06-08 00:53:59 -04:00
ed	57285d048b	feat(run_tests_batched): add --plan and --audit modes (Phase 1 stub)	2026-06-08 00:50:37 -04:00
ed	29ac64adc6	test(conftest): register tests.pytest_collection_order as pytest plugin	2026-06-08 00:49:11 -04:00
ed	f240504f0e	feat(collection_order): implement opt-in per-test sort via conftest hook	2026-06-08 00:47:21 -04:00
ed	6287005ad1	test(collection_order): add red tests for opt-in sort_items_by_order	2026-06-08 00:47:03 -04:00
ed	e07036ad5d	feat(batcher): implement Batch dataclass and plan() function	2026-06-08 00:46:12 -04:00
ed	246f293c56	test(batcher): add red tests for plan() function	2026-06-08 00:41:20 -04:00
ed	9c5ad3fb8d	config	2026-06-08 00:40:33 -04:00
ed	f778ef509e	feat(categorizer): implement load_registry, merge_registry, categorize_all	2026-06-08 00:33:21 -04:00
ed	2b56ab3c5c	conductor(track): initialize test_batching_post_refactor_polish_20260607 spec/plan/state	2026-06-08 00:27:32 -04:00
ed	828050ae4f	test(categorizer): add red tests for registry merge and full classification	2026-06-08 00:27:04 -04:00
ed	9e5fed56a5	feat(categorizer): implement subsystem/speed/batch_group inference	2026-06-08 00:22:22 -04:00
ed	7aaac7d586	test(categorizer): add red tests for subsystem/speed/batch_group inference	2026-06-08 00:21:03 -04:00
ed	b2e8cce9f6	feat(categorizer): implement auto_classify using AST scan (no regex)	2026-06-08 00:19:43 -04:00
ed	fb54737f45	test(categorizer): add red tests for auto_classify fixture_class rules	2026-06-08 00:16:18 -04:00
ed	dd48c095b8	refactor(tests): move test_categorizer library from scripts/ to tests/	2026-06-08 00:15:19 -04:00

1 2 3 4 5 ...