# Track Specification: Video Analysis De-obfuscation Campaign (2026-06-21) **Status:** Active (spec approved 2026-06-21) **Initialized:** 2026-06-21 **Owner:** Tier 1 Orchestrator (umbrella spec + synthesis); Tier 2 Tech Lead (per-track execution) **Priority:** A (user-blocking; Pass 2 of the 3-pass research campaign) **Type:** Multi-track research campaign (1 warmup + 1 umbrella + 3 phase children = 5 folders total) **Domain:** Meta-tooling (research deliverable + LLM operational spec; no `src/` changes) > **Purpose.** This umbrella organizes Pass 2 of the user's 3-pass research campaign: **de-obfuscation** of the Pass 1 video reports via the user's constructive type-theoretic re-encoding DSL. The de-obfuscation reduces standard math notation + verbose DSL/verbiage into a bounded, constructive, type-theoretic form that bridges the conceptual gap and crystallizes the formal language into the reader's mind. > **Multi-pass context.** Pass 1 produced 12 deep-dive reports (1000-10000 LOC each) + 1 cross-cutting synthesis. Pass 2 takes those and produces a multi-layer de-obfuscated version per video. Pass 3 (future, user-led) projects the de-obfuscated content to the user's applied domain (handmade/data-oriented/GPGPU + own caveats). > **Companion docs.** The warmup track (`video_analysis_deob_warmup_20260621/`) is the precursor that produces the initial lexicon + LLM prompt template. The 3 phase children (`video_analysis_deob_{lexicon,pilot,apply}_20260621/`) consume the warmup's output and apply it to the Pass 1 reports. --- ## 1. Overview ### 1.1 The user's de-obfuscation philosophy (foundational) The user curates knowledge unorthodoxy, especially formal math/sciences. Their position: | Position | Take | |---|---| | **Form requires bounds** | "To be known is to project a form." Boundedness is required for direct knowledge. | | **Indefinite is not directly knowable** | What is unbounded is indefinite; what is indefinite is indiscernible, unobserved, unsubject, unknowable. | | **Cycles/iteration/repetition are allowed** | Indefinite *operations* on bounded *forms* are expressible. `Stream A = nat -> A` is fine; `∞_val` is not. | | **The agent is bounded by necessity** | An agent is "envesseled in the soup of the universe," separated from the indefinite to discern. The agent cannot be indefinite. | | **Standard math notation is "noise"** | Too compressed, error-prone, ASCII-hostile, not programmatic, not verifiable, not visualizable. Lots of synonyms that mean the same thing (Curry-Howard: proofs=programs, types=propositions, etc.). | | **Constructive type theory is the foundation** | Proofs = programs (Curry-Howard); every value is a bounded form; operations are transformations. | | **Lexicon is etymology-aware** | Each term's word origin + definitional history is documented. Words are chosen to match modern subjective experience. | | **Inspiration** | Modern PL design — concatenative (Forth/KYRA/CoSy), data-oriented imperative (Lottes), immediate-mode DAG-building DSLs (O'Donnell's IMGUI). | ### 1.2 What Pass 2 produces For each of the 12 Pass 1 reports + 1 cross-cutting synthesis, Pass 2 produces a **3-layer de-obfuscated deliverable**: 1. **Translation** (`_translation.md`) — side-by-side table: original expression ↔ re-encoded form 2. **Replacement** (`_deobfuscated.md`) — the re-encoded form replaces the original; the report is read as a bounded, constructive, type-theoretic document 3. **Decoder index** (`_decoder.md`) — per-term decoder: form anchor, etymology, definition history, link to the original section Plus a per-track **pilot_report.md** or **apply_report.md** capturing lexicon refinements. ### 1.3 The 2-stage Pass 2 flow ``` Stage 1 (Warmup - precursor): Stage 2 (Apply - 3 phases): ┌─ Phase 1: Lexicon (refine warmup's draft) User's past notes ──► Warmup report.md + │ prompt_template.md ───────┤─ Phase 2: Pilot (apply to 2 videos, refine) │ └─ Phase 3: Apply (apply to 10 + synthesis) ``` --- ## 2. Current State Audit (as of 2026-06-21) ### 2.1 Already Available (DO NOT re-derive) | Asset | Location | Use in Pass 2 | |---|---|---| | Pass 1 reports (12 + 1 synthesis) | `conductor/tracks/video_analysis__20260621/report.md` + `summary.md` | The input to de-obfuscate | | Pass 1 transcripts + OCR | `conductor/tracks/video_analysis__20260621/artifacts/` | Source material for re-encoding context | | `intent_dsl_survey_20260612` report | `conductor/tracks/intent_dsl_survey_20260612/report_v1.2.md` | Sibling DSL (a tool-verb DSL for AI agents); not the math re-encoding, but shares the philosophy | | 4-tier vocab + 14-primitive grammar | `intent_dsl_survey_20260612/report_v1.2.md` §3, §4 | Reference for the PL-design vocabulary to use in the de-obfuscation DSL | | `conductor/code_styleguides/error_handling.md` | Conductor docs | `Result[T]` convention for any new Python tooling | | `conductor/code_styleguides/python.md` | Conductor docs | 1-space indent, type hints, no comments | | Reference scripts (bootslop) | `C:\projects\forth\bootslop\*.py` | yt-dlp / cv2 / winsdk OCR — NOT needed for Pass 2 (no video processing) | ### 2.2 Gaps to Fill (this campaign's scope) | # | Gap | Resolution | |---|---|---| | G1 | The user has no codified de-obfuscation DSL | Warmup produces `report.md` + `prompt_template.md` from the user's past samples | | G2 | The de-obfuscation lexicon is not yet finalized | Phase 1 (lexicon) refines the warmup's draft into a codified spec | | G3 | No pilot validation of the lexicon | Phase 2 (pilot) applies to 2 videos (1 foundational + 1 math-heavy) and captures refinements | | G4 | No application to the remaining 10 + synthesis | Phase 3 (apply) applies the refined lexicon to the remaining Pass 1 outputs | | G5 | No multi-layer deliverable structure | The 3-layer format (translation / replacement / decoder) is the new convention | --- ## 3. Goals 1. **Lexicon derived from the user's exemplars.** The de-obfuscation DSL is not invented from scratch; it is extracted from the user's past de-obfuscation notes via the warmup track. Evidence-based, not imposed. 2. **LLM-direct operational spec.** The de-obfuscation is performed by an LLM following the prompt template. The template is the "code" — the contract between the warmup and the apply phases. 3. **Lossless preservation (carries Pass 1's directive).** No Pass 1 concept is lost. The 3-layer output ensures every standard-math expression is represented (translation), replaced (replacement), and explained (decoder). 4. **Bounded, constructive, type-theoretic.** Every value is a bounded form. Iteration is explicit. "Infinity" is disambiguated lexically: `∞_val` (banned), `∞_proc` (allowed), `∞_card` (banned). 5. **Etymology + definitional history.** Each new term has a 1-line origin note + a 1-line definition history in the decoder. 6. **Multi-pass handoff.** Pass 3 (projection to applied domain) can consume the de-obfuscated outputs as its input. The handoff is clean: Pass 2 produces bounded, constructive forms; Pass 3 can apply them to the user's stylistic preferences. --- ## 4. Functional Requirements ### FR1. Umbrella folder + README **WHERE:** `conductor/tracks/video_analysis_deob_20260621/` **WHAT:** This folder contains the umbrella design (this spec) + 4 sibling files (`plan.md`, `metadata.json`, `state.toml`, `README.md`). The README is the index of the 4 sibling tracks (warmup + 3 phases) with their statuses. ### FR2. Warmup track (precursor) **WHERE:** `conductor/tracks/video_analysis_deob_warmup_20260621/` **WHAT:** Standalone research-style track. Produces: - `report.md` — the design philosophy + the curated lexicon (terms + re-encodings) + the 3 noise-dedup maps (Curry-Howard-style collapses) - `prompt_template.md` — the operational spec; an LLM can be prompted with this directly **Inputs:** The user provides samples in `samples/` (their past de-obfuscation notes). Format: markdown, txt, or any text the user has. **Process:** Tier 2 worker surveys the samples for term frequency, structural patterns, "form projection" heuristics, and noise-dedup maps. Produces a report + prompt template following the convention of `intent_dsl_survey_20260612/report_v1.2.md`. **`blocked_by`:** none (user must provide samples before warmup can start; the user's action item is the FIRST dependency). **`blocks`:** the 3 phase children (lexicon, pilot, apply) all depend on the warmup's output. ### FR3. Phase 1 — Lexicon refinement **WHERE:** `conductor/tracks/video_analysis_deob_lexicon_20260621/` **WHAT:** Consumes the warmup's `report.md` + `prompt_template.md`. Produces a codified `lexicon.md` (the operational spec for the de-obfuscation LLM) + `terms_catalog.md` (the machine-readable lexicon) + `dedup_map.md` (the 3 noise-dedup maps). The lexicon refinement adds: - Test cases (5-10 example transformations drawn from the user's samples) - A "form anchor" requirement: each re-encoding must project from an indefinite to a bounded form - Cross-references to the warmup's report sections ### FR4. Phase 2 — Pilot on 2 videos **WHERE:** `conductor/tracks/video_analysis_deob_pilot_20260621/` **WHAT:** Consumes the codified `lexicon.md` + 2 Pass 1 reports. The 2 pilot videos are: 1. `cs229_building_llms` (foundational ML/LLM coverage — wide scope, good test for "form projection" across many concepts) 2. `entropy_epiplexity` (math-heavy, focused on information-theoretic concepts — good test for "boundedness" + type-theoretic encoding of measure theory) For each pilot video, produces the 3-layer deliverable in `artifacts//`: - `translation.md` (side-by-side: original ↔ re-encoded) - `deobfuscated.md` (replacement: re-encoded form replaces the original) - `decoder.md` (per-term decoder: form anchor, etymology, definition history) Plus a `pilot_report.md` capturing: - Lexicon refinements discovered during the pilot - Concepts that didn't fit the lexicon (gaps) - Process improvements for Phase 3 ### FR5. Phase 3 — Apply to remaining 10 + synthesis **WHERE:** `conductor/tracks/video_analysis_deob_apply_20260621/` **WHAT:** Consumes the refined lexicon (from Phase 2) + 10 remaining Pass 1 reports + 1 cross-cutting synthesis. Produces the 3-layer deliverable for each, in `artifacts//`. Plus an `apply_report.md` capturing: - Final lexicon v2 - Final process refinements - Open questions for Pass 3 ### FR6. Multi-layer deliverable structure (per video) For each Pass 1 report, the de-obfuscation produces 3 files in `artifacts//`: **`_translation.md`** — side-by-side translation table: ```markdown # Translation: