conductor(deob_warmup): Distinguish principled scheme from user-specific preferences (6 surgical edits)

Per user 2026-06-23 review: the Tier 2 over-cited the user's specific implementations (Sectored Language V1, LLM session patterns, GA reinterpretations, classical Greek/Latin) as the canonical scheme, when they should be optional output conventions.

Changes:
1. report.md §3.4 — added Reading guide: Tier 4 mixes principled re-encodings (from the 5 rules) with user-specific re-encodings (from samples). The principled forms are scheme-canonical; the user-specific are optional output conventions.
2. report.md §3.5 — added Reading guide: Sectored Language operator terms are USER preferences, not scheme-canonical. The scheme produces principled re-encodings; the Sectored Language is one way to express them.
3. report.md §4.4 — added Reading guide: 'Real = Imaginary = Bivector' is the user's GA reinterpretation, not a scheme-canonical dedup. The principled forms are bivector (with grade annotation) + quantity(<value>) : <encoding>.
4. report.md §6.2 — added Reading guide: 4-layer output format is OPTIONAL (the user's preferred convention for etymological trails). The scheme's baseline is the 3-layer format.
5. prompt_template.md 'Your role' — removed 'Construct, not Invent' (was a user preference, not scheme-canonical). Added a 'Scheme-canonical vs. user-specific' bullet that makes the distinction explicit.
6. prompt_template.md 'The Sectored Language Operator Names' — labeled OPTIONAL; added Reading guide explaining it's one of several ways to express the scheme's principled re-encodings.
7. prompt_template.md verification checklist — replaced 'Sectored-language-named' with 'User-specific conventions applied only when appropriate'.

Phase 1 (lexicon child) will formalize this distinction further (e.g., moving §3.5 to Appendix B, marking each user-specific entry with [user-also-accepted]). The principled spine (5 rules + 6 noise-dedup maps + form-anchor examples + etymology rule + lossless preservation) is intact.
This commit is contained in:
ed
2026-06-23 15:39:16 -04:00
parent 6f21df7c7b
commit c7b6c6c920
2 changed files with 14 additions and 6 deletions
@@ -11,10 +11,10 @@ You are a de-obfuscator. Your task: take a Pass 1 report (full of standard math
Your operational stance:
- **Library Specification > Philosophy** (per Cluster 0, Pattern 9): prefer executable, debuggable, deterministic specifications over intuition pumps.
- **Decompression > Compression** (per Cluster 0, P1): the first step for any math is to decompress it.
- **Construct, not Invent** (per Cluster 0, Pattern 3): use the user's pseudo-code DSL, not free-form prose.
- **Bounded form required** (per §1.1 of `report.md`): no `∞_val`; use `Stream A = nat -> A` for processes.
- **Form anchor required** (per §5): every re-encoding has a form anchor — "what bounded form does this project from the indefinite?"
- **Honest epistemic hedging** (per §1.10): if uncertain, flag it; do not guess.
- **Scheme-canonical vs. user-specific:** the de-obfuscation scheme's principled re-encodings come from the 5 rules (boundedness, form-anchor, etymology, lossless, encoding-explicit). Some terms ALSO have user-specific preferred forms (Sectored Language V1, GA reinterpretations, classical Greek/Latin); these are *optional* output conventions documented in `report.md` §3.4 + §3.5, not scheme-canonical. Apply the user-specific forms only when the reader would prefer them.
## Input
@@ -167,9 +167,11 @@ Reference: `report.md` §3 for the full lexicon (~70 terms after Phase 1 expansi
- **Tier 3 (Type-theoretic primitives, 18 terms):** `Type``kind`; `Type of types``Kind`; `Constructor``intro`; `Eliminator``elim`; `Computation rule` (value-level) → `comp`; `Type-level Computation``getType(...) === T`; `Pair<A, B>` with `Build<A>/Build<B>`; `Dependent<x : A>(B)`; `lambda.x.M`; `objects : m : A, n : B ;`; etc.
- **Tier 4 (AI-fuzzing tolerance, 21 terms):** "invent" → `construct`; "real number" → `encodable quantity`; "imaginary number" → `bivector`; "dot product" → `length-projection product` (or `'scalar product'`); "cross product" → `wedge product`; "anti-wedge" → `regressive product` / `contraction` / `interior product`; "negative" → `F²` operator; "infinity" → **BANNED**; "point" → `Punctum` / `σημεῖον`; "kernel" (cross-domain) → `discrete subsystem that holds a continuous process up`; "Bourbaki" / "Standard GA" → **FOIL**; etc.
## The Sectored Language Operator Names (per `report.md` §3.5, from Cluster 9)
## The Sectored Language Operator Names (per `report.md` §3.5, from Cluster 9) — OPTIONAL
For linear algebra and CAS, use the Sectored Language naming:
> **Reading guide.** This is the **user's preferred output convention** for linear-algebra and CAS operations (per Cluster 9, the FGED V1 .sectr files). It is OPTIONAL — the de-obfuscation scheme does not require it. The scheme's principled re-encodings (e.g., `scalar product`, `magnitude`, `normalize`) are what the LLM produces; the Sectored Language names are one of several ways the reader can express those principled re-encodings. Apply the Sectored Language names when (a) the user requests it, (b) the term appears in a context where the user's prior de-obfuscation work used Sectored Language, or (c) the reader's preference is to use Sectored Language output. Otherwise, use conventional math with explicit type annotations.
For linear algebra and CAS, the user's preferred Sectored Language naming is:
- `magnitude(v)` for `||v||`
- `normalize(v) -> UnitVector` for unit vector
- `transpose(M) -> Matrix` for matrix transpose
@@ -203,7 +205,7 @@ After producing the 3 files, verify each:
- [ ] **Etymology-cited** — every new term has the 1-line origin + 1-line definition history.
- [ ] **Form-anchored** — every re-encoding has a form anchor.
- [ ] **Noise-deduped** — the 6 noise-dedup maps applied where applicable.
- [ ] **Sectored-language-named** — linear algebra and CAS use the Sectored Language names (per §3.5).
- [ ] **User-specific conventions applied only when appropriate**Sectored Language names + classical Greek/Latin/Sanskrit forms + GA reinterpretations are USER preferences, not scheme-canonical. Apply them only when the reader would prefer them (per §3.4 reading guide + §3.5 note).
- [ ] **EPP-formatted** — the fully-expanded pseudo-code follows the EPP format (per Cluster 1, Pattern 5).
- [ ] **Univalence-flagged** — when the LLM encounters the univalence axiom, flag it as "compression (lossless?**)" so the user can opt in or out.
@@ -277,6 +277,8 @@ The lexicon is the heart of the de-obfuscation. It is organized in 4 tiers. The
### §3.4 Tier 4: AI-fuzzing tolerance terms (22 terms, expanded in Phase 1)
> **Reading guide.** This tier mixes **principled re-encodings** (derived from the 5 load-bearing rules) with **user-specific re-encodings** (the user's personal preferences, drawn from Cluster 0/1/6/7/8/9 — including the Sectored Language V1, the GA reinterpretations, and the classical Greek/Latin/Sanskrit forms). The principled entries are the scheme's canonical forms; the user-specific entries are *one possible output convention* the scheme can produce. The 4-layer output format (§6.2) is similarly optional. For each entry below, the "Re-encoded" column is the form the *scheme* produces; entries that ALSO accept the user's preferred output are marked **[user-also-accepted]**. Phase 1 (lexicon child) will formalize this distinction.
| # | Conventional (fuzzy) | Re-encoded (precise) | Source cluster |
|---|---|---|---|
| 4.1 | "invent" | `construct` | Cluster 0 |
@@ -306,7 +308,7 @@ The lexicon is the heart of the de-obfuscation. It is organized in 4 tiers. The
### §3.5 Sectored Language operator terms (Phase 1, from Cluster 6 + Cluster 9)
These are the user's preferred operator names from the Sectored Language (per Cluster 9, the FGED V1 .sectr files):
> **Reading guide.** This section documents the **user's preferred output convention** for linear-algebra and CAS operations — the Sectored Language V1 (FGED V1) naming. The de-obfuscation scheme does NOT require this convention. It is one example of how the scheme's principled re-encodings (e.g., `scalar product`, `magnitude`) can be realized in an executable form. Other readers may use different conventions (Standard GA, conventional math with explicit type annotations, etc.); the scheme's output is the re-encoded form, not the Sectored Language names. The 7 example transformations in §7 demonstrate how the scheme produces these specific names. **Phase 1 (lexicon child) will move this table to Appendix B ("User's preferred output conventions, optional") to make the principled/user-specific distinction explicit.**
| Conventional | Sectored Language name | Source |
|---|---|---|
@@ -360,6 +362,8 @@ All 6 are documented below. Sources: Cluster 0, 2, 3, 4, 7, 8 (see report §4 of
### §4.4 Map 4: "Real" = "Imaginary" = "Bivector" (geometric algebra)
> **Reading guide.** This map is the **user's geometric-algebra reinterpretation** (per Cluster 0, Pattern 2; Cluster 8 GeoAlg). It is the user's preference, not a scheme-canonical dedup. The scheme's principled form for "imaginary number" is `bivector` (or `Bivector` with explicit grade annotation); "real number" re-encodes to `quantity(<value>) : <encoding>` (per Rule 5). The collapse of "real" + "imaginary" into the GA framework is the user's contribution; the scheme's job is to produce the principled re-encoding, and the user (or the LLM at their direction) can additionally adopt the GA collapse if they want.
**Source cluster:** Cluster 0 (Pattern 2 — Descartes-rejection / Clifford-affirmation); Cluster 8 (GeoAlg — `Point`, `Circle`, `Line`, `Plane` types).
### §4.5 Map 5: "Invent" = "Create" = "Imagine" → "Construct"
@@ -423,7 +427,9 @@ The etymology rule is the second operational requirement, derived from §1.5.
If the term is a user coinage, the etymology is the user's reason for the coinage, and the definition history is "user-specific; see `samples/`."
### §6.2 The 4-layer output format (per Cluster 7)
### §6.2 The 4-layer output format (per Cluster 7) — OPTIONAL
> **Reading guide.** This 4-layer format is the **user's preferred output convention** for terms with rich etymological trails (per Cluster 7, `Notiones.txt`). It is NOT a scheme-canonical requirement. The de-obfuscation scheme's baseline output is the 3-layer format (§5.2): translation (side-by-side) + replacement (re-encoded) + decoder (per-term). The 4-layer form adds an optional etymological layer (Greek + Latin + English + Sanskrit) that the LLM can apply **if the user requests it** or **if the term has a rich etymological trail**. Readers who don't care about Greek/Latin/Sanskrit get the 3-layer output; readers who do get the 4-layer output. The deob-warmup's `prompt_template.md` should make this an OPTION, not a REQUIREMENT.
For terms with rich etymological trails, the output should be **trilingual or 4-language** (per Cluster 7, Pattern 6):
1. **Original** (e.g., the Greek or Latin).