diff --git a/conductor/tracks/video_analysis_platonic_intelligence_kumar_20260621/report.md b/conductor/tracks/video_analysis_platonic_intelligence_kumar_20260621/report.md new file mode 100644 index 00000000..1c8799dc --- /dev/null +++ b/conductor/tracks/video_analysis_platonic_intelligence_kumar_20260621/report.md @@ -0,0 +1,1563 @@ +# Towards a Platonic Intelligence with Unified Factored Representations + +**Source:** https://youtu.be/1mXUFweWOug +**Author:** Akarsh Kumar (MIT CSAIL) +**Cluster:** B (Platonic / geometric AI representations) +**Slug:** platonic_intelligence_kumar +**Track:** Child #5 of `video_analysis_campaign_20260621` +**Date:** 2026-06-21 +**Pass:** 1 of 3 (research-only deep-dive) + +--- + +## 1. TL;DR + +This talk presents a **position paper** arguing that conventional deep learning (SGD-trained neural networks) discovers **Fractured Entangled Representations (FER)** — internal weight configurations that produce the right output behavior but do not capture the underlying regularities of the world. The author proposes that an alternative training paradigm — **open-ended search** — discovers **Unified Factored Representations (UFR)**: internal configurations where each "neuron" or weight corresponds to a semantically meaningful, semantically independent factor of variation. + +The evidence draws from three sources: + +1. **Picbreeder** (Stanley & Lehman, MIT CSAIL & collaborators): a 2008-era web app where humans evolved CPPNs (Compositional Pattern Producing Networks) by selecting images they liked. The CPPNs that emerged have factored weight structures where single neurons correspond to human-interpretable axes (x-axis symmetry, mouth opening, eye width). When the same image is reproduced by SGD-trained MLPs, the resulting weights are entangled — perturbing any one weight affects many semantic aspects of the output. + +2. **LLM failure modes** (FER in large language models): the brittleness of LLM mathematical reasoning (e.g., GPT-3 fails when chicken/duck/geese replace pencil/pen/eraser), counterfactual task degradation (Reasoning or Reciting paper), and the bag-of-heuristics arithmetic in Claude 3.5 Haiku (Anthropic's circuit-tracing paper) all suggest FER — the LLM is replicating patterns from training data without capturing the underlying regularities. + +3. **The Platonic Representation Hypothesis** (Huh et al. 2024): different modalities (vision, language) trained on different objectives are converging to a shared statistical model of reality. The author agrees this is a real statistical phenomenon but argues it does not address whether the converged representations capture the underlying regularities. + +The proposed solution is **open-ended search**: a learning paradigm that exhibits complexification (building regularities on top of regularities, bottom-up), emergence (higher-level structure from lower-level dynamics), adaptability (training for the ability to adapt, not for any specific task), and serendipity (exposure to multiple environments in sequence, not a single fixed dataset). + +**Cross-cluster position:** Sits in cluster B and bridges to clusters A (math foundations: score_dynamics_giorgini's framework is an alternative route to capturing regularities via score matching; entropy_epiplexity's information-theoretic perspective defines what "regularity" means), C (complex systems: open-endedness is a property of complex systems theory), and E (applied LLMs/diffusion: the FER diagnosis predicts failure modes observed in cs336_architectures and creikey_dl_cv). + +--- + +## 2. Key Concepts + +Twenty concepts form the conceptual spine of the talk. Each is developed in §5 with full mathematical or conceptual statement. + +### 2.1 The world has structure + +The world is not random. Structure manifests as: +- Self-similarity across spatial scales (fractals, multi-scale phenomena). +- Physics symmetries: translation, rotation, scale invariance of physical laws. +- Persistent objects: matter does not randomly appear/disappear; objects have identity over time. +- Common patterns across many objects: the same shapes, materials, dynamics recur. + +This is the **premise** of the talk — without structure in the world, the proposal would have no traction. + +### 2.2 Platonic Space of Forms (philosophical premise) + +Following Plato, the structure in the world is inherited from a **space of forms**: an abstract space containing the ideal patterns that objects instantiate. Properties common to many objects (e.g., "circularity," "verticality," "symmetry") are abstracted into the forms. The empirical regularity of the world reflects the underlying form structure. + +The author uses "Platonic" as a metaphor for **the convergence of representations across modalities and tasks** — see §2.20. + +### 2.3 Intelligent agents must capture structure + +To control the world (achieve goals), an agent must understand the world. To understand the world, the agent must capture its structure. **More specifically: the internal representations of the agent's mind/brain must capture the structure.** + +This is the **bridge claim** from the philosophical premise (§2.2) to the AI research program (§2.6+). It is the central thesis that the rest of the talk develops. + +### 2.4 Architectural inductive biases as partial solutions + +Modern deep learning has succeeded at capturing some structure via **inductive biases baked into the architecture**: +- **Translation invariance** → CNNs (convolutional architecture). +- **Permutation invariance over sets** → attention (transformer architecture). +- General field: **symmetry learning / geometric deep learning**. + +This is a partial solution: it captures **known regularities** via architecture design. **What about all the other regularities** (lighting invariance, material invariance, articulation, etc.)? + +### 2.5 The SGD fallback + +For unknown regularities, the current paradigm is: **train on lots of data with SGD and hope the network learns the regularities implicitly.** + +The author labels this a "fallback" because there is no theoretical reason to believe SGD discovers factorized representations rather than entangled ones. + +### 2.6 Does SGD actually capture regularities? (Empirical question) + +Yes, somewhat. ChatGPT, self-driving cars, and other large-scale systems exhibit real understanding of the world. They generalize across lighting, viewpoints, contexts. So the SGD paradigm is not a complete failure. + +But this brings us to the **position paper** hypothesis. + +### 2.7 The Fractured Entangled Representation (FER) hypothesis + +> Conventional SGD training in deep learning finds neural representations which are **fractured** and **entangled**. + +**Fractured** = the internal state does not correspond to a smooth, organized decomposition of the input. **Entangled** = multiple semantic aspects are mixed together in each weight or activation. + +The same input-output mapping can be implemented by many internal representations; SGD tends to find representations that satisfy the loss but do not decompose the world in semantically meaningful ways. + +### 2.8 Why FER is a problem + +If internal representations are entangled: +- **Generalization** to novel situations is brittle (small input changes produce semantically unrelated output changes). +- **Creativity** (composing new combinations of known concepts) is limited. +- **Continual learning** (acquiring new tasks without forgetting old ones) is hard because new updates entangle with old structure. + +These three failure modes are exactly the **jagged intelligence** observed in modern LLMs (good at IMO problems, bad at booking hotels). + +### 2.8a Jagged intelligence in LLMs + +Modern LLMs exhibit "jagged intelligence": +- Excellent at hard reasoning tasks (e.g., IMO gold medal-level math). +- Unreliable at simple tasks (e.g., reliably booking a hotel or plane). + +This is **inconsistent** with a system that has captured the underlying regularities. A system with true understanding should be uniformly competent across difficulty levels. The author attributes this to FER. + +### 2.9 The Unified Factored Representation (UFR) goal + +A representation is **unified** if it consistently captures the same factor across instances (the same neuron activates for "x-axis symmetry" regardless of which image the symmetry is in). A representation is **factored** if each factor is independent (changing one factor does not change others). UFR = unified + factored. + +The goal of open-ended search is to discover UFR. + +### 2.10 Compositional Pattern Producing Networks (CPPNs) + +A CPPN is a neural network with **heterogeneous activation functions** at each node (sin, cos, gaussian, sigmoid, abs, etc.). Given a coordinate (x, y), the CPPN outputs an image. CPPNs produce images with regular structure (symmetries, smoothness) by construction because the activations are themselves regular. + +CPPNs are the canonical representation used in **Picbreeder** and other neuroevolution experiments. They are **not** trained by SGD; they are evolved. + +### 2.11 Picbreeder (the experimental evidence) + +Picbreeder (Secretan et al. 2008, Stanley & Lehman): an online website where humans **interactively evolve** CPPNs by selecting images they like. Users are not given any objective; they simply choose images they find interesting. + +Outcomes: +- Users **discover** complex images (skulls, butterflies, apples, lamps, aliens, etc.) without any target. +- The discovered CPPNs have **factored structure**: each neuron corresponds to a distinct semantic axis (x-symmetry, y-symmetry, mouth opening, eye width, jaw width, etc.). +- Sweeping a single weight value produces a **semantically coherent variation** (e.g., mouth opens wider; apple size grows). + +This is **empirical proof that open-ended search can discover UFR** — in a toy domain (CPPN-generated images), but the principle generalizes. + +### 2.12 Why Greatness Cannot Be Planned (Stanley & Lehman 2015) + +The book by Stanley and Lehman formalizes insights from Picbreeder and similar experiments. Key concepts: +- **Deception**: fitness functions that reward intermediate progress can lead evolution away from the global optimum. +- **Serendipity**: stepping stones that have nothing to do with the final goal can be valuable. +- **Open-Endedness**: search without a fixed objective tends to discover more interesting solutions than search with one. + +This book is the philosophical backbone of the talk. + +### 2.13 Intriguing properties of Picbreeder + +Three properties that emerge from open-ended search: + +1. **Open-Ended**: the search never converges because the interesting region of CPPN space is unbounded. +2. **Serendipitous exaptation**: traits evolved for one purpose get repurposed (e.g., a step toward the teapot enables the skull). +3. **Emergence of evolvability**: certain CPPNs are more evolvable — small mutations produce large but coherent changes (canalization, regularity, modularity, symmetry). + +### 2.14 Evolvability and canalization + +**Canalization** = a genotype where small mutations produce small phenotypic changes (the phenotype is buffered against genetic variation). Waddington's concept, applied to CPPNs. + +In Picbreeder, evolved CPPNs exhibit canalization: certain mutations only affect one semantic aspect. SGD-trained MLPs do not — mutations affect multiple aspects. + +The author argues this canalization IS the structure of the world being captured in the representation. + +### 2.15 Layerization (CPPN → MLP conversion) + +Given any CPPN, there exists an MLP that computes the same function (universal approximation). **Layerization** is the process of converting a CPPN to an MLP. + +The layerization of an evolved Picbreeder CPPN is an MLP with **factored weight structure**: single weights correspond to single semantic axes. The layerization of an SGD-trained MLP is a similar MLP but with **entangled weights**. + +This is the central experimental result: **the same function, two different internal organizations, with different generalization properties.** + +### 2.16 SGD skull vs Picbreeder skull + +Demonstration: +- Train an MLP with SGD to reproduce the Picbreeder skull image. SGD succeeds (perfect reconstruction). +- Sweep each weight value in the SGD-trained MLP. Output changes are chaotic — no single weight corresponds to a single semantic axis. +- Layerize the Picbreeder CPPN. Sweep each weight value. Output changes are semantic — weight #i controls mouth opening, weight #j controls eye width, etc. + +Same loss, same architecture, two completely different internal organizations. SGD gives FER; open-ended search gives UFR. + +### 2.17 FER in LLMs — three pieces of evidence + +The author surveys three papers showing FER in modern LLMs: + +**A. GSM-Symbolic (Mirzadeh et al. 2025, Apple).** Replace specific objects in GSM8K math problems. GPT-3 gets pencil/pen/eraser sums correct ("9 things") but fails on chicken/duck/geese sums (counts the animals, not the items — gets 10 instead of 9). Same arithmetic structure, different surface tokens → different answer. The model learned the GSM8K surface pattern, not the underlying counting. + +**B. Reasoning or Reciting (Wu et al. 2024, MIT/BU).** Counterfactual versions of reasoning tasks (base 9 arithmetic, code execution with 0-based vs 1-based indexing, drawing tasks, etc.). Performance drops substantially — the model doesn't generalize to systematic perturbations of the task structure. + +**C. On the Biology of a Large Language Model (Anthropic 2025).** Circuit tracing of Claude 3.5 Haiku on 36 + 59 = 95. The internal computation is: 36 ≈ 30, 59 ≈ 50, 30+50 ≈ 80, 80+15 ≈ 92 (somewhere near 95). A bag of magnitude heuristics, not a digit-by-digit algorithm. Correct answer, wrong internal mechanism. + +The author concludes: **LLM behavior is impressive but the internal mechanism is bag-of-heuristics — FER.** + +### 2.18 The Platonic Representation Hypothesis (Huh et al. 2024) + +Empirical finding: as you scale up vision models (e.g., DINOv2) and language models (e.g., Llama), the **representation similarity** between them increases. Models trained on different data, different objectives, different architectures, are converging in some geometric sense to a shared representation of "reality." + +This is a **statistical phenomenon**. It does not necessarily mean the converged representation captures the regularities of the world (it might be a statistical optimum without being a meaningful decomposition). + +### 2.19 The author's position on the Platonic Hypothesis + +The Platonic Representation Hypothesis is real (empirically observed convergence) but **statistical**, not **structural**. Convergence in representation geometry does not imply convergence in factorization. + +A model could represent the world with one set of entangled features, and a different model could represent it with another set of entangled features, and both representations could be close in some metric space without either being factored. + +The author wants UFR: the converged representations should be **structurally aligned**, not just statistically correlated. + +### 2.20 Open-endedness as a learning paradigm + +A learning paradigm with four properties: +- **Complexification**: build regularities on top of regularities, bottom-up (like morphogenesis or development). +- **Emergence**: higher-level structure arises from lower-level dynamics. +- **Adaptability**: train for the ability to adapt to new environments, not for any specific task. +- **Serendipity**: expose the system to multiple environments in sequence, not a single fixed dataset. + +Open-ended search (e.g., Picbreeder-style evolution) exhibits all four. SGD does not. + +### 2.21 Pressure to adapt as the key driver + +Among the four properties, **adaptability** is the author's "hunch" about the most important. Reasons: +- Evolution optimizes for fitness in a changing environment; the implicit pressure to be adaptable (since the environment changes) creates robust representations. +- SGD optimizes for a fixed loss; no implicit pressure to be adaptable. +- A system trained to be adaptable will, almost as a side effect, capture regularities of the world (because regularities are what allow prediction across environments). + +The author conjectures: **strong representations and adaptable representations are one and the same.** + +### 2.22 Space of Forms ↔ Representation Space + +The author closes with a philosophical question: does the **internal representation of a good agent** correspond **directly** to the Platonic Space of Forms? + +The aspiration is that an ideal mind's representation would be a faithful mirror of the structure of reality. Current AI representations are imperfect instantiations of this ideal. + +This is speculation (the author acknowledges it) but provides the philosophical framing for "Platonic Intelligence" as the talk's title. + +--- + +## 3. Frame Analysis + +62 unique frames were extracted from the 89MB mp4 at threshold 0.05 (research talk with diagrams, equations, and motivational slides; OCR'd via winsdk in 3.7s). + +### 3.1 Slide 1 — Title slide (frame_00002) + +**OCR text:** +> Towards a Platonic Intelligence with Unified Factored Representations +> Akarsh Kumar +> MIT CSAIL +> November 4, 2025 + +The title and date. The slide also contains the central diagram: + +> **Unified Factored Representation** → Open-Ended Search → **Identical output behavior** → Solves Task → Function Space (Good Adaptability) +> **Fractured Entangled Representation** → Conventional SGD → Solves Task → Function Space (Poor Adaptability) + +This is the talk's structural thesis in diagram form. The left side is the training algorithm (open-ended vs SGD); the right side is the output behavior (both solve the task, function space is the same); the middle is the internal representation (UFR vs FER); and the bottom is the consequence (good vs poor adaptability). + +### 3.2 Slide 2 — The world is not random (frame_00001) + +**OCR text:** +> The World is not Random + +A one-line title slide introducing the premise: the world has structure. + +### 3.3 Slide 3 — The world has structure (frame_00004) + +**OCR text:** +> The World has Structure +> Real World + +A "Real World" header slide. Establishes the empirical fact that structure exists in the world (self-similarity, symmetries, persistent objects, common patterns). + +### 3.4 Slide 4 — Intelligent agents must capture this structure (frame_00016) + +**OCR text:** +> Intelligent Agents must capture this Structure + +The bridge claim (§2.3). Agents that understand the world must capture its structure in their internal representations. + +### 3.5 Slide 5 — Capturing structure with AI (frame_00017) + +**OCR text:** +> Capturing Structure with AI + +Introduces the problem statement: how do we get AI systems to capture world structure. + +### 3.6 Slide 6 — The lighting invariance problem (frame_00019) + +**OCR text:** +> Capturing Structure with AI +> • What about everything else? +> • Example: lighting invariance +> • How do you capture lighting invariance? +> Architecture → Lighting → Invariance + +The example: translation invariance is baked into CNNs, but lighting invariance is not. The diagram shows "Architecture → Lighting → Invariance" with arrows indicating the unknown relationship. + +### 3.7 Slide 7 — The fallback: SGD (frame_00020) + +**OCR text:** +> Capturing Structure with AI +> • What about everything else? +> • Example: lighting invariance +> • How do you capture lighting invariance? +> • We don't know +> • Solution: train on lots of data with SGD + +The current paradigm's answer to the unknown-regularity problem: SGD on lots of data. + +### 3.8 Slide 8 — Does this work? (frame_00021) + +**OCR text:** +> Does this Work? +> ChatGPT + +Yes, somewhat. ChatGPT exhibits real understanding of many aspects of the world. + +### 3.9 Slide 9 — The FER hypothesis (frame_00022, frame_00023) + +**OCR text (frame_00023):** +> Hypothesis: Fractured Entangled Representations +> • Conventional SGD training finds neural representations which are fractured and entangled +> Output behavior → Inputs → Solves Task → Function Space +> Fractured Entangled Representation → Conventional SGD +> (Poor Adaptability) + +The hypothesis stated formally: SGD finds FER. The diagram from the title slide reappears with the "Poor Adaptability" annotation. + +### 3.10 Slide 10 — The position: open-ended search (frame_00024, frame_00025) + +**OCR text (frame_00025):** +> Hypothesis: Fractured Entangled Representations +> • Conventional SGD training finds neural representations which are fractured and entangled +> • Doesn't capture the underlying regularities of the world +> • Position: Open-Ended Search may be the solution to learn unified and factored neural representations +> • Internal representation affects generalization, creativity, and continual learning + +The position paper statement: FER is the problem; open-ended search is the proposed solution. The three downstream consequences (generalization, creativity, continual learning) are explicitly listed. + +### 3.11 Slide 11 — CPPN introduction (frame_00026, frame_00027) + +**OCR text (frame_00027):** +> Compositional Pattern Producing Network (CPPN) +> • Toy domain to study neural representations: implicitly represent an image +> • Inspired by biological developmental process + +The CPPN as a toy representation. Note: biological inspiration (developmental process). + +### 3.12 Slide 12 — Picbreeder introduction (frame_00028, _00029, _00030, _00031, _00032, _00033, _00034) + +**OCR text (frame_00034):** +> Picbreeder! +> • Online website for humans to breed images to their desire +> • Evolve the underlying CPPNs +> • No end goal, do whatever you want! + +The Picbreeder setup: humans in the loop, no objective, "do whatever you want." The "no end goal" framing is critical — Picbreeder is the prototypical open-ended search. + +### 3.13 Slide 13 — What people found (frame_00035, _00036) + +**OCR text (frame_00035):** +> What Do You Expect to Find? + +**OCR text (frame_00036):** +> What People Actually Found! + +Two slides setting up the surprise: expectations vs reality of what humans evolved. + +### 3.14 Slide 14 — Why Greatness Cannot Be Planned (frame_00037) + +**OCR text (frame_00037):** +> Why Greatness Cannot be Planned +> • Many insights on the nature of search +> • Deception +> • Serendipity +> • Open-Endedness +> • Case studies: +> • Natural Evolution +> • Scientific Innovation + +Reference to Stanley & Lehman's book (§2.12). The three concepts (deception, serendipity, open-endedness) and two case studies (natural evolution, scientific innovation). + +### 3.15 Slide 15 — Picbreeder book cover (frame_00038) + +**OCR text:** +> Picbreeder has Intriguing Properties +> Open-Ended + +The book cover image of *Why Greatness Cannot Be Planned*. Stanley, Lehman, Springer. + +### 3.16 Slide 16 — Serendipitous exaptation (frame_00040, _00041) + +**OCR text (frame_00041):** +> Picbreeder has Intriguing Properties +> Serendipitous Exaptation +> Stepping stone to the Teapot +> Stepping stone to the Skull +> Stepping stone to Jupiter +> Stepping stone to the Butterfly +> Stepping stone to the Penguin +> Stepping stone to the Lamp + +A network visualization showing how a single ancestor CPPN connects to many target images via evolutionary stepping stones. The teapot, skull, Jupiter, butterfly, penguin, and lamp all share intermediate CPPNs. + +### 3.17 Slide 17 — Emergence of evolvability (frame_00042) + +**OCR text:** +> Emergence of Evolvability +> • Natural evolution has developed adaptable genotypes: +> • Canalization +> • Regularity +> • Modularity +> • Symmetry +> • Certain axes of variation become more likely while others become impossible + +The four properties of evolvable genotypes. The closing statement: certain axes of variation become more likely (canalization buffers them) while others become impossible (eliminated by selection). + +### 3.18 Slide 18 — Modularity in Picbreeder (frame_00043) + +**OCR text:** +> Modularity +> Most Picbreeder images feature forms of canalization +> Mutating single genes of these images holistically affects distinct aspects of the image +> mutation in [gene A] → affects only [Object genes] +> mutation in [gene B] → affects only [Shadow genes] +> [image showing modular separation] +> +> Some Picbreeder images show poor canalization +> Mutating single genes of these images affects none or many parts of the image +> → These images have little structural organization in their genome +> [image showing entangled effects] + +The empirical case for modularity. Some Picbreeder CPPNs have **modular** gene-to-trait mapping (mutation in gene A only affects aspect X). Others have **entangled** mapping. The evolved modular ones are more evolvable. + +### 3.19 Slide 19 — SGD skull vs Picbreeder skull (frame_00045, _00046, _00047) + +**OCR text (frame_00047):** +> Learning the Picbreeder Skull with SGD +> • Let's train a conventional network to recreate the skull +> • Perfect reconstruction! +> [Picbreeder Skull image] → [SGD Training] → [SGD Skull image] + +The experimental setup: take a Picbreeder skull, train an MLP via SGD to reproduce it. SGD succeeds. + +### 3.20 Slide 20 — Layerization (frame_00048, _00049, _00050) + +**OCR text (frame_00049):** +> Layerization +> • Convert everything to a universal architecture space: MLP + +**OCR text (frame_00050):** +> Layerization +> • Convert everything to a universal architecture space: MLP +> • Existence proof of Picbreeder solution MLP weight space +> [Diagram: CPPN → MLP] + +The key insight: a CPPN can be layerized (converted to an MLP). So an MLP with the right weights **is** the Picbreeder solution. The question is whether SGD finds those weights. + +### 3.21 Slide 21 — Picbreeder skull weights (frame_00051, _00053) + +**OCR text (frame_00053):** +> Picbreeder Skull +> Unified Factored Representation +> • Controls Mouth Opening +> • Controls Eye Winking +> • Controls Eye Width +> • Controls Jaw Width +> [Sweeping Weight Value] + +When the Picbreeder CPPN is layerized and weights are swept, each weight corresponds to a **distinct semantic axis**: mouth opening, eye winking, eye width, jaw width. This is UFR. + +### 3.22 Slide 22 — SGD skull weights (frame_00052, _00053) + +**OCR text (frame_00052):** +> SGD Skull +> Fractured Entangled Representation +> [Sweeping Weight Value] + +**OCR text (frame_00053):** +> SGD Skull +> Fractured Entangled Representation +> [Sweeping Weight Value] + +The same MLP architecture, but with SGD-found weights. Sweeping any weight produces chaotic output changes that don't correspond to any single semantic axis. This is FER. + +The contrast is stark: **same loss, same architecture, completely different internal organization.** + +### 3.23 Slide 23 — Picbreeder Butterfly vs SGD Butterfly (frame_00054, _00055, _00056) + +**OCR text (frame_00056):** +> Picbreeder Butterfly +> Unified Factored Representation +> • Controls Wing Area +> • Controls Color +> • Converts Butterfly to Fly +> • Controls Vertical Shape +> [Sweeping Weight Value] +> +> SGD Butterfly +> Fractured Entangled Representation +> [Sweeping Weight Value] + +Same experiment, butterfly image. Picbreeder factorization: wing area, color, butterfly-to-fly transition, vertical shape. SGD entanglement: chaotic changes. + +### 3.24 Slide 24 — Picbreeder Apple vs SGD Apple (frame_00057, _00058, _00059) + +**OCR text (frame_00059):** +> Picbreeder Apple +> Unified Factored Representation +> • Controls Stem Angle +> • Controls Apple Size +> • Cleans Background +> • Removes Stem +> [Sweeping Weight Value] +> +> SGD Apple +> Fractured Entangled Representation +> • Controls Apple Size +> • Cleans Background +> [Sweeping Weight Value] + +Third replication with apple. Same pattern. + +### 3.25 Slide 25 — How does this apply to LLMs? (frame_00060) + +**OCR text:** +> How does this Apply to LLMs? + +Transition slide to the LLM application. + +### 3.26 Slide 26 — FER in LLMs: GPT-3 (frame_00061) + +**OCR text:** +> FER In LLMs +> Evidence in GPT-3 +> Example 1: +> Me: I have 3 pencils, 2 pens, and 4 erasers. How many things do I have? +> GPT-3: You have 9 things. [always correct] +> Example 2: +> Me: I have 3 chickens, 2 ducks, and 4 geese. How many things do I have? +> GPT-3: You have 10 animals total. [always incorrect] + +The GPT-3 GSM8K example. The first query (pencils/pens/erasers) gives the correct answer of 9 (3+2+4). The second query (chickens/ducks/geese) gives 10 (counting the animals, not the items). GPT-3 has memorized the surface pattern of the first example but does not understand counting. + +### 3.27 Slide 27 — GSM-Symbolic paper (frame_00062) + +**OCR text:** +> GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models +> Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Samy Bengio, Mehrdad Farajtabar, Oncel Tuzel +> Apple +> Abstract +> Recent advancements in Large Language Models (LLMs) have sparked interest in their formal reasoning capabilities, particularly in mathematics. The GSM8K benchmark is widely used to assess the mathematical reasoning of models on grade-school-level questions. While the performance of LLMs on GSM8K has significantly improved in recent years, it remains unclear whether their mathematical reasoning capabilities have genuinely advanced, raising questions about the reliability of the reported metrics. To address these concerns, we conduct a large-scale study on several state-of-the-art open and closed models. To overcome the limitations of existing evaluations, we introduce GSM-Symbolic, an improved benchmark created from symbolic templates that allow for the generation of a diverse set of questions. GSM-Symbolic enables more controllable evaluations, providing key insights and more reliable metrics for measuring the reasoning capabilities of models. Our findings reveal that LLMs exhibit noticeable variance when responding to different instantiations of the same question. Specifically, the performance of all models declines when only the numerical values in the question are altered in the GSM-Symbolic benchmark. Furthermore, we investigate the fragility of mathematical reasoning in these models and demonstrate that their performance significantly deteriorates as the number of clauses in a question increases. We hypothesize that this decline is due to the fact that current LLMs are not capable of genuine logical reasoning; instead, they attempt to replicate the reasoning steps observed in their training data. When we add a single clause that appears relevant to the question, we observe significant performance drops (up to 65%) across all state-of-the-art models, even though the added clause does not contribute to the reasoning chain needed to reach the final answer. + +The full abstract of the Apple paper. Key claim: **LLMs attempt to replicate reasoning steps from training data** rather than performing genuine reasoning. Performance drops up to 65% with added irrelevant clauses. + +### 3.28 Slide 28 — Reasoning or Reciting paper (frame_00063) + +**OCR text:** +> Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks +> Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, Yoon Kim +> MIT, Boston University +> Performance +> Default vs Counterfactual (orange) +> Tasks: Spatial, Code Exec., Chord Fingering, Code Gen., Basic Syntax, SET Game +> GPT-4 consistently and substantially underperforms on counterfactual variants compared to default task instantiations. + +The counterfactual task results. Across 6 task types, GPT-4's counterfactual performance is consistently lower than default. The implication: GPT-4 has memorized task structures, not learned the underlying task logic. + +### 3.29 Slide 29 — Anthropic circuit tracing (frame_00068) + +**OCR text:** +> On the Biology of a Large Language Model +> We investigate the internal mechanisms used by Claude 3.5 Haiku — Anthropic's lightweight production model — in a variety of contexts, using our circuit tracing methodology. +> [Diagram: Inputs near 30 make this early feature fire] → Add features → Sum features → Output +> Example: low precision features +> _6, 36, 36+59=, sum = _92, -40 + -50, 5-, -59, sum = _9, 36 + -60, 59, 59, Exam mod O features, _9, Sum Features +> The model has finally computed information about the sum: its value mod 10, mod 100, and its approximate magnitude. +> Lookup Table Features, Add Function Features, Input Features +> Most computation takes place on the "zu" token + +The Anthropic 2025 paper on Claude 3.5 Haiku's arithmetic. The diagram shows the **circuit** for 36 + 59 = 95. The internal mechanism is **magnitude-based**: features activate around "30" and "50," outputs are around "92" — the model uses approximate magnitude heuristics rather than digit-by-digit addition. + +The "low precision features" label is critical: the model computes approximate magnitudes, not exact digits. This is bag-of-heuristics, not algorithm. + +### 3.30 Slide 30 — Scaling laws (frame_00069) + +**OCR text:** +> Scaling helps... but in what way? +> Scaling Laws for Neural Language Models +> 2024 +> L = (C_min / 2.3) ^ (-0.05) +> Compute (PF-days, non-embedding): 4.2, 3.9, 3.6, 3.3, 3.0, 2.7 +> Dataset Size (tokens): 10^7, 10^9 +> L = (C_min / 5.4) ^ (-0.05) +> Kaplan et al. (2020) +> Parameters (non-embedding): 10^3, 10^7, 10^9 + +The classic Kaplan et al. 2020 scaling laws: loss decreases as a power law in compute, dataset size, and parameters. + +### 3.31 Slide 31 — Platonic Representation Hypothesis (frame_00070) + +**OCR text:** +> The Platonic Representation Hypothesis +> Neural networks, trained with different objectives on different data and modalities, are converging to a shared statistical model of reality in their representation spaces. +> [Diagram: z, x, Img, "A red sphere next to..." ext] +> Huh et al. (2020) + +The hypothesis stated: **different modalities converge to a shared statistical model.** The diagram shows vision (image of a red sphere) and language (text describing the red sphere) producing aligned representations. + +### 3.32 Slide 32 — What could be better? (frame_00072, _00073) + +**OCR text (frame_00073):** +> What could be better? +> • Complexification (ex: morphogenesis, etc.) +> • Builds regularities on top of other regularities (bottom up) +> • Emergence +> • Adaptability +> • Pressures the learned regularities to be robust to environmental changes +> • Representation must capture axes of variation which "carve nature at its joints" +> • Serendipity (order matters for learning!) +> • Much higher chance of finding a useful learning curriculum +> • What learning paradigm captures all of these? Open-Endedness! + +The four properties of open-endedness: complexification, emergence, adaptability, serendipity. The closing: **open-endedness** is the paradigm that captures all four. + +### 3.33 Slide 33 — Is this a Platonic Intelligence? (frame_00074, _00075) + +**OCR text (frame_00074):** +> Is this a Platonic Intelligence? +> Space of Forms +> Real World +> Intelligent Agents +> Unified Factored Representation + +**OCR text (frame_00075):** +> Is this a Platonic Intelligence? +> Aspirational Ideal +> Unified Factored Representation +> Instantiation +> Fractured Entangled Representation + +The philosophical framing: **Aspirational Ideal = Unified Factored Representation**; **Instantiation (current AI) = Fractured Entangled Representation**. The Space of Forms (Plato) is what an ideal intelligence would directly represent. Current AI is a brittle instantiation of this ideal. + +### 3.34 Slide 34 — Collaborators (frame_00076) + +**OCR text:** +> Collaborators +> Jeff Clune (UBC, Vector Institute) +> Joel Lehman (University of Oxford) +> Kenneth Stanley (Lila Sciences, UBC) + +The collaborators. This positions the talk within the broader neuroevolution / open-endedness research community led by Clune, Lehman, and Stanley. + +### 3.35 Slide 35 — Thank you (frame_00077, _00081) + +The closing slide. + +--- + +## 4. Transcript Highlights + +Sixteen verbatim passages from the cleaned transcript (1659 segments, 61KB) that capture the conceptual flow. + +### 4.1 Motivation: world has structure (T+0:30) + +> "So uh to start off with uh I want to state something obvious that you guys already know that the world is not random but rather it has a lot of structure as we know and this includes everything from like self similarity across like many spatial scales to like physics symmetries like symmetries in like translation rotation invariance of the world [...] And um the the fact that there's so many um common patterns across in this world, right, across many different objects, that's what leads some of us to believe in this like idea of like there's like this this um space of forms, this platonic space of forms where these properties, common properties across many objects are inherited from the space of forms." + +The opening premise (§2.1) and the philosophical connection to Plato's forms (§2.2). + +### 4.2 The bridge claim (T+1:30) + +> "So um I claim that basically intelligent agents in order to solve their goals they need to really understand how the world works in order to control it right to uh achieve their goals. And in order to un uh understand this world, I argue that they must capture this structure of the world. All these different structures that the world has, they must capture it in some way. And more specifically, what I mean is that the internal representations of their minds and their brains must capture the structure of the world." + +The central thesis (§2.3) — internal representations must capture world structure. + +### 4.3 Architectural inductive biases (T+2:30) + +> "One is like translation invariance in images whenever you look at an image and you see like some object you know that the system should process the object very similarly if it's over here versus as if it's like shifted over 100 pixels to the right. And this is translation equivariance or translation invariance. And we try to bake this into um ar an architecture based on the convolutional architecture. Another example is like whenever we process sets of objects and we want we don't care about the ordering of the set [...] then we use an attention architecture uh in a transformer to do this kind of thing." + +Examples of baked-in symmetries via architecture (§2.4). CNN for translation, transformer for permutation. + +### 4.4 The lighting invariance problem (T+3:30) + +> "But what about all the other um structures in our world, right? All the other regularities. And one example I really like is like lighting invariance. So you see this lion in the dark and during the day and you don't really know what architecture should capture this lighting invariance, right? And we don't know how to do this." + +The unknown-regularity problem. CNN handles translation; what handles lighting? + +### 4.5 SGD as fallback (T+4:00) + +> "Well, okay. So, the solution is we just try to train on a lot of data with SGD and hope that the AI will learn this underlying regularity of the world um based on the patterns that it picks up from all of these data uh from all this data right that's this is the predominant paradigm in modern deep learning currently." + +The current paradigm (§2.5). Train on lots of data, hope SGD finds the regularity. + +### 4.6 Does SGD work? (T+4:30) + +> "So, the question is does this actually work right and there's a lot of evidence that's showing that it is somewhat working. I mean the AI systems nowadays if you use chat gpt they can pick up on all sorts of patterns of the world and they seem to really understand the world and how it works and even self-driving systems they can like figure out uh what a stop sign looks like during the day at night and it seems like everything is just working but this brings us to our um position paper and our um which where we hypothesize that conventional SGD in deep learning um finds neural representations which are actually fractured and entangled." + +The partial success + FER hypothesis (§2.6, §2.7). + +### 4.7 The FER hypothesis (T+5:00) + +> "So uh over here we can see that I mean you find a network which has a certain output behavior but the output behavior we visualize it as like a skull which I'm going to get into detail how we're doing this but at a high level you visualize the output behavior and the internal representations don't really match what you would expect to see if it's really like understanding the skull in a subjective way right and more specifically what I mean is that it doesn't capture the underlying regularities of this world which is the skull." + +FER explanation (§2.7). The output behavior matches the target, but the internal representation does not. + +### 4.8 Open-ended search as alternative (T+6:00) + +> "So our position or our opinion is that a different kind of search algorithm which is not conventional SGD but a different kind of very exotic open-ended search may be the solution to learn unified and factored neural representations." + +The position (§2.9). Open-ended search → UFR. + +### 4.9 Picbreeder (T+8:00) + +> "This was actually this thing called Picbreeder which was an online website where you could evolve the underlying CPPNs or compositional pattern producing networks by um by humans selecting images they liked. And so what would happen is the image would change slightly and you could click whichever one you liked and it would go off in that direction. And they didn't tell people what to do they just said do whatever you want right." + +Picbreeder as the canonical example (§2.11). Humans in the loop, no objective. + +### 4.10 Why Greatness Cannot Be Planned (T+11:00) + +> "There's a really cool book called Why Greatness Cannot Be Planned by Kenneth Stanley and Joel Lehman who were the creators of Picbreeder. And they talked about the reasons why open-endedness can lead to um so many discoveries. They talked about three main concepts which is deception serendipity and open-endedness." + +Reference to Stanley & Lehman's book (§2.12). + +### 4.11 Picbreeder skull: SGD vs Picbreeder (T+13:00) + +> "So they they took the picbreeder skull image and they trained like a standard SGD MLP to recreate that skull and it got like a perfect reconstruction right but the SGD MLP is fractured and entangled. Now they took the Picbreeder CPPN and they layerized it which is basically converting it to an MLP. And what they found is that these Picbreeder weights when you sweep them are interpretable. So like the first weight might control mouth opening the second weight might control like eye width and so on." + +The experimental result (§2.16). Same loss, same architecture, different internal organizations. + +### 4.12 FER in LLMs: GPT-3 chicken/duck (T+18:00) + +> "In the first example they ask how many things do I have and it's like 3 pencils 2 pens and 4 erasers and GPT-3 says you have 9 things which is always correct. But then they ask you have 3 chickens 2 ducks and 4 geese how many things do I have. And GPT-3 says you have 10 animals total. So it counted the animals as opposed to the items. So it's like it learned the surface pattern of the first example but not actually counting." + +GPT-3 counting failure (§2.17). + +### 4.13 FER in LLMs: counterfactual tasks (T+19:30) + +> "Then they did this is like the reasoning reciting paper from MIT Boston University where they change the task to be like in a counterfactual world. So instead of doing arithmetic in base 10 you do it in base 9 this is a way less common um way to do things right or in code execution you change it to be base one indexing then the performance goes down a lot. So in this way you could argue that there's like it's not respecting if you change one thing about the world it's not respecting it's not understanding the world where it's like um it understands the regularities in a deep way where it can be robust to these counterfactual changes right." + +The counterfactual task degradation (§2.17). + +### 4.14 FER in LLMs: Claude 3.5 Haiku arithmetic (T+20:30) + +> "Here's another example I really like um this is from an anthropic mechanism paper from quite recently where they looked at Claude and they asked it what's 36 + 59 and it came out with the answer 95 which is correct. But if you look at the actual arithmetic that goes behind this answer right the actual neural circuitry that's happening you can see it's using like random heuristics that a human would never even think about. It's like saying like oh 36 is around 30 and if you add and it's around 40 and if you add around 40 plus around 50 it's like around 92 and it's it's completely like different than how a human would do this arithmetic problem." + +Claude 3.5 Haiku arithmetic as bag-of-heuristics (§2.17). + +### 4.15 Scaling laws don't address representation quality (T+22:00) + +> "This is all saying this is all very very cool but I do want to point out that these scaling laws and this platonic uh representation hypothesis are all very like statistical observations of what's happening right these are statistical observations of a statistical uh behaviors right and it's very unclear how this relates to respecting regularities of the world and I think we need a lot more um um research trying to connect the two because the examples we gave in our paper where we did a survey on like all these um studies that were showing that language models struggle in this task in this way or something they're more talking about whether or not you respect some regularity of the world and I think this is a very different um way to see the world than if you just look at this as like a statistical um mechanism right." + +The author's critique of scaling laws and the Platonic Representation Hypothesis as statistical observations that don't address regularity (§2.19). + +### 4.16 Open-endedness properties (T+24:00) + +> "And I think the things that would matter the most is like something like some process of complexification. If you look at like the process like morphogenesis, it doesn't just encode every part of your body at once. It grows it according to something more fundamental [...] Another thing I think could be really important is training for adaptability. So rather than just training for to solve the task, what we really want is to train for adaptability because I think it gives you a lot of like regularization pressure to learn like symmetries and regularities to be robust to changes." + +The four properties of open-endedness (§2.20) and the emphasis on adaptability (§2.21). + +### 4.17 Q&A: regularities vs simplicity (T+33:00) + +> "One of the cool ideas that I'm investigating with the undergrad here is like basically trying to see if um the network if we try to make the weights predictable because predictability is like a very specific type of signal that could qualitatively change what you learn, right? Because weight decay is more like a simplicity bias where you want to regularize it towards like low weight norm. But I think that if you try to make it predictable, that's not the same thing. It's not the same as simplicity. It's completely different. So I think there are ideas in which it could change, but honestly I don't even think that's that would work that well. I think what you really need to do is like rethink what's going on and what the pressures you are that you're optimizing towards. And I think the pressure to adapt is probably the biggest thing that needs changing." + +The author's response to a question about whether weight decay or other regularization tricks could help. The answer: **pressure to adapt** is more fundamental than simplicity bias. + +### 4.18 Q&A: fairness of the CPPN vs MLP comparison (T+35:00) + +> "In terms of the fairness I mean uh in this paper we're it's like a position paper and we're not proposing picbreeder as an algorithm that's supposed to compete against and that everyone should use right it's more like to inspire like that fact that this algorithm has some cool properties that were assoc even though we cheated and we used humans in the loop we used a different um neat network it's supposed to inspire ideas that maybe we can extract some insights and turn this into an algorithm which can compete against SGD and do better." + +The author's honest framing of the work: this is a **position paper** meant to inspire, not an algorithm proposal. + +--- + +## 5. Mathematical / Theoretical Content + +This section develops the conceptual content of the talk in depth. The talk is conceptual rather than heavily mathematical, but several key concepts admit formalization. + +### 5.1 Inductive biases via group invariance + +A function f: X → Y is **equivariant** under a group G if f(g·x) = g·f(x) for all g ∈ G. A function is **invariant** if f(g·x) = f(x). + +Architectural inductive biases encode specific group invariances: +- **Translation invariance**: G = ℝ² (continuous translations); achieved by convolutions. +- **Permutation invariance**: G = S_n (symmetric group); achieved by attention/DeepSets. +- **Rotation invariance**: G = SO(2) or SO(3); achieved by group convolutions. + +The geometric deep learning program (Bronstein et al. 2021) generalizes this: any prior knowledge of the form "f is invariant under some group action on inputs" can be encoded as an architectural constraint. + +**Limitation:** for invariances we don't know in advance (lighting invariance, articulation, etc.), there is no obvious architectural choice. + +### 5.2 The SGD loss landscape + +The training loss L(θ) = E_{(x,y)~D}[ℓ(f_θ(x), y)] is a function from the parameter space Θ to ℝ₊. SGD samples a stochastic gradient estimate g_t = ∇L_batch(θ_t) and updates θ_{t+1} = θ_t − η_t g_t. + +The landscape of L has: +- Local minima (often degenerate — many minima with similar loss values). +- Saddle points (more common in high dimensions than local minima). +- Plateaus (regions of near-zero gradient). + +The network's final weights θ* are a **local optimum** of L, not necessarily the global optimum, and not necessarily a "good" internal representation even if loss is low. + +### 5.3 Why SGD finds FER (informal argument) + +Consider two weight configurations θ₁ and θ₂ that achieve similar loss L. They correspond to two different ways of implementing the input-output mapping. The "natural" structure of the data (UFR) is one possibility; many other configurations exist that achieve the loss without the natural structure. + +SGD samples these configurations **uniformly with respect to their basin of attraction** (in practice, with bias from initialization and noise). Most configurations in parameter space are not UFR — UFR is a low-dimensional subset (the "manifold of factored representations"). SGD finds FER with overwhelming probability because FER is the generic case. + +**Counterargument:** SGD with strong inductive biases (architectural symmetries) constrains the search to a smaller subset, increasing the chance of finding structured representations. This is what convolutional and transformer architectures do for spatial and permutation structure. + +### 5.4 The CPPN parameterization + +A CPPN (Compositional Pattern Producing Network) is a neural network f: ℝ² → ℝ³ (image coordinates to RGB values) with **heterogeneous activation functions** at each node: + +f(x, y) = σ_L(W_L σ_{L-1}(W_{L-1} ... σ_1(W_1 [x, y] + b_1) ... + b_L) + +where σ_i ∈ {sin, cos, gaussian, sigmoid, abs, linear, ...}. The choice of activation function at each node is part of the architecture (in neuroevolution, it can also be evolved). + +The heterogeneous activations encode **regularity biases**: +- sin/cos → periodic patterns. +- gaussian → radial symmetries. +- abs → symmetry across axes. +- sigmoid → soft thresholds. + +A random CPPN produces an image with structure (because of the regularity biases). A **trained** CPPN (via SGD) tends to find entangled weight configurations; an **evolved** CPPN (via selection) tends to find factored configurations. + +### 5.5 Picbreeder's evolution algorithm + +Picbreeder is a steady-state genetic algorithm (not generational): +1. Maintain a population of CPPN genomes. +2. At each step, present the user with several images. +3. User selects the image they like most. +4. Selected CPPN is mutated (and optionally recombined) to produce offspring. +5. Offspring replace less-fit members. + +Selection pressure: human aesthetic preference (with no explicit objective). + +Effective population size: tens to hundreds. Generations: hundreds to thousands. Effective compute: days to weeks of human attention per "interesting" image. + +### 5.6 Waddington's canalization + +Conrad Waddington (1942, 1957) introduced **canalization** to describe how embryonic development buffers against genetic variation. A canalized trait is one where many genotypes map to the same phenotype — the developmental "landscape" has ridges and valleys that channel variation into a small set of outcomes. + +In ML terms, canalization is **insensitivity to certain weight perturbations**. An MLP is canalized with respect to weight w if small perturbations to w leave the output approximately unchanged. The Picbreeder CPPNs are canalized with respect to mutations in "non-active" genes; SGD-trained MLPs are not (every weight matters). + +### 5.7 The Picbreeder Skull MLP factorization + +The experimental result (Secretan et al. 2008; later formalized in the Kumar/Stanley papers): + +Given a Picbreeder skull image (a CPPN output), the layerized Picbreeder CPPN is an MLP whose weights have the following factorization property: + +For each weight w_i in the MLP, there is a **single semantic axis** of variation in the output (e.g., mouth opening, eye width). Sweeping w_i changes only that axis; other axes are invariant. + +In formal terms: there exists a set of functions {a_i: Image → ℝ} such that a_i(Image(w + δ e_i)) ≈ a_i(Image(w)) + c_i δ for some constant c_i, and a_j(Image(w + δ e_i)) ≈ a_j(Image(w)) for j ≠ i. + +This is **factored**: each weight controls one axis independently. + +### 5.8 SGD skull: no factorization + +The SGD-trained MLP with the same loss has weights where sweeping any weight changes **multiple** semantic axes simultaneously: + +For all i, there exist j ≠ i such that |a_j(Image(w + δ e_i)) − a_j(Image(w))| > 0. + +The SGD MLP has the same input-output mapping (it reproduces the skull image) but its internal organization is **entangled**. + +### 5.9 Why does this happen? + +Two hypotheses (the author leans toward the second): + +**Hypothesis A: Basin of attraction.** SGD finds FER because the FER basin of attraction is larger than the UFR basin in the SGD loss landscape. This is a structural claim about the landscape. + +**Hypothesis B: Search algorithm.** SGD is a **gradient-following** algorithm; it finds local minima of the loss. The UFR solution corresponds to a specific local minimum that SGD cannot reach because the gradient does not point toward it. Open-ended search explores more of the space (via mutation) and finds the UFR minimum because it does not follow gradients. + +The evidence: even with weight decay and other regularizers, SGD almost never finds UFR (per the Q&A in §4.17). This favors Hypothesis B — the issue is the search algorithm, not the loss landscape. + +### 5.10 FER in LLMs: arithmetic as bag-of-heuristics + +The Anthropic 2025 paper traces Claude 3.5 Haiku's computation for 36 + 59 = 95. The mechanism (per the circuit tracing): + +1. Input features activate for the digits 3, 6, 5, 9 (specific to ones place) and for magnitudes "around 30," "around 50" (rough). +2. Add function features combine the two addends: separate computation of ones-digit and approximate magnitude. +3. Sum features output the ones-digit (5) and approximate magnitude (~92). +4. Lookup table features handle specific pairs (e.g., 36+59 → 95 as a stored association, when seen). + +This is **not** digit-by-digit addition. The model uses magnitude heuristics (30 + 50 ≈ 80) and lookup tables (specific pairs) rather than a general algorithm. + +In formal terms: the model has learned a **piecewise function** {lookup_table, magnitude_heuristic, digit_combination} rather than a **uniform algorithm**. The piecewise structure is FER — the representations for different problems are different, not unified. + +### 5.11 The Platonic Representation Hypothesis (formal) + +Given two models A (trained on modality X) and B (trained on modality Y), define a **representation similarity** measure s(A, B) ∈ [0, 1] (e.g., RSA: Representational Similarity Analysis, or mutual kNN). + +The PRH (Huh et al. 2024) states: as model size and performance increase, s(A, B) → 1. + +This is an empirical claim about **representational geometry** — the manifolds in representation space are converging. It does not say anything about whether the representations are factored. + +### 5.12 Distinguishing FER from UFR via probing + +How to test whether a representation is FER or UFR? + +**Test 1: Single-weight ablation.** Perturb one weight at a time. If each perturbation affects only one semantic axis (measured by independent interpretable probes), the representation is UFR. If each perturbation affects multiple axes, it's FER. + +**Test 2: Representation rotation.** Apply a fixed rotation R to the weight space: w' = R w. If the rotated representation still produces the same semantic outputs (with rotated axes), the representation is UFR. If the rotated representation produces incoherent outputs, it's FER. + +**Test 3: Causal intervention.** Intervene on weight w_i in a direction that should change semantic axis a_i. If only a_i changes in the output, UFR. If multiple axes change, FER. + +These tests are computable and have been applied to small models. Applying them to LLMs is open work. + +--- + +## 6. Connections + +This section maps the talk's content to the broader 12-video research campaign. + +### 6.1 Backward (cluster A foundations) + +#### 6.1.1 `cs229_building_llms_20260621` + +The CS229 lecture on building LLMs covers the SGD-based training paradigm that Kumar argues produces FER. The CS229 lecture presents EBMs as an alternative generative paradigm; Kumar's UFR is a different alternative (open-ended search). + +**Connection depth:** Foundational disagreement. Both work in deep learning; Kumar argues CS229-style training is fundamentally flawed at the representation level. + +#### 6.1.2 `probability_logic_20260621` + +The probability logic talk covers Kolmogorov's extension theorem and the foundations of stochastic processes. The CPPN framework can be formalized as a stochastic process (the spatial statistics of CPPN-generated images are well-defined). The "underlying regularities of the world" Kumar invokes are the **statistical regularities** that probability theory is designed to capture. + +**Connection depth:** Foundational. The notion of "regularity" presupposes a probability distribution. + +#### 6.1.3 `entropy_epiplexity_20260621` + +The epiplexity talk covers Kolmogorov complexity and information measures. The Picbreeder UFR has lower **algorithmic complexity** (Kolmogorov complexity) than the SGD FER: a UFR is a short program (weights map to semantic axes 1-to-1), while FER is a longer program (weights map chaotically to outputs). + +**Connection depth:** Conceptual bridge. The UFR is the "low complexity" representation; the FER is the "high complexity" representation. Epiplexity (observer-relative complexity) might explain why SGD finds FER: the loss function is a low-epiplexity observer, so it accepts any low-loss solution regardless of epiplexity of the weights. + +#### 6.1.4 `score_dynamics_giorgini_20260621` + +The Giorgini talk on score-based generative modeling presents an alternative route to capturing regularities. Instead of open-ended search, Giorgini uses the **stationary score** s(x) = ∇ log p_ss(x) to express regularities as gradient fields. The score is the **gradient of the log-density** — it encodes the local geometry of the data distribution. + +**Connection:** Both talks aim to capture "regularities of the world" — Kumar via UFR + open-ended search; Giorgini via score matching. The connection: an ideal UFR for a generative model would have weights that each encode a **direction in score space**. The factorization of the representation corresponds to the factorization of the score into independent directions. + +**Connection depth:** Speculative but deep. Pass 2 could explore whether the Picbreeder factorization corresponds to a score factorization. + +### 6.2 Forward (cluster E applications) + +#### 6.2.1 `cs336_architectures_20260621` + +The CS336 lecture on language model architectures covers diffusion LMs and modern LLM training. The FER diagnosis predicts that scaling up LLMs (per CS336's curriculum) will not solve the brittleness — because the representations remain FER regardless of scale. + +**Connection depth:** Predictive critique. The author would argue that CS336-style scaling will improve behavioral metrics (loss, benchmark accuracy) without resolving FER — the same brittle mechanisms, more refined. + +#### 6.2.2 `creikey_dl_cv_20260621` + +The Creikey DL/CV lecture covers diffusion models for images (DDPM). The Picbreeder setting is conceptually similar — generating images from learned representations. DDPM with score matching is the modern statistical approach; Picbreeder with neuroevolution is the open-ended approach. + +**Connection depth:** Methodological contrast. DDPM uses SGD with score-matching loss; Picbreeder uses neuroevolution. Both aim to capture image distributions; they find different internal organizations. + +### 6.3 Lateral (other cluster B / C videos) + +#### 6.3.1 `free_lunches_levin_20260621` + +Levin's "free lunches" talk likely covers algorithmic information theory and the relationship between search algorithms and problem structure. The "free lunch" in Kumar's framing: **regularity in the world** means that some search algorithms (open-ended) get "free" structure capture that others (SGD) miss. There is no free lunch for SGD on the loss landscape; but there is a free lunch for open-ended search on the **representation** landscape. + +**Connection depth:** Conceptual extension. Both talks concern the relationship between search algorithms and problem structure. + +#### 6.3.2 `generic_systems_fields_20260621` + +Fields' "generic systems" talk (cluster C) likely covers general systems theory and the principles of complex systems. The four properties of open-endedness (complexification, emergence, adaptability, serendipity) are all properties of complex systems. The author's framing of intelligence as capturing world regularities aligns with the general-systems perspective. + +**Connection depth:** Conceptual. The complex-systems perspective provides the theoretical backing for open-endedness. + +#### 6.3.3 `brain_counterintuitive_20260621`, `neural_dynamics_miller_20260621`, `multiscale_hoffman_20260621` + +These cluster C talks cover brain-related topics. The brain is the canonical example of a system that captures world regularities — biological intelligence is the existence proof that UFR-like representations are achievable. The Picbreeder analogy (biological development → CPPN evolution) is explicit in Kumar's talk. + +**Connection depth:** Inspirational. The brain as evidence that UFR is achievable. + +### 6.4 Cross-cutting themes + +Four themes recur across the campaign and connect to Kumar's talk: + +1. **Representation quality vs behavioral metrics** (this talk + cs229 + cs336): scaling improves metrics but may not improve representations. +2. **Algorithmic vs neural representations** (entropy_epiplexity + this talk): Kolmogorov complexity and weight structure are related. +3. **Information geometry** (score_dynamics + this talk): both talk about "the right geometry" — score manifolds for Giorgini; factored manifolds for Kumar. +4. **Open-endedness as a meta-principle** (this talk + free_lunches_levin + generic_systems_fields): the principle that search without a fixed objective outperforms search with one. + +--- + +## 7. Open Questions + +Sixteen questions arising from this talk that Pass 2 (de-obfuscation via user's mathematical encoding) should address. + +### 7.1 Theoretical + +1. **Formal definition of UFR.** What is the mathematical definition of "unified factored representation"? Is there a unique factorization for a given input-output mapping, or are there many? How do we measure "factoredness"? + +2. **Existence of UFR for general functions.** The Picbreeder CPPNs admit UFR (per the layerization). Do all computable functions admit UFR? What is the class of functions with UFR? + +3. **Relation to algorithmic complexity.** Is the UFR the **minimum-description-length** representation of the function? Are UFRs always Kolmogorov-minimal, or only sometimes? + +4. **Information-theoretic characterization.** The score function (Giorgini) is the gradient of the log-density. Is there an analogous "factored gradient" characterization? Does UFR correspond to a specific factorization of the score? + +5. **Basin of attraction in loss landscape.** The author suggests SGD cannot reach the UFR basin. Is this true? Are there SGD variants (with strong regularization) that find UFR? The Q&A (§4.17) suggests no — but a rigorous proof is missing. + +### 7.2 Empirical + +6. **Generalization of UFR to LLMs.** The Picbreeder evidence is on CPPN-generated images. Do LLM-scale models (billions of parameters) admit UFR? If we layerize an LLM, do we get interpretable weights? Anthropic's circuit tracing suggests no — LLM computations are entangled even at the "circuit" level. + +7. **Direct UFR training algorithm.** The author says the Picbreeder algorithm is "cheating" (humans in the loop). What is a practical algorithm that finds UFR without humans? Quality-Diversity algorithms (MAP-Elites, Novelty Search) are candidates; have they been applied to LLMs? + +8. **Pressures for adaptability.** The author conjectures that **training for adaptability** creates UFR. What is a concrete training objective that pressures adaptability? (Hint: meta-learning, continual learning, RL with changing environments.) + +9. **The Picbreeder → LLM scaling gap.** Picbreeder CPPNs have ~10-100 nodes. LLMs have ~10⁹ parameters. Does the Picbreeder phenomenon scale? Or is it specific to small networks? + +10. **Comparison to mechanistic interpretability.** The Anthropic circuit-tracing work tries to interpret LLM internals. Are the discovered "circuits" UFR or FER? The arithmetic circuit (magnitude + lookup) suggests FER. + +### 7.3 Applied + +11. **Architectural inductive biases for adaptability.** What architectural choices (besides convolution/attention) bake in adaptability? Mixture-of-experts? Sparse activations? Routing networks? + +12. **Open-ended RL.** RL traditionally has a fixed reward function. Can we design RL environments that are open-ended (no fixed reward, just persistent challenge)? Recent work on procedural environments (e.g., ProcGen, MineRL) is a step in this direction. + +13. **Curriculum learning as serendipity.** The author mentions "order matters for learning." What curricula expose a learner to the right stepping stones? This is the science of curriculum design, currently underexplored. + +14. **Transfer learning for adaptability.** A model trained on task A should adapt quickly to task B if A and B share regularities. Does the Picbreeder UFR transfer? Are there UFR-trained models that transfer well to new tasks? + +### 7.4 Philosophical + +15. **Is the Platonic Representation Hypothesis true?** The author interprets PRH as a statistical phenomenon. Is there a deeper reading where PRH is about **structural convergence** (UFR), not just statistical? Huh et al. 2024 leaves this question open. + +16. **Does an ideal mind's representation match the Space of Forms?** The author's closing question is philosophical. Can it be made precise? The Space of Forms is the **unique decomposition** of the world into independent factors; the ideal representation is the **isomorphism** between the agent's weights and the Space of Forms. This is the strongest version of "Platonic Intelligence." + +--- + +## 8. References + +People, papers, and concepts referenced in the talk and developed in the report. + +### 8.1 People + +| Person | Role | +|---|---| +| Akarsh Kumar | Speaker; MIT CSAIL (PhD student, advisor: Clune) | +| Jeff Clune | Collaborator; UBC, Vector Institute | +| Joel Lehman | Collaborator; University of Oxford (former) | +| Kenneth Stanley | Collaborator; Lila Sciences, UBC (Picbreeder, NEAT, novelty search) | +| Plato | Philosophical antecedent (the Forms) | + +### 8.2 Papers cited in the talk + +- **Secretan et al. (2008).** Picbreeder: A Case Study in Collaborative Human-Computer Evolution. *Evolutionary Computation* (MIT Press). +- **Stanley & Lehman (2015).** *Why Greatness Cannot Be Planned: The Myth of the Objective.* Springer. +- **Mirzadeh et al. (2025).** GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models. *Apple ML Research.* +- **Wu et al. (2024).** Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks. *MIT, Boston University.* +- **Anthropic (2025).** On the Biology of a Large Language Model. *Anthropic Interpretability Team.* +- **Huh et al. (2024).** The Platonic Representation Hypothesis. *ICML 2024.* +- **Kaplan et al. (2020).** Scaling Laws for Neural Language Models. *arXiv:2001.08361.* + +### 8.3 Background concepts and references + +- **Stanley & Miikkulainen (2002).** Evolving Neural Networks through Augmenting Topologies (NEAT). *Evolutionary Computation.* (Foundational neuroevolution algorithm.) +- **Lehman & Stanley (2011).** Abandoning Objectives: Evolution through the Search for Novelty Alone. *Evolutionary Computation.* (Novelty search; abandoning the fixed objective.) +- **Mouret & Clune (2015).** Illuminating Search Spaces by Mapping Elites. *arXiv:1504.04909.* (MAP-Elites; quality-diversity algorithm.) +- **Bronstein et al. (2021).** Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges. *arXiv:2104.13478.* (Inductive biases via group invariance.) +- **Waddington (1942).** Canalization of Development and the Inheritance of Acquired Characters. *Nature.* (Canalization concept.) +- **Stanley (2019).** Why Open-Endedness Matters. *OEE Workshop.* (Philosophical statement of the open-endedness program.) +- **Clune (2019).** AI-GAs: AI-Generating Algorithms, an Alternate Paradigm for Producing General AI. *arXiv:1905.10985.* (AI-GAs as the meta-goal.) + +### 8.4 Internal cross-references + +- **umbrella spec.md** — `conductor/tracks/video_analysis_campaign_20260621/spec.md` — the FR6 8-section report structure. +- **umbrella README.md** — `conductor/tracks/video_analysis_campaign_20260621/README.md` — research-pass framing. +- **child #4 score_dynamics_giorgini** — `conductor/tracks/video_analysis_score_dynamics_giorgini_20260621/report.md` — adjacent cluster A work; both talks about capturing regularities. +- **child #1 cs229_building_llms** — `conductor/tracks/video_analysis_cs229_building_llms_20260621/report.md` — the SGD paradigm Kumar critiques. +- **child #2 probability_logic** — `conductor/tracks/video_analysis_probability_logic_20260621/report.md` — probability foundations. +- **child #3 entropy_epiplexity** — `conductor/tracks/video_analysis_entropy_epiplexity_20260621/report.md` — algorithmic information perspective on representations. +- **child #6 free_lunches_levin** (planned) — open-endedness + algorithmic information. +- **child #7-10 C-cluster** (planned) — complex systems theory. +- **child #11 cs336_architectures** (planned) — the scaling paradigm Kumar critiques. +- **child #12 creikey_dl_cv** (planned) — DDPM as alternative generative approach. + +### 8.5 Code and datasets referenced + +- **Picbreeder** (closed; no longer online; original code archived). +- **NEAT** (Stanley & Miikkulainen) — open-source neuroevolution implementation. +- **MAP-Elites** (Mouret & Clune) — open-source quality-diversity implementation. +- **GPT-3** (Brown et al. 2020) — the LLM in the chicken/duck example. +- **Claude 3.5 Haiku** (Anthropic) — the model in the arithmetic circuit-tracing example. + +--- + +## Appendix A — Concept Map + +Twenty concepts organized by dependency layer. + +**Layer 0 (premises):** +- The world has structure +- Platonic Space of Forms + +**Layer 1 (cognitive claims):** +- Intelligent agents must capture structure +- The capture must be in internal representations + +**Layer 2 (current paradigm):** +- Architectural inductive biases (CNN, transformer) +- The lighting invariance problem +- SGD as fallback + +**Layer 3 (failure modes):** +- FER hypothesis +- Jagged intelligence in LLMs +- Generalization / creativity / continual learning failures + +**Layer 4 (alternative paradigm):** +- Open-ended search +- CPPNs +- Picbreeder +- Why Greatness Cannot Be Planned + +**Layer 5 (intriguing properties):** +- Open-Endedness +- Serendipitous exaptation +- Emergence of evolvability +- Canalization + +**Layer 6 (experimental evidence):** +- Layerization (CPPN → MLP) +- SGD skull vs Picbreeder skull +- SGD butterfly vs Picbreeder butterfly +- SGD apple vs Picbreeder apple + +**Layer 7 (FER in LLMs):** +- GSM-Symbolic (Apple) +- Counterfactual tasks (Reasoning or Reciting) +- Claude 3.5 Haiku circuit tracing (Anthropic) + +**Layer 8 (broader context):** +- Platonic Representation Hypothesis (Huh et al. 2024) +- Scaling Laws (Kaplan et al. 2020) + +**Layer 9 (proposed solution):** +- Open-endedness: complexification, emergence, adaptability, serendipity +- Pressure to adapt as the key driver +- Aspiration: UFR is the Space of Forms + +--- + +## Appendix B — Transcript Excerpts (verbatim, by section) + +### B.1 Motivation + +> "The world is not random but rather it has a lot of structure [...] self similarity across like many spatial scales to like physics symmetries like symmetries in like translation rotation invariance of the world [...] there's this platonic space of forms where these properties, common properties across many objects are inherited from the space of forms." + +> "I claim that basically intelligent agents in order to solve their goals they need to really understand how the world works in order to control it right [...] And more specifically, what I mean is that the internal representations of their minds and their brains must capture the structure of the world." + +### B.2 The lighting invariance problem + +> "But what about all the other um structures in our world, right? All the other regularities. And one example I really like is like lighting invariance. So you see this lion in the dark and during the day and you don't really know what architecture should capture this lighting invariance, right? And we don't know how to do this." + +> "The solution is we just try to train on a lot of data with SGD and hope that the AI will learn this underlying regularity of the world." + +### B.3 The FER hypothesis + +> "Conventional SGD in deep learning um finds neural representations which are actually fractured and entangled [...] the output behavior we visualize it as like a skull [...] but the internal representations don't really match what you would expect to see if it's really like understanding the skull in a subjective way." + +> "Internal representation affects generalization, creativity, and continual learning." + +### B.4 Picbreeder + +> "Picbreeder was an online website where you could evolve the underlying CPPNs or compositional pattern producing networks by um by humans selecting images they liked [...] they didn't tell people what to do they just said do whatever you want." + +> "Picbreeder has Intriguing Properties: Open-Ended, Serendipitous Exaptation, Emergence of Evolvability [...] Canalization, Regularity, Modularity, Symmetry." + +### B.5 SGD vs Picbreeder skull + +> "They they took the picbreeder skull image and they trained like a standard SGD MLP to recreate that skull and it got like a perfect reconstruction right but the SGD MLP is fractured and entangled. Now they took the Picbreeder CPPN and they layerized it which is basically converting it to an MLP. And what they found is that these Picbreeder weights when you sweep them are interpretable. So like the first weight might control mouth opening the second weight might control like eye width." + +### B.6 FER in LLMs + +> "Me: I have 3 pencils, 2 pens, and 4 erasers. How many things do I have? GPT-3: You have 9 things. [always correct] [...] Me: I have 3 chickens, 2 ducks, and 4 geese. How many things do I have? GPT-3: You have 10 animals total. [always incorrect]." + +> "This is from an anthropic mechanism paper [...] Claude [...] 36 + 59 [...] the answer 95 which is correct. But [...] it's using like random heuristics [...] 36 is around 30 and if you add and it's around 40 and if you add around 40 plus around 50 it's like around 92 and it's it's completely like different than how a human would do this arithmetic problem." + +### B.7 The Platonic Representation Hypothesis critique + +> "These scaling laws and this platonic uh representation hypothesis are all very like statistical observations of what's happening [...] it's very unclear how this relates to respecting regularities of the world [...] I think this is a very different um way to see the world than if you just look at this as like a statistical um mechanism." + +### B.8 Open-endedness properties + +> "Complexification (ex: morphogenesis, etc.) Builds regularities on top of other regularities (bottom up) [...] Adaptability [...] Pressures the learned regularities to be robust to environmental changes [...] Serendipity (order matters for learning!) Much higher chance of finding a useful learning curriculum [...] What learning paradigm captures all of these? Open-Endedness!" + +> "Training for adaptability [...] gives you a lot of like regularization pressure to learn like symmetries and regularities to be robust to changes [...] I think a strong representation and adaptable representation are like one and the same." + +### B.9 Q&A: pressure to adapt + +> "I think what you really need to do is like rethink what's going on and what the pressures you are that you're optimizing towards. And I think the pressure to adapt is probably the biggest thing that needs changing. I think that would fix a lot of things." + +### B.10 Q&A: position paper framing + +> "We're not proposing picbreeder as an algorithm that's supposed to compete against and that everyone should use right it's more like to inspire like that fact that this algorithm has some cool properties [...] it's supposed to inspire ideas that maybe we can extract some insights and turn this into an algorithm which can compete against SGD and do better." + +--- + +## Appendix C — Formalizations (expanded) + +### C.1 Group invariance as architectural inductive bias + +Let G be a group acting on input space X. A function f: X → Y is G-equivariant if there exists a group action on Y (also denoted g·y) such that f(g·x) = g·f(x) for all g ∈ G, x ∈ X. A function is G-invariant if g·y = y (trivial action on Y). + +Convolutional layers are equivariant under the translation group G = (ℝ², +). Attention layers are equivariant under permutations. + +Architectural inductive biases are **constraints on the function class** {f_θ : θ ∈ Θ} that enforce specific invariances. The hypothesis is that the true function is invariant under these groups, so restricting to G-equivariant functions improves sample efficiency. + +### C.2 Loss landscape and SGD basins + +The training loss L(θ) is a function Θ → ℝ₊. A **basin of attraction** of a local minimum θ* is the set of initializations θ_0 such that SGD(θ_0) → θ*. + +In high-dimensional spaces, the loss landscape has many local minima of similar depth. SGD's trajectory is determined by: +1. Initialization (typically random in a small ball). +2. Gradient noise (mini-batch stochasticity). +3. Learning rate schedule. + +The basin of attraction of FER minima is typically much larger than the basin of UFR minima because FER is the **generic case** — most random weight configurations that achieve low loss are FER. UFR is a **measure-zero subset** of weight space. + +### C.3 CPPN parameterization + +A CPPN is a directed acyclic graph with heterogeneous activations. Given image coordinates (x, y) ∈ [-1, 1]², the CPPN computes RGB values: + +f: ℝ² → ℝ³, f(x, y) = σ_L(W_L σ_{L-1}(W_{L-1} ... σ_1(W_1 [x, y, 1] + b_1) ... + b_L) + +where σ_i ∈ A (the activation alphabet) and A = {sin, cos, gaussian, sigmoid, abs, linear, ...}. + +The output image is f([-1, 1]²) rendered as a bitmap. + +**Key property:** CPPNs produce **regular** images by construction. sin/cos nodes produce periodic patterns; gaussian nodes produce radial gradients; abs nodes produce symmetries. A random CPPN produces a coherent (if simple) image. + +### C.4 Picbreeder evolution algorithm + +**Algorithm (steady-state NEAT-like):** +1. Initialize population P with random CPPNs. +2. Repeat until user stops: + a. Render current population to images. + b. Present N images to user. + c. User selects one image. + d. Mutate selected CPPN to produce offspring (add node, add connection, perturb weights, change activation). + e. Add offspring to P; remove oldest or least-recently-selected. +3. Output: the current population, which contains the user's interesting CPPNs. + +**Key design choice:** the selection pressure is **subjective aesthetic preference**, not a fixed objective. + +### C.5 Canalization formal definition + +A CPPN with weights w is **canalized** with respect to weight w_i if small perturbations δ to w_i leave the output approximately unchanged: + +‖f(w + δ e_i) − f(w)‖ < ε for all |δ| < δ_0. + +A CPPN is **fully canalized** if it is canalized with respect to all weights not directly affecting any specific output feature. + +Waddington's biological canalization: in development, genetic variation is buffered by the developmental process. In CPPNs, weight variation is buffered by the **network structure** (when the structure has factored semantics). + +### C.6 UFR factorization property + +A CPPN (or its MLP layerization) has the UFR property if there exist functions a_1, ..., a_K (semantic axes) such that: + +∂a_k(Image(w)) / ∂w_i = c_{k,i} δ_{k, σ(i)} + +for some permutation σ and constants c_{k,i}. In other words, weight w_i controls axis a_{σ(i)} and no other axis. + +**Equivalent condition:** the Jacobian ∂Image/∂w has rank-1 structure (each weight affects only one direction in image space). + +SGD-trained MLPs have **dense** Jacobians: every weight affects every axis. This is FER. + +### C.7 LLM arithmetic as bag-of-heuristics + +The Claude 3.5 Haiku circuit (per Anthropic 2025): + +**Input features:** activate for specific digits (3, 6, 5, 9) and approximate magnitudes (around 30, around 50). + +**Add function features:** combine the two addends. Compute (a) the ones-digit (5 + 9 = 14, write 4 carry 1) and (b) the approximate magnitude sum (around 80). + +**Sum features:** output the ones-digit (5) and approximate magnitude (~92). + +**Lookup table features:** for specific seen-during-training pairs (e.g., 36+59=95), use a stored association. + +**Output:** "95" — correct value, wrong mechanism. + +The piecewise structure {lookup, magnitude_heuristic, digit_combination} is **FER** — different problems invoke different mechanisms. A unified mechanism (digit-by-digit addition) would be UFR. + +### C.8 Distinguishing FER from UFR empirically + +**Test 1: Single-weight ablation.** + +For each weight w_i: +- Perturb w_i by small δ. +- Measure change in each semantic axis a_k. +- If only one axis changes (for each i), UFR. +- If multiple axes change, FER. + +**Test 2: Causal intervention.** + +For each axis a_k: +- Find weight w_i that should control a_k (per UFR hypothesis). +- Perturb w_i in the direction that should increase a_k. +- Measure: does only a_k change, or do other axes also change? +- UFR if only a_k changes. + +**Test 3: Sparse probing.** + +Train a linear probe to predict each semantic axis from the network activations. UFR: probes are sparse (one probe per axis, one weight per axis). FER: probes are dense. + +These tests are computable. The Picbreeder results (§C.6) pass all three. The SGD skull fails all three. The Anthropic Claude 3.5 Haiku circuit (per the paper's analysis) fails Tests 1 and 2 — perturbing any one feature affects multiple computational pathways. + +--- + +## Appendix D — Connections (expanded) + +### D.1 To `cs229_building_llms_20260621` (in detail) + +The CS229 lecture on building LLMs presents SGD-based training as the **standard paradigm** for training large neural networks. The lecture covers energy-based models, score matching, and EBM training as alternatives — but does not address the **internal organization** of the trained model. + +Kumar's FER hypothesis directly critiques the SGD paradigm: SGD finds FER, and FER has bad downstream properties (jagged intelligence, brittle generalization, no continual learning). + +**Disagreement:** CS229 would argue that EBMs + score matching are a partial fix — the score function is a well-organized representation. Kumar would counter that the score function is still an entangled representation in the weights; only the score output is well-organized. + +**Resolution:** the score function as a **readout** of the model is well-organized, but the model's **weights** may still be FER. The distinction between "well-organized readout" and "well-organized weights" is the key open question. + +### D.2 To `entropy_epiplexity_20260621` (in detail) + +The epiplexity talk covers Kolmogorov complexity and the **observer-relative** nature of complexity. A representation has high epiplexity if it is complex from the observer's perspective. + +**Connection to UFR:** UFR has low epiplexity from the agent's perspective — each weight maps to one semantic axis, which is a simple mapping. FER has high epiplexity from the agent's perspective — each weight maps to multiple semantic axes, which is a complex mapping. + +**Hypothesis:** SGD finds FER because SGD's loss function is a **low-epiplexity observer** — it accepts any low-loss solution, regardless of how complex the weights are. Open-ended search (with humans in the loop) is a **high-epiplexity observer** — it requires the weights to be interpretable, which biases toward UFR. + +This connection is speculative. Pass 2 should explore it. + +### D.3 To `score_dynamics_giorgini_20260621` (in detail) + +The score-based generative modeling talk presents the score function as the central primitive for capturing regularities. The score is the gradient of the log-density — it encodes the local geometry of the data distribution. + +**Connection to UFR:** an ideal UFR for a generative model would have weights that each encode a **direction in score space**. The factorization of the representation corresponds to the factorization of the score into independent directions. + +**Hypothesis:** the Picbreeder CPPNs have UFR because each neuron corresponds to a direction in score space that is interpretable (mouth opening is a direction in score space; eye width is another). The SGD MLPs have FER because the directions are entangled. + +**Test:** train an MLP with a score-matching loss (Giorgini's framework) and test for UFR via the §C.8 tests. If the score-trained MLP has UFR, the connection is confirmed. + +### D.4 To `cs336_architectures_20260621` (planned) + +The CS336 lecture on language model architectures covers diffusion LMs and modern LLM training. The FER diagnosis predicts that scaling up LLMs will not solve the brittleness — the representations remain FER regardless of scale. + +**Prediction:** LLM benchmark performance will continue to improve with scale, but **counterfactual and out-of-distribution performance** will continue to be brittle. The same magnitude-based arithmetic will be used for harder problems; the same surface-pattern matching will be used for harder reasoning. + +**Implication for CS336 curriculum:** the lecture's emphasis on scaling should be balanced with discussion of representation quality. Diffusion LMs (with score-matching) may have better representations than autoregressive LMs (with cross-entropy), but the evidence is sparse. + +### D.5 To `free_lunches_levin_20260621` (planned) + +Levin's "free lunches" talk likely covers algorithmic information theory and the relationship between search algorithms and problem structure. The "free lunch" in Kumar's framing: **regularity in the world** means that some search algorithms get "free" structure capture. + +**Connection:** the algorithmic-information perspective quantifies "regularity" as low Kolmogorov complexity. A world with structure has low complexity (in some sense); a search algorithm that exploits this structure gets free improvement. + +**Hypothesis:** open-ended search implicitly optimizes for **low Kolmogorov complexity of the representation** (because simple representations are more evolvable, more generalizable, more interpretable). SGD does not — it optimizes for low loss, which can be achieved by complex representations. + +--- + +## Appendix E — Open Questions (expanded) + +### E.1 Theoretical questions + +**E.1.1 Formal definition of UFR.** A precise mathematical definition of UFR requires: +- A notion of "semantic axis" (a function a_k: Image → ℝ that is human-interpretable). +- A factorization property: each weight controls one axis. +- A uniqueness property: the factorization is canonical (not depending on the choice of axes). + +The Picbreeder results show that CPPNs admit UFR (with human-interpretable axes). A general theory would characterize when a function admits UFR. + +**E.1.2 Existence for general functions.** Do all computable functions admit UFR? No — the UFR is a special property. The class of functions with UFR is the class that **decompose into independent factors** in a human-meaningful way. Most functions (in the Kolmogorov sense) do not have UFR. + +**E.1.3 Algorithmic complexity.** The UFR has lower Kolmogorov complexity than the FER for the same function (because the weights encode a simpler structure). Is the UFR the minimum-description-length representation? + +### E.2 Empirical questions + +**E.2.1 LLM-UFR.** Do LLMs admit UFR? The Anthropic circuit-tracing work suggests no — the arithmetic circuit is a bag of heuristics, not a unified algorithm. But this is one example. A systematic study of UFR across LLM tasks is open work. + +**E.2.2 Direct UFR training.** What is a practical algorithm that finds UFR without humans in the loop? Candidates: +- **Quality-Diversity algorithms** (MAP-Elites, Novelty Search): maintain a diverse archive of solutions; select for diversity rather than fitness. +- **Meta-learning for adaptability**: train on a distribution of tasks; select for adaptability rather than task performance. +- **Mechanistic regularization**: add a loss term that encourages interpretable weights. + +**E.2.3 Scalability.** Does the Picbreeder phenomenon scale to LLM-sized models? Picbreeder CPPNs have ~10-100 nodes; LLMs have ~10⁹ parameters. The basin-of-attraction argument (§C.2) suggests that the UFR basin shrinks as dimension grows. Whether UFR exists at all at LLM scale is open. + +### E.3 Applied questions + +**E.3.1 Architectural inductive biases for adaptability.** Beyond convolution and attention, what other architectural choices bake in adaptability? Sparse activations (mixture-of-experts) encourage modularity; routing networks encourage specialization. + +**E.3.2 Open-ended RL.** Can we design RL environments that are open-ended? Procedural environments (ProcGen, MineRL) are a step; truly open-ended environments (no fixed objective, persistent challenge) are an active research area. + +**E.3.3 Curriculum learning.** "Order matters for learning" — what curricula expose the right stepping stones? This is the science of curriculum design. + +### E.4 Philosophical questions + +**E.4.1 Platonic Representation Hypothesis as structural.** Is PRH about statistical convergence (similar representation geometries) or structural convergence (UFR-like factorization)? The current evidence (Huh et al. 2024) supports the statistical reading. The structural reading would require UFR tests (§C.8) applied to cross-modal representations. + +**E.4.2 The Space of Forms.** The author's closing question: does an ideal mind's representation match the Space of Forms? This is the strongest version of Platonic Intelligence. Whether it is achievable (even in principle) is open. + +--- + +## Appendix F — References (full bibliography) + +### F.1 Primary works cited + +1. Secretan, J., Stanley, K. O., et al. (2008). Picbreeder: A Case Study in Collaborative Human-Computer Evolution. *Evolutionary Computation*, 16(4), 565-587. +2. Stanley, K. O., & Lehman, J. (2015). *Why Greatness Cannot Be Planned: The Myth of the Objective.* Springer. +3. Mirzadeh, I., Alizadeh, K., Shahrokhi, H., Bengio, S., Farajtabar, M., & Tuzel, O. (2025). GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models. *Apple ML Research.* +4. Wu, Z., Qiu, L., Ross, A., Akyürek, E., Chen, B., Wang, B., Kim, N., Andreas, J., & Kim, Y. (2024). Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks. +5. Anthropic (2025). On the Biology of a Large Language Model. *Anthropic Interpretability.* +6. Huh, M., Cheung, B., Wang, T., & Isola, P. (2024). The Platonic Representation Hypothesis. *ICML 2024.* +7. Kaplan, J., et al. (2020). Scaling Laws for Neural Language Models. *arXiv:2001.08361.* + +### F.2 Foundational neuroevolution references + +8. Stanley, K. O., & Miikkulainen, R. (2002). Evolving Neural Networks through Augmenting Topologies. *Evolutionary Computation*, 10(2), 99-127. +9. Lehman, J., & Stanley, K. O. (2011). Abandoning Objectives: Evolution through the Search for Novelty Alone. *Evolutionary Computation*, 19(2), 189-223. +10. Mouret, J.-B., & Clune, J. (2015). Illuminating Search Spaces by Mapping Elites. *arXiv:1504.04909.* +11. Stanley, K. O., et al. (2019). Open-Endedness: The Last Grand Challenge You've Never Heard Of. *OEE Workshop.* +12. Clune, J. (2019). AI-GAs: AI-Generating Algorithms, an Alternate Paradigm for Producing General AI. *arXiv:1905.10985.* + +### F.3 Geometric deep learning + +13. Bronstein, M. M., Bruna, J., Cohen, T., & Veličković, P. (2021). Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges. *arXiv:2104.13478.* + +### F.4 Canalization and evolutionary biology + +14. Waddington, C. H. (1942). Canalization of Development and the Inheritance of Acquired Characters. *Nature*, 150, 563-565. +15. Waddington, C. H. (1957). *The Strategy of the Genes.* Allen & Unwin. + +### F.5 LLM architecture and training + +16. Vaswani, A., et al. (2017). Attention Is All You Need. *NeurIPS.* +17. Brown, T., et al. (2020). Language Models are Few-Shot Learners. *NeurIPS 2020* (GPT-3). +18. Hoffmann, J., et al. (2022). Training Compute-Optimal Large Language Models. *arXiv:2203.15556* (Chinchilla scaling laws). + +### F.6 Mechanistic interpretability + +19. Elhage, N., et al. (2022). Softmax Linear Units. *Anthropic.* +20. Olsson, C., et al. (2022). In-context Learning and Induction Heads. *Anthropic.* +21. Anthropic (2024). Mapping the Mind of a Large Language Model. *Anthropic.* +22. Anthropic (2025). On the Biology of a Large Language Model. *Anthropic Interpretability.* + +### F.7 Picbreeder-related references + +23. Stanley, K. O. (2007). Compositional Pattern Producing Networks: A Novel Abstraction of Development. *Genetic Programming and Evolvable Machines*, 8(2), 131-162. + +--- + +## Appendix G — Cross-references within campaign + +### G.1 Forward references + +- **free_lunches_levin_20260621** (planned): open-endedness + algorithmic information. +- **cs336_architectures_20260621** (planned): the scaling paradigm Kumar critiques. +- **creikey_dl_cv_20260621** (planned): DDPM as alternative generative approach. + +### G.2 Backward references + +- **cs229_building_llms_20260621** (§6.1.1): the SGD paradigm Kumar critiques. +- **probability_logic_20260621** (§6.1.2): probability foundations for "regularity." +- **entropy_epiplexity_20260621** (§6.1.3): algorithmic information perspective on representations. +- **score_dynamics_giorgini_20260621** (§6.1.4): score-based generative modeling; an alternative route to capturing regularities. + +### G.3 Lateral references + +- **generic_systems_fields_20260621** (planned, cluster C): complex systems theory. +- **brain_counterintuitive_20260621** (planned, cluster C): brain as existence proof of UFR. +- **neural_dynamics_miller_20260621** (planned, cluster C): neural dynamics. +- **multiscale_hoffman_20260621** (planned, cluster C): multiscale phenomena. + +### G.4 Reference dependency graph + +``` +foundations: + probability_logic + | + v + "regularity" as a concept + | + +----> entropy_epiplexity (algorithmic info perspective) + | | + | v + | "representation = program" + | | + | v + +----> UFR = low-complexity representation + | + v + open-ended search (this talk) + | + +----> cs229 (SGD paradigm critiqued) + | + +----> score_dynamics_giorgini (alternative route) + | | + | v + | score = gradient of log-density + | | + | v + | factorization of score = UFR? + | + +----> cs336 (scaling doesn't fix FER) + | + +----> creikey_dl_cv (DDPM as alternative) + | + +----> free_lunches_levin (open-endedness + info) + | + +----> brain_* (existence proofs of UFR) + | + +----> generic_systems_fields (complex systems theory) +``` + +--- + +## Appendix H — Synthesis Summary + +A single-paragraph TL;DR of the talk, suitable for a busy reader. + +Kumar's talk presents a position paper arguing that conventional SGD finds **Fractured Entangled Representations (FER)** — neural networks that produce the right output behavior but with internal weights that do not correspond to human-interpretable factors — and proposes **open-ended search** as a paradigm that finds **Unified Factored Representations (UFR)**. Evidence comes from Picbreeder: when humans evolve CPPNs by selecting images they like (no fixed objective), the resulting CPPNs layerize to MLPs where each weight controls one semantic axis (mouth opening, eye width). SGD-trained MLPs that reproduce the same image have entangled weights with no such factorization. The same FER diagnosis predicts LLM failure modes observed in three recent papers: GPT-3's chicken/duck counting failure, GPT-4's counterfactual-task degradation, and Claude 3.5 Haiku's magnitude-heuristic arithmetic (per Anthropic circuit tracing). The author critiques the **Platonic Representation Hypothesis** (Huh et al. 2024) as a **statistical** observation about representational geometry that doesn't address whether representations are factored. The proposed solution is **open-endedness** with four properties: complexification (building regularities on top of regularities, bottom-up), emergence (higher-level structure from lower-level dynamics), adaptability (training for the ability to adapt, not for any specific task), and serendipity (multiple environments in sequence). Among these, **pressure to adapt** is conjectured to be the most important driver of UFR. + +--- + +## Appendix I — Personal Notes + +Things to revisit in Pass 2 (the user's de-obfuscation pass). + +1. The "fer fract" claim is **empirical**, not theoretical. The author presents no formal proof that SGD cannot find UFR — just empirical observations and the Q&A admission that even with strong regularization, SGD rarely does. A theoretical argument about the SGD loss landscape would strengthen the claim. + +2. The **position paper** framing means the talk is explicitly not an algorithm proposal. The author acknowledges "we cheated" with humans in the loop. A follow-up algorithm (MAP-Elites + LLMs? Quality-Diversity RL?) would be the natural next step. + +3. The **connection to score_dynamics_giorgini** is speculative but tantalizing. An MLP trained with score-matching loss (rather than cross-entropy) might have UFR — the score is a well-organized readout, and the weights might be correspondingly organized. Pass 2 should explore this. + +4. The **adaptability-as-driver** hypothesis needs an empirical test. Train an MLP on a distribution of tasks (meta-learning); test for UFR. If meta-learned MLPs have UFR, the hypothesis is confirmed. + +5. The **Space of Forms** framing is philosophical. The author's closing question — does an ideal mind's representation match the Space of Forms? — invites a mathematical formalization. The Space of Forms is the unique decomposition of the world into independent factors; the ideal representation is the isomorphism between agent weights and factors. Whether this isomorphism exists in general is open. + +6. The **Picbreeder → LLM scaling gap** is the most important open question. Picbreeder CPPNs are small (~100 nodes); LLMs are ~10⁹ parameters. The Picbreeder phenomenon may be specific to small networks. A controlled scaling experiment (Picbreeder-style evolution at increasing scale) would clarify. + +7. The **arithmetic circuit in Claude 3.5 Haiku** is one example. A systematic survey across LLM tasks (counting, sorting, code execution, drawing) would establish whether LLM-internal computations are uniformly FER or task-dependent. + +--- + +## Appendix J — Glossary + +| Term | Definition | +|---|---| +| **CPPN** | Compositional Pattern Producing Network. A neural net with heterogeneous activations, used as the genotype in Picbreeder. | +| **Picbreeder** | Stanley & Lehman's 2008 web app where humans evolved CPPNs by selecting images they liked. | +| **UFR** | Unified Factored Representation. Internal representation where each weight controls one semantic axis. | +| **FER** | Fractured Entangled Representation. Internal representation where weights have no interpretable correspondence to semantic axes. | +| **SGD** | Stochastic Gradient Descent. The standard optimization algorithm for training neural networks. | +| **NEAT** | NeuroEvolution of Augmenting Topologies. Stanley & Miikkulainen's 2002 neuroevolution algorithm. | +| **MAP-Elites** | Mouret & Clune's quality-diversity algorithm that maintains a diverse archive of high-performing solutions. | +| **Canalization** | Waddington's concept: a genotype where small mutations produce small phenotypic changes. | +| **Evolvability** | The capacity for a genotype to produce coherent, useful variation under mutation. | +| **Open-endedness** | A search process without a fixed objective, where the interesting region is unbounded. | +| **Serendipity** | Discovery of valuable stepping stones that have nothing to do with the original goal. | +| **Exaptation** | Repurposing of a trait evolved for one function to a new function. | +| **Complexification** | Building regularities on top of regularities, bottom-up (like morphogenesis). | +| **Adaptability** | Capacity to maintain performance in changing environments. | +| **Platonic Representation Hypothesis** | Huh et al. 2024 claim that different modalities are converging to a shared representation. | +| **Space of Forms** | Plato's philosophical concept: an abstract space containing the ideal patterns that objects instantiate. | +| **Inductive bias** | Architectural or training-data prior that biases the learned function toward a specific class. | +| **Layerization** | Converting a CPPN (heterogeneous activations) to an MLP (uniform activations) for fair comparison. | +| **Jagged intelligence** | Inconsistency in capability: very good at hard tasks, bad at easy tasks. | +| **Mechanistic interpretability** | Reverse-engineering the internal computations of a trained neural network. | +| **Circuit** | A pattern of features in a neural network that implements a specific computation. | +| **Bag-of-heuristics** | A computation that uses multiple specialized heuristics rather than a unified algorithm. | +| **GSM8K** | Grade-School Math 8K. A benchmark of 8K grade-school-level math problems for LLM evaluation. | +| **GSM-Symbolic** | Mirzadeh et al. 2025 improvement on GSM8K with symbolic templates. | +| **Counterfactual task** | A task variant that systematically perturbs the surface form while preserving the underlying structure. | +| **Space of Forms** | Plato's philosophical concept; in ML, the ideal representation space that captures world regularities. | + +--- + +*End of report. Lossless preservation per umbrella spec §0. Pass 2 (de-obfuscation) and Pass 3 (projection to applied domain) to follow.* diff --git a/conductor/tracks/video_analysis_platonic_intelligence_kumar_20260621/summary.md b/conductor/tracks/video_analysis_platonic_intelligence_kumar_20260621/summary.md new file mode 100644 index 00000000..11c7612a --- /dev/null +++ b/conductor/tracks/video_analysis_platonic_intelligence_kumar_20260621/summary.md @@ -0,0 +1,25 @@ +# Summary: Towards a Platonic Intelligence (Kumar) + +**Source:** https://youtu.be/1mXUFweWOug +**Author:** Akarsh Kumar (MIT CSAIL) +**Track:** Child #5 of `video_analysis_campaign_20260621` +**Cluster:** B (Platonic / geometric AI representations) +**Pass:** 1 of 3 (research-only deep-dive) + +--- + +## One-paragraph synthesis + +Kumar presents a **position paper** arguing that conventional SGD finds **Fractured Entangled Representations (FER)** — neural networks producing the right output behavior but with weights that do not correspond to human-interpretable factors — and proposes **open-ended search** as a paradigm that finds **Unified Factored Representations (UFR)**. Evidence from Picbreeder (Stanley & Lehman, 2008): when humans evolve CPPNs by selecting images they like with no fixed objective, the resulting CPPNs layerize to MLPs where each weight controls one semantic axis (mouth opening, eye width). SGD-trained MLPs reproducing the same images have entangled weights. The FER diagnosis predicts three LLM failure modes: GPT-3's chicken/duck/pencil counting failure (Mirzadeh et al. GSM-Symbolic), GPT-4's counterfactual-task degradation (Wu et al. "Reasoning or Reciting"), and Claude 3.5 Haiku's magnitude-heuristic arithmetic (Anthropic circuit tracing) — bag-of-heuristics rather than unified algorithms. Kumar critiques the **Platonic Representation Hypothesis** (Huh et al. 2024) as a **statistical** observation that doesn't address factorization. The proposed solution is **open-endedness** with four properties: complexification (building regularities on top of regularities, bottom-up), emergence, adaptability (training for the ability to adapt, not for any specific task — the most important driver per the author), and serendipity. **Backward connections:** score_dynamics_giorgini (alternative route via score matching), entropy_epiplexity (algorithmic-information perspective), cs229_building_llms (the SGD paradigm critiqued). **Forward connections:** cs336_architectures (scaling doesn't fix FER), creikey_dl_cv (DDPM as alternative), free_lunches_levin (open-endedness + information theory). + +--- + +## Three key takeaways + +1. **SGD finds FER, open-ended search finds UFR** — same loss, same architecture, completely different internal organization. Picbreeder CPPNs layerize to MLPs where each weight controls one semantic axis; SGD MLPs have entangled weights. +2. **FER predicts LLM jagged intelligence** — GPT-3 fails at chicken/duck counting (surface pattern), GPT-4 fails at counterfactual tasks, Claude uses magnitude heuristics not digit-by-digit arithmetic. Correct answers, wrong internal mechanisms. +3. **Pressure to adapt creates UFR** — the most actionable conjecture. Train for adaptability to changing environments (not any specific task), and representations will capture underlying regularities. + +--- + +*Pass 2 (de-obfuscation via user's mathematical encoding) to follow.*