All 5 phases done, 34 tests passing, VC1-VC7 pass. Task 5.3 (live 8-URL smoke) + 5.6 (user manual verification) deferred to user (network/cookies + merge decision).
17 KiB
Plan: Twitter/X Thread Extraction Tooling
Track: twitter_threads_extraction_20260705
Branch: master (scripts + tests only; no src/ changes, no GUI changes)
Spec: conductor/tracks/twitter_threads_extraction_20260705/spec.md
Standalone tooling track. The deliverable is a scripts/twitter_threads/ package + tests + README. No application code changes.
Phase 1: Scaffold + Error Types + Data Classes [checkpoint: 161d8da]
Focus: create the package directory, the shared error type, and the typed dataclasses (ThreadData, PostData) that the rest of the pipeline passes around.
- Task 1.1: Create the
scripts/twitter_threads/directory +__init__.py
WHERE: scripts/twitter_threads/__init__.py
WHAT: A docstring documenting the namespace, the per-module responsibilities, and the standalone-usage note (no src/ imports; copy-pasteable to another repo).
HOW: Mirror scripts/video_analysis/__init__.py.
- Task 1.2: Write
error_types.py(the shared Result[T, ErrorInfo] shape)
WHERE: scripts/twitter_threads/error_types.py
WHAT: Copy the shape of scripts/video_analysis/error_types.py: ErrorInfo dataclass (frozen, slots) + make_error factory. Standalone — no import from scripts.video_analysis.
HOW:
from __future__ import annotations
from dataclasses import dataclass
@dataclass(frozen=True, slots=True)
class ErrorInfo:
kind: str
source: str
detail: str
def make_error(kind: str, source: str, detail: str) -> ErrorInfo:
return ErrorInfo(kind=kind, source=source, detail=detail)
VERIFY: python -c "from scripts.twitter_threads.error_types import ErrorInfo, make_error; print(make_error('X','y','z'))" prints the dataclass repr.
- Task 1.3: Write the typed dataclasses (
ThreadData,PostData)
WHERE: scripts/twitter_threads/error_types.py (same file; keeps the standalone file count low) OR a new scripts/twitter_threads/types.py (if the user prefers separation). Default: same file.
WHAT:
@dataclass(frozen=True, slots=True)
class PostData:
post_id: str
author: str
handle: str
text: str
timestamp: str # ISO 8601
media_urls: tuple[str, ...]
reply_to_id: str | None
quote_of_id: str | None
metrics: PostMetrics
@dataclass(frozen=True, slots=True)
class PostMetrics:
reply_count: int
repost_count: int
like_count: int
view_count: int | None
@dataclass(frozen=True, slots=True)
class ThreadData:
root_post_id: str
posts: tuple[PostData, ...]
source_url: str
HOW: Per conductor/code_styleguides/data_oriented_design.md §8.5 — typed, frozen, slots. No dict[str, Any].
VERIFY: python -c "from scripts.twitter_threads.error_types import PostData; print(PostData.__slots__)" prints the slots tuple.
- Task 1.4: Write tests for the error types + dataclasses
WHERE: tests/test_twitter_threads_types.py
WHAT: Construct ErrorInfo, PostData, ThreadData instances. Assert field access, frozen-ness (mutation raises FrozenInstanceError), slots-ness (no __dict__).
HOW: Per the TDD red-first protocol — write the tests first, run them (they fail because the types don't exist yet), then implement Task 1.2 + 1.3 to make them pass.
VERIFY: uv run pytest tests/test_twitter_threads_types.py -v passes.
- Task 1.5: Commit Phase 1
git add scripts/twitter_threads/__init__.py scripts/twitter_threads/error_types.py tests/test_twitter_threads_types.py
git commit -m "feat(twitter_threads): scaffold package + error types + typed dataclasses"
Phase 2: render_markdown.py (the pure-function module; TDD-first) [checkpoint: ef98f6d1]
Focus: the Markdown emission module. This is pure (no network, no subprocess), so it's the easiest to TDD and the highest-confidence module.
- Task 2.1: Write failing tests for
render_markdown.py
WHERE: tests/test_twitter_threads_render.py
WHAT: 5+ tests:
test_render_single_post— aThreadDatawith 1 post; assert the YAML front-matter has the right fields; assert the post body is rendered; assert media links are./media/<filename>.test_render_thread— aThreadDatawith 3 posts (a thread); assert 3## Post Nsections in chronological order; assert each has the timestamp + reply-to marker.test_render_quote_tweet— aThreadDatawith a post that hasquote_of_idset; assert the quote-tweet renders as a nested blockquote with a link.test_render_media_links— a post with 2 image URLs + 1 video URL; assert 3 media links, named<post_id>_img1.jpg,<post_id>_img2.jpg,<post_id>_vid1.mp4.test_render_metrics_in_frontmatter— assertreply_count,repost_count,like_count,view_countappear in the YAML front-matter; assertview_count: nullwhenview_countisNone.test_render_title_from_first_post— the thread title is the first post's first line (truncated to 80 chars). HOW: ConstructThreadDatafixtures inline (no network). Assert on the returned Markdown string. VERIFY:uv run pytest tests/test_twitter_threads_render.pyFAILS (module doesn't exist yet).
- Task 2.2: Implement
render_markdown.py
WHERE: scripts/twitter_threads/render_markdown.py
WHAT: The render_markdown(thread: ThreadData, media_paths: dict[str, list[Path]], output: Path) -> Result[Path, ErrorInfo] function.
HOW:
-
Build the YAML front-matter from
thread.posts[0](the root post's metadata). -
Build the body: iterate
thread.postsin order; for each, emit## Post N (<timestamp>)+ optional— reply to Post Mmarker. -
For each
media_urlinpost.media_urls: emit[Media K](./media/<post_id>_<kind><K>.<ext>). -
For quote-tweets: emit a
> [Quoting @<handle>](<quoted_url>)blockquote. -
Write to
outputviaPath.write_text(content, encoding="utf-8"). -
Return
Result.ok(output)on success;Result.err(...)on file-write failure. VERIFY:uv run pytest tests/test_twitter_threads_render.pyPASSES. -
Task 2.3: Commit Phase 2
git add scripts/twitter_threads/render_markdown.py tests/test_twitter_threads_render.py
git commit -m "feat(twitter_threads): render_markdown.py — YAML front-matter + per-post sections + media links"
Phase 3: download_media.py (the media downloader) [checkpoint: f503eb5d]
Focus: download the media files referenced by a ThreadData. Mock the HTTP in tests; real network only at runtime.
- Task 3.1: Write failing tests for
download_media.py
WHERE: tests/test_twitter_threads_media.py
WHAT: 4+ tests:
test_download_naming— a post with 2 image URLs (...?format=jpg&name=large,...?format=png&name=large); assert the downloaded files are named<post_id>_img1.jpgand<post_id>_img2.png.test_download_video_naming— a post with a video URL (.../vid/.../1234567890.mp4); assert the file is named<post_id>_vid1.mp4.test_download_idempotent— pre-create the target file with the expected byte size; assert the download is skipped (no HTTP call made).test_download_http_error— mockurlopento raiseURLError; assertResult.errwithErrorInfo.kind == "HttpError". HOW: Mockurllib.request.urlopenviaunittest.mock.patch(this is a boundary test — the HTTP boundary — so mocking is allowed per the structural testing contract). Usetests/artifacts/twitter_threads_media_<test_name>/as the output dir per the workspace-paths convention. VERIFY:uv run pytest tests/test_twitter_threads_media.pyFAILS.
- Task 3.2: Implement
download_media.py
WHERE: scripts/twitter_threads/download_media.py
WHAT: The download_media(thread: ThreadData, output_dir: Path) -> Result[list[Path], ErrorInfo] function.
HOW:
-
For each post in
thread.posts, for eachmedia_urlinpost.media_urls:- Derive the kind (
img/vid/gif) from the URL or Content-Type. - Derive the extension (
.jpg,.png,.mp4). - Compute the target filename:
<post_id>_<kind><index>.<ext>. - If the target file exists and has byte size > 0: skip (idempotent).
- Else:
urllib.request.urlopen(media_url)→ read bytes → write to target.
- Derive the kind (
-
Return
Result.ok(list_of_paths)on success;Result.err(ErrorInfo)on the first HTTP failure. -
Use only stdlib (
urllib.request,pathlib,dataclasses). Norequests. VERIFY:uv run pytest tests/test_twitter_threads_media.pyPASSES. -
Task 3.3: Commit Phase 3
git add scripts/twitter_threads/download_media.py tests/test_twitter_threads_media.py
git commit -m "feat(twitter_threads): download_media.py — stdlib urllib + idempotent + typed naming"
Phase 4: fetch_thread.py (the acquisition module; the network boundary) [checkpoint: 06ff9299]
Focus: acquire a ThreadData from either a URL (via gallery-dl subprocess) or a local HTML file (Strategy C parse). The URL path is the network boundary; the HTML path is the offline fallback.
- Task 4.1: Write failing tests for the local-HTML-parse path (Strategy C)
WHERE: tests/test_twitter_threads_fetch.py
WHAT: 3+ tests:
test_parse_html_single_post— a small fixture HTML file (intests/artifacts/) containing one tweet's text + author + timestamp + 1 image URL; assert the parsedThreadDatahas 1 post with the right fields.test_parse_html_thread— a fixture HTML file with a 3-post thread; assert 3PostDatainstances in chronological order.test_parse_html_no_media— a post with no media; assertmedia_urls == ().test_parse_html_malformed— a malformed HTML file; assertResult.errwithErrorInfo.kind == "ParseError". HOW: Usehtml.parser.HTMLParser(stdlib) — the test fixtures are real HTML strings. No network. VERIFY:uv run pytest tests/test_twitter_threads_fetch.pyFAILS.
- Task 4.2: Implement the local-HTML-parse path
WHERE: scripts/twitter_threads/fetch_thread.py
WHAT: A fetch_thread_from_html(html_path: Path) -> Result[ThreadData, ErrorInfo] function that parses a local HTML file using html.parser.HTMLParser (stdlib, no BeautifulSoup).
HOW:
-
Subclass
HTMLParser; walk the DOM; extract post text (the tweet-text container), author, handle, timestamp, media URLs (theimg[src]andvideo[src]tags in the tweet container). -
Construct
PostDatainstances; chain them into aThreadData. -
Return
Result.err(ErrorInfo("ParseError", "fetch_thread_from_html", detail))on malformed HTML. VERIFY:uv run pytest tests/test_twitter_threads_fetch.pyPASSES for the Strategy C tests. -
Task 4.3: Implement the URL path (the
gallery-dlsubprocess wrapper)
WHERE: scripts/twitter_threads/fetch_thread.py (same file; add the URL path)
WHAT: A fetch_thread_from_url(url: str, cookies_path: Path | None = None) -> Result[ThreadData, ErrorInfo] function.
HOW:
-
Build the
gallery-dlargs:["gallery-dl", "--dump-json", "--write-metadata", url]. Ifcookies_pathis provided, add--cookies <cookies_path>. -
subprocess.run(args, capture_output=True, text=True). -
Parse the JSON stdout:
gallery-dlemits one JSON object per post. Parse each into aPostData. -
Chain into
ThreadData. -
On
gallery-dlnon-zero exit: returnResult.err(ErrorInfo("GalleryDlError", "fetch_thread_from_url", stderr[:500])). -
On JSON parse failure: return
Result.err(ErrorInfo("JsonParseError", ...)). VERIFY: Manual smoke test against a real X.com URL (the user provides a URL + cookies file). The automated tests cover the HTML path only (Strategy C); the URL path is the network boundary and is smoke-tested manually. -
Task 4.4: Implement the CLI dispatch +
--help
WHERE: scripts/twitter_threads/fetch_thread.py (the if __name__ == "__main__": block)
WHAT: A CLI that takes a URL or a local HTML path + --output + optional --cookies, dispatches to the right function, and writes the ThreadData to a JSON file (intermediate) for downstream download_media.py + render_markdown.py to consume.
HOW:
if __name__ == "__main__":
import argparse
parser = argparse.ArgumentParser(description="Fetch a Twitter/X thread into a ThreadData JSON.")
parser.add_argument("source", help="X.com URL or local HTML file path")
parser.add_argument("--output", type=Path, required=True, help="Output directory")
parser.add_argument("--cookies", type=Path, default=None, help="cookies.txt for gallery-dl auth")
args = parser.parse_args()
...
VERIFY: python scripts/twitter_threads/fetch_thread.py --help works (standalone, no src/ imports).
- Task 4.5: Commit Phase 4
git add scripts/twitter_threads/fetch_thread.py tests/test_twitter_threads_fetch.py
git commit -m "feat(twitter_threads): fetch_thread.py — gallery-dl subprocess (URL) + html.parser (local HTML fallback)"
Phase 5: README + End-to-End CLI + Verification [checkpoint: fe207c1e]
Focus: the standalone-usage README, an end-to-end CLI test, and the final verification.
- Task 5.1: Write
README.md
WHERE: scripts/twitter_threads/README.md
WHAT: Per spec FR7 — prerequisites (gallery-dl, cookies.txt, Python 3.11+), usage (both uv run and standalone python), output layout, "copy to another repo" instructions, and the Strategy B/D documentation as alternatives.
HOW: Markdown; no code.
- Task 5.2: Verify the standalone requirement (VC6 + VC7)
WHAT:
-
python scripts/twitter_threads/fetch_thread.py --helpworks from the repo root with nouv run(standalone invocation). -
grep -r "from src\." scripts/twitter_threads/returns nothing. -
grep -r "import src\." scripts/twitter_threads/returns nothing. -
grep -r "from conductor\." scripts/twitter_threads/returns nothing. -
grep -r "from scripts\.video_analysis" scripts/twitter_threads/returns nothing. VERIFY: All greps return empty; the--helpworks. -
[~] Task 5.3: End-to-end smoke test against the 8-thread corpus (@NOTimothyLottes + @VPCOMPRESSB) — DEFERRED to user (needs network + cookies.txt; pipeline mechanics verified by test_twitter_threads_pipeline.py)
WHAT: Run the full pipeline against all 8 reference URLs from the spec's "Reference Project + Test Corpus" section. These are the acceptance corpus for the C:\projects\forth\bootslop reference-generation pipeline.
URLs:
https://x.com/NOTimothyLottes/status/1757198624818168210https://x.com/NOTimothyLottes/status/1653570742762479620https://x.com/NOTimothyLottes/status/1917646466417381426https://x.com/NOTimothyLottes/status/1917645859791200562https://x.com/NOTimothyLottes/status/1917644904055910502https://x.com/NOTimothyLottes/status/1917642786804785230https://x.com/VPCOMPRESSB/status/1991383117571957052(note:?s=20query suffix must be stripped byfetch_thread.pybefore acquisition)https://x.com/VPCOMPRESSB/status/1987744335333622188(note:?s=20query suffix must be stripped) HOW (per URL):
uv run python -m scripts.twitter_threads.fetch_thread "<url>" --output ./tests/artifacts/twitter_threads_corpus/ --cookies ./cookies.txt
uv run python -m scripts.twitter_threads.download_media --input ./tests/artifacts/twitter_threads_corpus/<id>/thread_data.json --output ./tests/artifacts/twitter_threads_corpus/<id>/media/
uv run python -m scripts.twitter_threads.render_markdown --input ./tests/artifacts/twitter_threads_corpus/<id>/thread_data.json --media-dir ./tests/artifacts/twitter_threads_corpus/<id>/media/ --output ./tests/artifacts/twitter_threads_corpus/<id>/thread.md
VERIFY: 8 thread.md files exist with YAML front-matter + post sections + media links; 8 media/ directories have the downloaded assets. This is the acceptance corpus for the track — if all 8 extract cleanly, the track ships and the corpus is handed off to bootslop. This is a manual verification (network boundary; the user provides the cookies.txt).
- Task 5.4: Run the full test suite for the new tests
WHAT: uv run pytest tests/test_twitter_threads_types.py tests/test_twitter_threads_render.py tests/test_twitter_threads_media.py tests/test_twitter_threads_fetch.py -v
VERIFY: All tests pass. This is the batched verification (the only verification that matters per the Isolated-Pass Verification Fallacy rule).
- Task 5.5: Commit Phase 5 + the README
git add scripts/twitter_threads/README.md
git commit -m "docs(twitter_threads): README — standalone usage + copy-to-another-repo instructions"
- [~] Task 5.6: Conductor — User Manual Verification (Protocol in workflow.md) — handed off via TRACK_COMPLETION report (autonomous mode)
Present the verification results to the user. PAUSE for user confirmation before marking the track complete.
- Task 5.7: Mark the track complete + write the TRACK_COMPLETION report
WHERE: docs/reports/TRACK_COMPLETION_twitter_threads_extraction_20260705.md
WHAT: A short report (per the tier2_autonomous_sandbox_20260616 precedent): what was done, files created, tests passing, the standalone-requirement verification, and any deferred items.
HOW: Commit + update conductor/tracks/twitter_threads_extraction_20260705/state.toml to status = "completed".