From 186a71321a037d993863f7e55ecfc26aa53756bb Mon Sep 17 00:00:00 2001 From: Ed_ Date: Sun, 26 Jul 2026 21:36:15 -0400 Subject: [PATCH] docs(twitter): add 2078729662021111958 corpus (NOTimothyLottes 64KiB aligned jump-window + @noop_dev exchange) 11-post thread. Root: 4-byte overhead interpreter with 64KiB aligned window of directly-jumpable words (write to ax doesn't change other 48 bits), cuts source size in half. Tangent with @noop_dev covering cold-cache misses, runtime-macroassembler idea, 4K Atari 2600 emu precedent. End: GPU code generation for AMD where 8+ bitfields in an opcode means the simple interpreter won't work. --- .../media/2078728600912630031_1.png | Bin 0 -> 3604 bytes docs/twitter/2078729662021111958/thread.md | 61 ++++++ .../2078729662021111958/thread_data.json | 184 ++++++++++++++++++ 3 files changed, 245 insertions(+) create mode 100644 docs/twitter/2078729662021111958/media/2078728600912630031_1.png create mode 100644 docs/twitter/2078729662021111958/thread.md create mode 100644 docs/twitter/2078729662021111958/thread_data.json diff --git a/docs/twitter/2078729662021111958/media/2078728600912630031_1.png b/docs/twitter/2078729662021111958/media/2078728600912630031_1.png new file mode 100644 index 0000000000000000000000000000000000000000..7946820c0b54afef2136b476c34628b7bee30040 GIT binary patch literal 3604 zcmaJ@c{CJ`6R(gG`HCo#%6%-k&!yZex5YYgUyCJ2$XQBqBrMjoi!73xMV7UWP_}jK zMy#uxx$o6F^7H%i_ulvY=8t(ZpEvV)GjHC^n*>v1Jr*WzrgP`cv4Hfo&Ci{?06nua zFVUT8IVWVr8L($+2-Z10J)NJQKLh{y=N}~{rOeFCzkmOFd3g;D4G9PcaBy%udGh4* z=g*feU3&1~fr^TXsj2DC&d!}XcZfuykdTmpfdPd=>FDU#+uPgP+T!Em`}*~3YHI4i z!NJMN$>!$fty{NzeSLX&c$k=&L`6j-BO|xBxA*t=BO)SLSy``LyCx$elai8hczF2X z!v}qReJ(C8c6RoPiV6aOU}$J~;lc$51_nh%#k{;c6B83|Zf+wZqlt-$u&}VJSFb87 zE62yjA0HoQWo5m7{rbj@8&|Gefj}VV&!7ME=g-HF9{~V>wzl@h#s)t>zo4L?uCA`9 zr>Co{tGc>63X4(9qCWUS6i7qm!1F#$vG!4h|9$6653J z;^N{LFJ7e4XtJ`hWHLE1G11S@Z+(5eva)h+Ztnj5`*n47w{PEGSXe-#(J?VGKYsim zkx0LO{Ypzq)6&wKoSbxYbj0CsU%q_#`ST|m8{5>>R99D5c6RpC(h?8|Tw7b~?d|pU z_AV|i9vB$N$;o;1<_(ofWnp3A<>hsAa|;d*uC1+ITwLtz>?|xSghHV=Z{BQfZeCqo zm6MYT4-Y>&I+B-{my(jo&(DvJj=p~Vx}Kiiy?ggQefp%PruOjR!`a!{>FMe8^z^2t zrf1KdJ%0RncXu}_De3OryE8L0-QC^7!oq`tgQcaVVq#)GK0dj*xjH&Jt*xyiBO~SI zfwQ+SG1IIyLa!RqN3W{+fgXg+qZ8=M@L&)TFT1G`uh5q znVFN5lY4r4N=iy%V`JOe+AJ+C8yXslii$WnIaO6v8yg#4TwLIAczu2S%a<=NU%vb& zG~VglIi?Aa_Cs)3{`ypeJ=mR_aVwq+&A`^fQ$JKV$t-k@rOnCtg9u?nu6tNlnEtzp zcW$mtsf$Oc7bWp|$@>uxF)yx1tuHeHVdbOs`}I~6xb=EitKHFH`=+vN&_HL)1HLD( zXi=xx&j!}_fu7O%gfB-gE}iY!ze^$ima~i_-uSP}4imN23Z{O?v`V$1&YLjtbW0<2Oon!|P5U1c-TC(gbG zqgx_b#6e@_772eyG&hePc04!T&#irYu6UgO8i=rqgU%9NuL$M%F*`}CIh^1A`Z7Lv zy-P?~43rm`WOHm`F)pV`YD*G=a)wl-W4e}c35kg|mxq}Z;8Ho)Gt;W`szm)==>p&Q zA;x^uFv>v*8&1Br@rjFdG&QKiit&4fYj$3^a7mZY?TcIutqIk3f$x^6Jd*N8qc3Sm(uXg-jgHApJD2>oo{Ka1m6tW#SnUlNL@Szyh#L$YIBry${K% zfiYoJ4Kq4~;UDT$#)enBMiEB+YX|lnx3+qX3hLUSr25$M+S) zyi@>SIDJP^&;mqR%p=f{yc2<<2W*mz-lQ%_#(qMfi<>)>tEJt^Gum(0WI@=3U<>zaX6j9VA=W@tfaWk>(MpZ|(hefXVo*uPJ--Tf5_ z)h`grbY)>}JjJuA_RdCp89t^q_*aKx{aOQBb^%d5nFpT8TQ6rkx>z`@Z6=bhTO%Ap z%@zy;BWQB=&?k`@cCeWgcSyC6%B1XH+xN7-h7}VIHii*5h(uOTNt=ODBgD{uxhWN~ zsE|MJUf2`38A8aHA#wtim$8D?;Z3uWb_?@2jyYR{*b_$atI<8_+W}mb#&@B_R0l)1 ztun@Geid-cn?w%@nSW?*yMjlK^HQ5q{I)@Jj=_0vN^*`0<}9=Fvn<>x0NWvpA0UfH z7m|0+%vC};c;)-A0qcBPXHO>5`H{Qx0pmhos;JM>6t5UGzG=jcNmILIVS>gCQ$^d7^ zMmsO1werQ?Mj?Qj_~Crl4mJq8KRseEYbOm=o$vLTKjcU<2*u6@;otzejbZDL<<1YPh<$cs|&#t>kQduLl!_i+Ad z?oPcLE;u_qk60i8mxl(Ah^}+GKnr9?+aTokeno+k7Q`hVjWWxn7z%!G(}?<^s0yyz zjQu2GKoi6a4i;AIsUPLz!3jdMv#O?~rl`9Pp6%sRiZO(xCC{|t2!6)-InCFvkBE*d zTx}$!l?v~I*r>hzh#gE(;?g2dxo3a#EPaW9(X+;1k18&Bh{b47P;y6{V!}X64M%yX zqv$4fpCSO=8Ii3gog|K&{*;PeS`};k6nofJne9(A(r!`1y{8kbGL1gjh?r{cVTr<7 zZ?pR4)$(!qZD>i!+t^a)|KuwUFt;KHRtBI6dNE))=aI=eXMZNi09ac#mEl^?{?!@HaVJ!XHSQ)HM!lk{3gOkHq!|p#5pjNbrtb(lJEqa+ad$jkGY^d2kplMdaUS~!I3R02 zmL$GGzy?v*7Z|GsQ7i>3l%QO>XIC=kP~Pe@N$I5%#Z4gFVWEa){*uZlf$^Tic+MO+ zmt47TJGCLBuRT-c9pXp!x1?x&^(T(2`$PPH0ier3L3r4XYJn7peLjN!sGa4ouD1Hf zC$t5s6=Xk4Wj92dr;QQiV^sS2N76CUUD~n3vr~-^0+48 zUd&C4SvP2d-rLwvo^ZqHAnEA_#|ODBdI6CgoKutIBl|tOG9V3a%wrz+I*9pgVopPE6tWyXF0sB?10u4W*!|XKi#l>KjtR+LtgXB9~*;6+vGpP zZx?(vv$oNKhkkNvhLS2lFGTJ^8#Us7@IpcCO=cO@eG%bE-G+4$?2XN35or3cUB9dP z#&b}ySoH1F=UJ6=AUIGGP+A{|Of&Jxv}~Cf@L<2rky0Xe~YC6pu(1>;IVs`rB)&R6uNX304}%I)4pIfyA)r=vC9 zJ3c0Dquk}UGmAWCxKB)%qiDK#azR4PnA}+&Fb(5+I-z~yk|hO#=&;%0SDR+FvnPn) z6%m8@IpBE?JZjiT*LUAkqwwVf?{1tQ_t~Sjg#)_}vmT$Qdsit>YT?fR&*1-$y@7nohC*0m`)oMF0Q* literal 0 HcmV?d00001 diff --git a/docs/twitter/2078729662021111958/thread.md b/docs/twitter/2078729662021111958/thread.md new file mode 100644 index 00000000..34717796 --- /dev/null +++ b/docs/twitter/2078729662021111958/thread.md @@ -0,0 +1,61 @@ +--- +title: "The aim of course, keep the stuff that doesn't need to go fast optimized instead" +author: "NOTimothyLottes" +handle: "@NOTimothyLottes" +post_url: "https://x.com/NOTimothyLottes/status/2078729662021111958" +post_id: "2078729662021111958" +timestamp: "2026-07-19 06:32:27" +post_count: 11 +reply_count: 1 +repost_count: 0 +like_count: 3 +view_count: 509 +--- + +# @NOTimothyLottes — The aim of course, keep the stuff that doesn't need to go fast optimized instead + +## Post 1 (2026-07-19 06:28:14) + +Probably shouldn't consider this -BUT- apparently writing to ax doesn't change the other 48-bits. So one could do a 4-byte overhead interpreter with a 64KiB aligned window of directly jumpable words like this below. [rsi]=addresses to jump to. You'd pay the false dependency stall + +![Media 1](./media/2078728600912630031_1.png) + +## Post 2 (2026-07-19 06:29:54) — reply to Post 1 + +It's interesting because it cuts interpreted source size in half. Ie a stream of 16-bit offsets instead of 32-bit addresses. + +## Post 3 (2026-07-19 06:32:27) — reply to Post 2 + +The aim of course, keep the stuff that doesn't need to go fast optimized instead for low complexity and low size (aka interpreted forth), and keep the stuff that needs to go fast, at peak, assembly. Hits 2 extremes well. + +## Post 4 (2026-07-19 08:46:33) — reply to Post 3 + +@NOTimothyLottes These instrs are anything but fast and if you want both "fast" and compact you run a separate decoding pass where you expand custom bytecode into target instructions that suck less. + +## Post 5 (2026-07-19 14:38:22) — reply to Post 4 + +@noop_dev Basic truth: if the program is tiny and dependency free (ie intrinsically not just glue for libraries) it’s probably already fast even if the machine isn’t. And for the things that need perf there is always assembly + +## Post 6 (2026-07-19 14:43:27) — reply to Post 5 + +@NOTimothyLottes Anyway, I am trying you to sell the idea of bytecode-driven macroassembler that expands "macros" on load. Like I did with 4K atari 2600 emu ~20 years ago.. some demosceners I knew also adopted the idea.. + +## Post 7 (2026-07-19 14:55:41) — reply to Post 6 + +@noop_dev Many of my other systems had been such that interpreted source generates raw code then executes the code (all at once after code is fully processed). Effectively forth as a macro language and instruction generator, runtime assembler. + +## Post 8 (2026-07-19 15:04:11) — reply to Post 7 + +@noop_dev But for cold cache stuff it’s easy to burn more time in binary generation than it would cost to do simple interpreter. + +## Post 9 (2026-07-19 15:10:25) — reply to Post 8 + +@NOTimothyLottes Not sure I can understand how cold cache matters for a linear transformation of a relatively small # of bytes. + +## Post 10 (2026-07-19 15:16:20) — reply to Post 9 + +@noop_dev If it’s a byte that indexes into say a fixed size physical instruction, yeah it’s just a decompression of source code, sure easy to do fast. Im talking more the multi-branch miss per instruction stuff. + +## Post 11 (2026-07-19 15:20:43) — reply to Post 10 + +@noop_dev Meaning “mov edi,[rax-0x32]” could be generated from 4 symbol lookups (2 for regs, one for offset, one for instruction). Also my intent here is for GPU code generation for AMD GPUs where youd have sometimes 8+ bitfields in an opcode. So the simple stuff won’t work there … diff --git a/docs/twitter/2078729662021111958/thread_data.json b/docs/twitter/2078729662021111958/thread_data.json new file mode 100644 index 00000000..48c519cc --- /dev/null +++ b/docs/twitter/2078729662021111958/thread_data.json @@ -0,0 +1,184 @@ +{ + "root_post_id": "2078729662021111958", + "posts": [ + { + "post_id": "2078728600912630031", + "author": "NOTimothyLottes", + "handle": "NOTimothyLottes", + "text": "Probably shouldn't consider this -BUT- apparently writing to ax doesn't change the other 48-bits. So one could do a 4-byte overhead interpreter with a 64KiB aligned window of directly jumpable words like this below. [rsi]=addresses to jump to. You'd pay the false dependency stall", + "timestamp": "2026-07-19 06:28:14", + "media_urls": [ + "https://pbs.twimg.com/media/HNkf6ntX0AA8d1M?format=png&name=orig" + ], + "reply_to_id": null, + "quote_of_id": null, + "metrics": { + "reply_count": 4, + "repost_count": 1, + "like_count": 14, + "view_count": 1564 + } + }, + { + "post_id": "2078729020900823089", + "author": "NOTimothyLottes", + "handle": "NOTimothyLottes", + "text": "It's interesting because it cuts interpreted source size in half. Ie a stream of 16-bit offsets instead of 32-bit addresses.", + "timestamp": "2026-07-19 06:29:54", + "media_urls": [], + "reply_to_id": "2078728600912630031", + "quote_of_id": null, + "metrics": { + "reply_count": 1, + "repost_count": 0, + "like_count": 2, + "view_count": 586 + } + }, + { + "post_id": "2078729662021111958", + "author": "NOTimothyLottes", + "handle": "NOTimothyLottes", + "text": "The aim of course, keep the stuff that doesn't need to go fast optimized instead for low complexity and low size (aka interpreted forth), and keep the stuff that needs to go fast, at peak, assembly. Hits 2 extremes well.", + "timestamp": "2026-07-19 06:32:27", + "media_urls": [], + "reply_to_id": "2078729020900823089", + "quote_of_id": null, + "metrics": { + "reply_count": 1, + "repost_count": 0, + "like_count": 3, + "view_count": 509 + } + }, + { + "post_id": "2078763409403683061", + "author": "Boris Chuprin", + "handle": "noop_dev", + "text": "@NOTimothyLottes These instrs are anything but fast and if you want both \"fast\" and compact you run a separate decoding pass where you expand custom bytecode into target instructions that suck less.", + "timestamp": "2026-07-19 08:46:33", + "media_urls": [], + "reply_to_id": "2078729662021111958", + "quote_of_id": null, + "metrics": { + "reply_count": 1, + "repost_count": 0, + "like_count": 0, + "view_count": 28 + } + }, + { + "post_id": "2078851948543910353", + "author": "NOTimothyLottes", + "handle": "NOTimothyLottes", + "text": "@noop_dev Basic truth: if the program is tiny and dependency free (ie intrinsically not just glue for libraries) it’s probably already fast even if the machine isn’t. And for the things that need perf there is always assembly", + "timestamp": "2026-07-19 14:38:22", + "media_urls": [], + "reply_to_id": "2078763409403683061", + "quote_of_id": null, + "metrics": { + "reply_count": 1, + "repost_count": 0, + "like_count": 1, + "view_count": 60 + } + }, + { + "post_id": "2078853226967781591", + "author": "Boris Chuprin", + "handle": "noop_dev", + "text": "@NOTimothyLottes Anyway, I am trying you to sell the idea of bytecode-driven macroassembler that expands \"macros\" on load. Like I did with 4K atari 2600 emu ~20 years ago.. some demosceners I knew also adopted the idea..", + "timestamp": "2026-07-19 14:43:27", + "media_urls": [], + "reply_to_id": "2078851948543910353", + "quote_of_id": null, + "metrics": { + "reply_count": 2, + "repost_count": 0, + "like_count": 0, + "view_count": 59 + } + }, + { + "post_id": "2078856306811691388", + "author": "NOTimothyLottes", + "handle": "NOTimothyLottes", + "text": "@noop_dev Many of my other systems had been such that interpreted source generates raw code then executes the code (all at once after code is fully processed). Effectively forth as a macro language and instruction generator, runtime assembler.", + "timestamp": "2026-07-19 14:55:41", + "media_urls": [], + "reply_to_id": "2078853226967781591", + "quote_of_id": null, + "metrics": { + "reply_count": 1, + "repost_count": 0, + "like_count": 2, + "view_count": 82 + } + }, + { + "post_id": "2078858446409978111", + "author": "NOTimothyLottes", + "handle": "NOTimothyLottes", + "text": "@noop_dev But for cold cache stuff it’s easy to burn more time in binary generation than it would cost to do simple interpreter.", + "timestamp": "2026-07-19 15:04:11", + "media_urls": [], + "reply_to_id": "2078856306811691388", + "quote_of_id": null, + "metrics": { + "reply_count": 1, + "repost_count": 0, + "like_count": 2, + "view_count": 81 + } + }, + { + "post_id": "2078860015323038149", + "author": "Boris Chuprin", + "handle": "noop_dev", + "text": "@NOTimothyLottes Not sure I can understand how cold cache matters for a linear transformation of a relatively small # of bytes.", + "timestamp": "2026-07-19 15:10:25", + "media_urls": [], + "reply_to_id": "2078858446409978111", + "quote_of_id": null, + "metrics": { + "reply_count": 1, + "repost_count": 0, + "like_count": 0, + "view_count": 30 + } + }, + { + "post_id": "2078861504284074431", + "author": "NOTimothyLottes", + "handle": "NOTimothyLottes", + "text": "@noop_dev If it’s a byte that indexes into say a fixed size physical instruction, yeah it’s just a decompression of source code, sure easy to do fast. Im talking more the multi-branch miss per instruction stuff.", + "timestamp": "2026-07-19 15:16:20", + "media_urls": [], + "reply_to_id": "2078860015323038149", + "quote_of_id": null, + "metrics": { + "reply_count": 1, + "repost_count": 0, + "like_count": 0, + "view_count": 207 + } + }, + { + "post_id": "2078862606517895324", + "author": "NOTimothyLottes", + "handle": "NOTimothyLottes", + "text": "@noop_dev Meaning “mov edi,[rax-0x32]” could be generated from 4 symbol lookups (2 for regs, one for offset, one for instruction). Also my intent here is for GPU code generation for AMD GPUs where youd have sometimes 8+ bitfields in an opcode. So the simple stuff won’t work there …", + "timestamp": "2026-07-19 15:20:43", + "media_urls": [], + "reply_to_id": "2078861504284074431", + "quote_of_id": null, + "metrics": { + "reply_count": 0, + "repost_count": 0, + "like_count": 2, + "view_count": 208 + } + } + ], + "source_url": "https://x.com/NOTimothyLottes/status/2078729662021111958" +} \ No newline at end of file