Private
Public Access
conductor(cs336_architectures): Phase 3 OCR - 39 frames OCR'd via winsdk in 2.3s
This commit is contained in:
@@ -0,0 +1,820 @@
|
||||
# OCR Results
|
||||
|
||||
## frame_00001.jpg
|
||||
|
||||
```
|
||||
Outline and goals
|
||||
Quick recap of a modern transformer (what you implement)
|
||||
What do most of the large LMs have in common?
|
||||
What are common variations to the architecture / training process?
|
||||
Today's theme: the best way to learn is hands-on experience
|
||||
the second best way is to try to learn from others' experience
|
||||
```
|
||||
|
||||
## frame_00002.jpg
|
||||
|
||||
```
|
||||
Stanf
|
||||
```
|
||||
|
||||
## frame_00004.jpg
|
||||
|
||||
```
|
||||
Architecture variations..
|
||||
Let's think about the core architecture piece
|
||||
High level view:
|
||||
• Dominance of 'LLaMA-
|
||||
like' architectures
|
||||
Trends over the years
|
||||
(QK-norm, Hybrid
|
||||
attention)
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00005.jpg
|
||||
|
||||
```
|
||||
Architecture variations..
|
||||
Let's think about the core architecture piece
|
||||
(17SB)
|
||||
Ant LM
|
||||
(2gogJ
|
||||
BLOOM (17 Sa)
|
||||
OPT
|
||||
paw
|
||||
17B)
|
||||
LL.MA2 (703)
|
||||
LLAMA
|
||||
Q— 2 (72b) -
|
||||
2 11B
|
||||
| 1481
|
||||
OLMO 3 (7B)
|
||||
SmelM7
|
||||
760000
|
||||
50257
|
||||
32000
|
||||
s02S7
|
||||
250680
|
||||
50272
|
||||
256000
|
||||
37000
|
||||
17S6ge
|
||||
32000
|
||||
256128
|
||||
256000
|
||||
152064
|
||||
6s024
|
||||
100362
|
||||
128000
|
||||
100000
|
||||
7 s 6000
|
||||
50304
|
||||
1S2064
|
||||
100000
|
||||
64000
|
||||
100278
|
||||
g36
|
||||
256000
|
||||
2621 "
|
||||
see
|
||||
see
|
||||
see
|
||||
tee
|
||||
2ßgrs
|
||||
'.egrs
|
||||
3. sog
|
||||
2,EBzS
|
||||
2.675
|
||||
2.gS71a2
|
||||
2.6875
|
||||
High level view:
|
||||
Dominance of 'LLaMA-
|
||||
like' architectures
|
||||
Trends over the years
|
||||
(QK-norm, Hybrid
|
||||
attention)
|
||||
```
|
||||
|
||||
## frame_00006.jpg
|
||||
|
||||
```
|
||||
.30
|
||||
28
|
||||
20
|
||||
Pre-vs-post-norm, the data
|
||||
English-Vietnamese development Rt_EU
|
||||
eNorm S.;nieNorxrj FixNorrn • NoW.rTOup
|
||||
PreNoitn Fix Norm
|
||||
PiVNorm
|
||||
— preNonn Layer N Orm
|
||||
PostNorm + LayerNor"i
|
||||
20
|
||||
40
|
||||
60
|
||||
epochs
|
||||
too
|
||||
Epochs
|
||||
(a) Validation (IWSLT)
|
||||
Stern
|
||||
(a) Validation on BERT
|
||||
Salazar and Neyuen 2019
|
||||
13
|
||||
Epochs
|
||||
(b) BLEU (IWSI.T)
|
||||
figure from Xiong 2020
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00007.jpg
|
||||
|
||||
```
|
||||
Pre-vs-post norm, explanations?
|
||||
Gradient attenuation [Xiong 2020]
|
||||
Pre-LN (init)
|
||||
Post-LN (init)
|
||||
Post-LN (after warm-up)
|
||||
0.5
|
||||
0.0
|
||||
Layer
|
||||
(a) WI in the FFN sub-layers
|
||||
Gradient spikes [Salazar and Ngyuen]
|
||||
0
|
||||
3.5
|
||||
3.0
|
||||
2.5
|
||||
2.0
|
||||
1.0
|
||||
0.5
|
||||
0.0
|
||||
—0.5
|
||||
Gradient global norm
|
||||
PostNorm+LayerNorm
|
||||
PreNorm+ScaleNorm+FixNorm+NoWarmup
|
||||
PreNorm+ScaleNorm+FixNorm
|
||||
— PreNorm+LayerNorm
|
||||
200
|
||||
400
|
||||
600
|
||||
800
|
||||
1000
|
||||
1200
|
||||
iteration (x100)
|
||||
Original stated advantage- removing warmup.
|
||||
Today - stability and larger LRS for large networks
|
||||
```
|
||||
|
||||
## frame_00010.jpg
|
||||
|
||||
```
|
||||
New things - 'double' norm or non-residual postnorm
|
||||
If putting LayerNorms in residual streams is bad.. Why not post-norm outside the stream?
|
||||
Recen
|
||||
k, Gemma 2. Olmo 2 only does non-residual post norm
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00011.jpg
|
||||
|
||||
```
|
||||
New things - 'double' norm or non-residual postnorm
|
||||
If putting LayerNorms in residual streams is bad.. Why not post-norm outside the stream?
|
||||
Layer Norm
|
||||
addition
|
||||
FFN
|
||||
Layer Norm
|
||||
addition
|
||||
Multi-Head
|
||||
Attention
|
||||
x
|
||||
Xl+l
|
||||
addition
|
||||
Layer Norm
|
||||
addition
|
||||
Multi-Head
|
||||
Attention
|
||||
Layer Norm
|
||||
x
|
||||
Recent models: Grok, Gemma 2. Olmo 2 only does non-residual post norm
|
||||
```
|
||||
|
||||
## frame_00012.jpg
|
||||
|
||||
```
|
||||
(no text extracted)
|
||||
```
|
||||
|
||||
## frame_00013.jpg
|
||||
|
||||
```
|
||||
RMSNorm - validation
|
||||
RMSNorm runtime (and surprisingly, perf) gains have been seen in papers
|
||||
Model
|
||||
Vanilla Transformer
|
||||
RMS Norm
|
||||
Rezero
|
||||
Rezero + LayerNorm
|
||||
Rezero + RMS Norm
|
||||
Fixup
|
||||
Params
|
||||
223M
|
||||
223M
|
||||
223M
|
||||
223M
|
||||
223M
|
||||
223M
|
||||
Ops
|
||||
11.1T
|
||||
11.1T
|
||||
11.1T
|
||||
11.1T
|
||||
11.1T
|
||||
11.11'
|
||||
Step/ s
|
||||
3.50
|
||||
3.68
|
||||
3.51
|
||||
3.26
|
||||
3.34
|
||||
2.95
|
||||
Early loss
|
||||
2.182 ± 0.005
|
||||
2.167 ± 0.008
|
||||
2.262 ± 0.003
|
||||
2.223 ± 0.006
|
||||
2.221 ± 0.009
|
||||
2.382 ± 0.012
|
||||
Final loss
|
||||
1.838
|
||||
1.821
|
||||
1.939
|
||||
1.858
|
||||
1.875
|
||||
2.067
|
||||
SGLUE
|
||||
71.66
|
||||
75.45
|
||||
61.69
|
||||
70.42
|
||||
70.33
|
||||
58.56
|
||||
XSum
|
||||
17.78
|
||||
17.94
|
||||
15.64
|
||||
17.58
|
||||
17.32
|
||||
14.42
|
||||
WebQ
|
||||
23.02
|
||||
24.07
|
||||
20.90
|
||||
23.02
|
||||
23.02
|
||||
23.02
|
||||
WMT EnDe
|
||||
26.62
|
||||
27.14
|
||||
26.37
|
||||
26.29
|
||||
26.19
|
||||
26.31
|
||||
Narang et al 2020
|
||||
```
|
||||
|
||||
## frame_00014.jpg
|
||||
|
||||
```
|
||||
Summary: architectures
|
||||
Pre-vs-post norm:
|
||||
Everyone does non-residual norm (except
|
||||
OPT350M), likely with good reason,
|
||||
Layer vs RMSnorm:
|
||||
RMSnorm has clear compute wins,
|
||||
sometimes even performance
|
||||
Gating:
|
||||
GLUs are consensus now
|
||||
Serial vs parallel layers:
|
||||
Most models now use serial layers
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00015.jpg
|
||||
|
||||
```
|
||||
Many variations in position embeddings
|
||||
Sine embeddings: add sines and cosines that enable localization
|
||||
Embed(x, i) = Vx + PE
|
||||
pos
|
||||
cos(pos/100002i/dmodel)
|
||||
Absolute embeddings: add a position vector to the embedding
|
||||
Embed(x, i) = Vx + uz
|
||||
Relative embeddings: add a vector to the attention computation
|
||||
xtvvQ(xjvvK + af)T
|
||||
Rope embeddings (next slides..)
|
||||
Notable models:
|
||||
Original transformer
|
||||
Notable models:
|
||||
GPT1/2/3, OPT
|
||||
Notable models:
|
||||
T5, Gopher, Chinchilla
|
||||
Notable models:
|
||||
GPTJ, PaLM, LLaMA
|
||||
Most 2024+ models
|
||||
```
|
||||
|
||||
## frame_00016.jpg
|
||||
|
||||
```
|
||||
RoPE: rotary position embeddings
|
||||
There are many rotations, which one do you pick?
|
||||
Anition • •
|
||||
/ Kev
|
||||
6
|
||||
Position
|
||||
Encodeo / Key
|
||||
et at 20211
|
||||
Just pair up the coordinates and rotate them in 2d (motivation: complex numbers)
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00017.jpg
|
||||
|
||||
```
|
||||
RoPE: rotary position embeddings
|
||||
There are many rotations, which one do you pick?
|
||||
original
|
||||
the
|
||||
query
|
||||
Enhanced
|
||||
Transformer
|
||||
with
|
||||
Rotary
|
||||
Position
|
||||
Embedding
|
||||
Constant
|
||||
Query / Key
|
||||
m
|
||||
Position
|
||||
Query / Key
|
||||
2
|
||||
3
|
||||
4
|
||||
6
|
||||
Position
|
||||
(x'l, X'2)
|
||||
Position Encoded Query / Key
|
||||
Position Encoded Query / Key
|
||||
the
|
||||
(rota
|
||||
[Su et al 2021]
|
||||
some
|
||||
inrormation
|
||||
only rotate high frequency palrs
|
||||
partially
|
||||
rotated query
|
||||
only
|
||||
information
|
||||
Gemma 4 alternative: just first 2
|
||||
Just pair up the coordinates and rotate them in 2d (motivation: complex numbers)
|
||||
```
|
||||
|
||||
## frame_00018.jpg
|
||||
|
||||
```
|
||||
Implementation and code for RoPE
|
||||
Usual
|
||||
attention stuff
|
||||
Get the RoPE
|
||||
matrix cos/sin
|
||||
Multiply
|
||||
query/key inputs
|
||||
query_states
|
||||
key. states
|
||||
value_states
|
||||
Flash attention requires the input to have the Shape
|
||||
seq_tength head din hldden_dim
|
||||
g therefore just need to keep the original shape
|
||||
query_states • q_ien, self 2)
|
||||
key _ states key _ q_ien. Self 2)
|
||||
value_states Q_len. 2)
|
||||
COS, sin • S*".
|
||||
query States. key_states
|
||||
cos. Sin)
|
||||
Same stuff as the usual multi.head self attention below
|
||||
te: embedding at each attention operation to enforce position invariance
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00019.jpg
|
||||
|
||||
```
|
||||
Hyperparameters
|
||||
Transformer hyperparameter questions you might have had in 224n..
|
||||
How much bigger should the feedforward size be compared to hidden size?
|
||||
How many heads, and should num_heads always divide hidden size?
|
||||
What should my vocab size be?
|
||||
And other model setting questions
|
||||
Do people even regularize these huge LMs?
|
||||
How do people scale these models - very deep or very wide?
|
||||
```
|
||||
|
||||
## frame_00020.jpg
|
||||
|
||||
```
|
||||
Stanf
|
||||
```
|
||||
|
||||
## frame_00022.jpg
|
||||
|
||||
```
|
||||
Stanf
|
||||
```
|
||||
|
||||
## frame_00023.jpg
|
||||
|
||||
```
|
||||
RoPE: rotary position embeddings
|
||||
High level thought process: a relative position embedding should be some f (x, i) s.t.
|
||||
That is, the attention function only gets to depend on the relative position (i-j). How do
|
||||
existing embeddings not fulfill this goal?
|
||||
Sine: Has various cross-terms that are not relative
|
||||
(Embed(x, i), Embed(y, i)) = (tax, vy) + (PEi, vy) ...
|
||||
Absolute: obviously not relative
|
||||
Relative embeddings:
|
||||
TivvQ(xjvvK + af)T
|
||||
is not an inner product
|
||||
```
|
||||
|
||||
## frame_00024.jpg
|
||||
|
||||
```
|
||||
Stanf
|
||||
```
|
||||
|
||||
## frame_00025.jpg
|
||||
|
||||
```
|
||||
RoPE: rotary position embeddings
|
||||
High level thought process: a relative position embedding should be some f (x, i) s.t.
|
||||
That is, the attention function only gets to depend on the relative position (i-j). How do
|
||||
ting embeddings not fulfill this goal?
|
||||
Sine: Has various cross-terms that are not relative
|
||||
Vx, vy) + (PEi, vy)
|
||||
(Embed(x, i), Embed(y, t)) (
|
||||
Absolute: obviously not relative
|
||||
Relative embeddings:
|
||||
is not an inner product
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00028.jpg
|
||||
|
||||
```
|
||||
Exception #1- GLU variants
|
||||
Remember that GLU variants scale down by 2/3"'. This means most GLU variants have
|
||||
This is mostly what happens. Some notable such examples.
|
||||
model •
|
||||
Model
|
||||
PaLM
|
||||
Mistral 7B
|
||||
LLaMA-2 70B
|
||||
LLaMA 70B
|
||||
Qwen 14B
|
||||
Deepseek 67B
|
||||
Yi 34B
|
||||
T5v1.1
|
||||
dff/dmodel
|
||||
3.5
|
||||
3.5
|
||||
2.68
|
||||
2.67
|
||||
2.68
|
||||
2.85
|
||||
2.5
|
||||
odels are roughly in this range, though PaLM, LLaMA2 and Mistral are slightly larger
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00029.jpg
|
||||
|
||||
```
|
||||
Exception #2 - T5
|
||||
As we have (and will) see, most LMs are have boring, conservative hyperparameters.
|
||||
One exception is T 5 [Raffel et al 2020] which has some very bold settings.
|
||||
In particular, for the 11B model, they set
|
||||
= 65,536
|
||||
d
|
||||
= 1024
|
||||
model
|
||||
For an astounding 64-times multiplier.
|
||||
for "11B" we use dff = 65,536 with 128-headed attention
|
||||
producing a model with about 11 billion parameters. We chose to scale up dff
|
||||
specifically because modern accelerators (such as the TPUs we train our models
|
||||
on) are most effcient for large dense matrix multiplications like those in the
|
||||
rlhansformer's feed-forward networks.
|
||||
Other, recent exceptions - Gemma 2 (8x), SmolLM/Gemma 3/Gemma 4 (4x, GLU)
|
||||
```
|
||||
|
||||
## frame_00030.jpg
|
||||
|
||||
```
|
||||
Evidence on aspect ratio scaling
|
||||
50M Params
|
||||
274M Params
|
||||
1-5B Params
|
||||
A wide range of architectures
|
||||
achieve similar performance
|
||||
Aspec
|
||||
102
|
||||
iO / mayor)
|
||||
[Kaplan et al 20201
|
||||
[Tay et al 20211
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00031.jpg
|
||||
|
||||
```
|
||||
What are typical vocabulary sizes?
|
||||
Monolingual models - 30-50k vocab
|
||||
Multilingual / production systems 100-250k
|
||||
Model
|
||||
Original
|
||||
transformer
|
||||
GPT
|
||||
GPT2/3
|
||||
T5/T5v1.1
|
||||
LLaMA
|
||||
Token count
|
||||
37000
|
||||
40257
|
||||
50257
|
||||
32128
|
||||
32000
|
||||
Model
|
||||
mT5
|
||||
PaLM
|
||||
GPT4
|
||||
Gemma 4
|
||||
DeepSeek
|
||||
Qwen 15B
|
||||
Yi
|
||||
Token count
|
||||
250000
|
||||
256000
|
||||
100276
|
||||
262144
|
||||
100000
|
||||
152064
|
||||
64000
|
||||
Monolingual vocabs don't need to be huge, but multilingual ones do
|
||||
```
|
||||
|
||||
## frame_00032.jpg
|
||||
|
||||
```
|
||||
What are typical vocabulary sizes?
|
||||
Monolingual models — 30-50k vocab
|
||||
Multilingual / production systems 100-250k
|
||||
Modet
|
||||
Original
|
||||
transformer
|
||||
GPT
|
||||
GPT2/3
|
||||
T5/TSv1.1
|
||||
LLaMA
|
||||
Token count
|
||||
37000
|
||||
40257
|
||||
50257
|
||||
32128
|
||||
Model
|
||||
CmT5
|
||||
PaUM
|
||||
GPT4
|
||||
Gemma 4
|
||||
OeepSeek
|
||||
Qwen ISB
|
||||
Token count
|
||||
250000
|
||||
256000
|
||||
100276
|
||||
262144
|
||||
100000
|
||||
152064
|
||||
64000
|
||||
abs don't need to be huge, but multilingual ones do
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00034.jpg
|
||||
|
||||
```
|
||||
Stanfo
|
||||
```
|
||||
|
||||
## frame_00035.jpg
|
||||
|
||||
```
|
||||
Dropout and other regularization
|
||||
Do we need regularization during pretraining?
|
||||
Arguments against:
|
||||
There is a lot of data (trillions of tokens), more than parameters.
|
||||
SGD only does a single pass on a corpus (hard to memorize)
|
||||
This is all quite reasonable.. but what do people do in practice?
|
||||
```
|
||||
|
||||
## frame_00037.jpg
|
||||
|
||||
```
|
||||
E Aspect ret*'
|
||||
Summary: hyperparameters
|
||||
Feedforward
|
||||
Factor-of-4 rule of thumb (8/3 for GLUs) is
|
||||
standard (with some evidence)
|
||||
Head dim
|
||||
Head dim*Num head = D model is standard
|
||||
- but low to no validation
|
||||
Aspect ratio
|
||||
Wide range of 'good' values (100-200).
|
||||
Systems concerns dictate the value
|
||||
Regularization
|
||||
You still 'regularize' LMs but its effects are
|
||||
primarily on optimization dynamics
|
||||
Aa Name
|
||||
transtorrner
|
||||
TS(XXL 11B) vl.l
|
||||
Anthropic LM eaude)
|
||||
(280B)
|
||||
(176B)
|
||||
OPT (175B)
|
||||
paw (5408)
|
||||
Bakhum 2
|
||||
Mistral (7B)
|
||||
LLaMA2 (70B)
|
||||
LLaMA (65B)
|
||||
2 (27B)
|
||||
Netmum-a (340B)
|
||||
Qwen 2 (72b) - 2.5
|
||||
Fakon2 11B
|
||||
Phi3 (mdl) - s.ne 'or
|
||||
3 (70B)
|
||||
Convnmd
|
||||
Qwen (143)
|
||||
Deepseek (67B)
|
||||
Yi (34B)
|
||||
Mixtrai of Experts
|
||||
Genna 3
|
||||
SmolLM2 (1.7B)
|
||||
2017
|
||||
201B
|
||||
2019
|
||||
26875
|
||||
2.6875
|
||||
2.75
|
||||
2.6875
|
||||
2.675
|
||||
2.6875
|
||||
2857142
|
||||
171
|
||||
001
|
||||
Stafifordo
|
||||
```
|
||||
|
||||
## frame_00038.jpg
|
||||
|
||||
```
|
||||
Stanfo
|
||||
```
|
||||
|
||||
## frame_00040.jpg
|
||||
|
||||
```
|
||||
Stability tricks
|
||||
Recently, lots of attention on stable training
|
||||
00000
|
||||
1 't train models that look like the blue curve!
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00041.jpg
|
||||
|
||||
```
|
||||
Where do the issues arise? Beware of softmaxes!
|
||||
Softmaxes - can be ill-behaved due to exponentials / divison by zero
|
||||
Output
|
||||
Probabilities
|
||||
Softmax
|
||||
Linear
|
||||
(Output Embedding)
|
||||
Norm
|
||||
Transformer Block
|
||||
Transformer Block
|
||||
Token
|
||||
Embedding
|
||||
Inputs
|
||||
(ba t ch
|
||||
(ba Ch
|
||||
Output tensor with shape
|
||||
size, seg_len, d model)
|
||||
Add
|
||||
Position-Wise
|
||||
Feed-Forward
|
||||
Norm
|
||||
Add
|
||||
Causal Multi-Head
|
||||
Self-Attention
|
||||
w/ RoPE
|
||||
Norm
|
||||
Input tensor with shape
|
||||
size, seg_len, d model)
|
||||
```
|
||||
|
||||
## frame_00042.jpg
|
||||
|
||||
```
|
||||
Where do the issues arise? Beware of softmaxes!
|
||||
Softmaxes — can be ill-behaved due to exponentials / divison by zero
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00043.jpg
|
||||
|
||||
```
|
||||
Output softmax stability - the 'z-loss'
|
||||
Recall the softmax calculation
|
||||
log(P@)) —
|
||||
log
|
||||
Ur(c) -
|
||||
E - - 0)2]
|
||||
— a log2 (Z@i))]
|
||||
[From Devlin 2014]
|
||||
This is useful for stability! PaLM used this 'z loss' trick.
|
||||
We additionally use an auxiliary loss of z -loss = 10
|
||||
• log Z to encourage the softmax
|
||||
normalizer log(Z) to be close to 0, which we found increases the stability of training.
|
||||
Other examples: Baichuan 2 (2023), DCLM (2024), OLMo 2 (2025), OLMo 3 (2025)
|
||||
```
|
||||
|
||||
## frame_00044.jpg
|
||||
|
||||
```
|
||||
Output softmax stability - the 'z-loss'
|
||||
Recall the softmax calculation
|
||||
log
|
||||
Ur(x) -
|
||||
IVI
|
||||
— — 0)2J
|
||||
[From Devlin 2014)
|
||||
u
|
||||
or stability! PaLM used this 'z loss' trick.
|
||||
We additionally use an auxiliary loss of to encourage the softmax
|
||||
normalizer log(Z) to be to O. which wc found increases tbc Stability of training.
|
||||
xamples: Baichuan 2 (2023), DCLM (2024), OLMo 2 (2025), OLMo 3 (2025)
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00046.jpg
|
||||
|
||||
```
|
||||
GQA/MQA - Reducing attention head cost
|
||||
What about the incremental case when we generate text?
|
||||
Key difference: can't parallelize the generation process — needs to be step by step
|
||||
In this case - we need to incrementally re-compute/update attention via the 'KV cache'
|
||||
Step
|
||||
(Animation from https://mrdium corn/@josolages/kwexhingoplsined 276.52020249)
|
||||
Stanford
|
||||
```
|
||||
|
||||
## frame_00047.jpg
|
||||
|
||||
```
|
||||
GQA/MQA - Reducing attention head cost
|
||||
What's the incremental arithmetic intensity?
|
||||
projection
|
||||
Total arithmetric operations (bnd2), total memory accesses (bn2d + nd2)
|
||||
-1
|
||||
Arithmetic intensity is not good O L +
|
||||
- need large batches + short seq length
|
||||
(n) or big model dimensions (d)
|
||||
Is there some way around this? The n/d term is difficult to reduce.
|
||||
```
|
||||
|
||||
## frame_00048.jpg
|
||||
|
||||
```
|
||||
Stanfo
|
||||
```
|
||||
|
||||
## frame_00049.jpg
|
||||
|
||||
```
|
||||
Sparse / sliding window attention
|
||||
Attending to the entire context can be expensive (quadratic).
|
||||
Build sparse / structured attention that trades off expressiveness vs runtime (GPT3, GPT-
|
||||
OSS, Gemma4)
|
||||
(a) Transformer
|
||||
(b) Sparse Transformer (strided)
|
||||
(c) Sparse Tyansformer (fixed)
|
||||
[Child et al 20191
|
||||
```
|
||||
@@ -0,0 +1,2 @@
|
||||
Phase 2 Keyframes for C:\projects\manual_slop\conductor\tracks\video_analysis_cs336_architectures_20260621\artifacts\video.mp4
|
||||
OK: kept 39 frames
|
||||
@@ -0,0 +1,2 @@
|
||||
Phase 3 OCR for C:\projects\manual_slop\conductor\tracks\video_analysis_cs336_architectures_20260621\artifacts\frames (winsdk)
|
||||
OK: OCR'd 39 frames in 2.3s
|
||||
Reference in New Issue
Block a user