On Chromium/linux, pressing pause doesn't pause, instead resetting the animation to it's pre-play state - the current attention highlighting disappears. Pressing play again, restarts at the beginning. Having a commonplace "pause pauses, and play resumes" UI, could allow more time to look over state. A youtube-like slow playback 0.25? option might similarly help. Or perhaps even better, buttons for single stepping. Tnx for your work.
Show HN: LLM Attention Visualization (ishamf.dev)
mncharity 9 hours ago
MCP123 12 hours ago
fuddle 15 hours ago
lhk931122 8 hours ago
scottcodie 8 hours ago
talhaanwar 4 hours ago
sva_ 16 hours ago
ifz 16 hours ago
To me, it's more of a neat visualization, not something that can be used to interpret LLM behavior. Even with a lot of simplification, it can show some interesting patterns.
smallmancontrov 16 hours ago
apnabhidu47 15 hours ago
wopak 15 hours ago
are you worried later-layer attention gets drowned out by earlier layers just because there are more of them contributing to the sum?
ifz 15 hours ago
Right now only simple correlations are visible.
asd000hh 6 hours ago
itsnasme 15 hours ago
stared 15 hours ago
I mean, there so many headers and layers, it is tricky to make a choice that will resonate with our intuition . Is it some weighted average? Or maybe ablation test?
ifz 15 hours ago
When I started, I expected I'd have to experiment a lot to find something comprehensible. But this simple computation can already show some patterns.
stared 14 hours ago
visarga 14 hours ago
ex-aws-dude 14 hours ago
acedTrex 14 hours ago
TomatoCo 13 hours ago
ex-aws-dude 13 hours ago
Or does it accumulate the relations like A relates to B, so also add in B's relations
libraryofbabel 12 hours ago
This is incorrect. The compute required per forward pass to generate each additional token during decode will scales as O(N), even with a KV cache (without a KV cache, it would scale as O(N^2)). Over generating N tokens, it's O(N^2) with the cache (and O(N^3) without).
It's O(N) for a forward pass because that new token still has to "attend to" to each previous token. That requires N dot products: between the cached key vectors and the new query vector for the new position. You also have N reads from memory (K and V) which is probably gonna be your actual bottleneck. (Decode is memory-bound.)
This is why you should avoid long contexts, if you can, even with a warm cache. You will get charged more, in "cache read" tokens.
fermlon30000 11 hours ago