Extracted from the tensor-formula-viz skill and rebuilt around the idea that the geometry rules should be enforced by construction rather than restated as prose an agent has to remember. - assets/supertensor.sty: faces, stacks, index faces, shared caption lanes, meaning box, signature. Macros take a declared axis and a declared role, so equal shapes get equal edges, a x a is square, a transpose swaps the face, and contracted axes share an edge length -- without any manual alignment. - scripts/preflight.sh: decide the TikZ/CJK path before drawing. - scripts/build.sh: compile and fail on silent corruption (missing CJK glyphs, overfull boxes, undeclared roles), then export pdf/svg/png/thumb. - scripts/test.sh: build every figure as a regression test for the package. - examples/: three golden figures (TP-FFN, causal MHA, MoE top-k gather) plus an anti-pattern gallery of figures that compile cleanly and still lie. - SKILL.md + references/: lean entry point, details loaded on demand.
3.8 KiB
Geometry invariants
A figure that is drawn to the wrong geometry is not a stylistic problem — it teaches the reader a false fact about the computation. These invariants are mandatory. Most of them are automatic if you declare the geometry ledger and never pass a raw number.
The ledger
\stdim{T}{6} % sequence length -> 6 cells, everywhere in the figure
\stdim{d}{9} % model width -> 9 cells, everywhere
\stdim{dh}{3} % head width -> 3 cells, and 3*dh = d holds visually
\stface{...}{T}{dh} resolves the names through the ledger, so one symbolic axis maps
to exactly one physical edge length across the whole figure. Equal shapes therefore form
an equivalence class automatically: Q and V at T×d_h come out identical without you
lining anything up by hand.
Raw integers are accepted (\stface{A}{(0,0)}{4}{4}) but they opt out of the guarantee.
Use them only for a face whose axis appears nowhere else.
The rules
-
Face orientation. A matrix face
a×bis heighta, widthb. Always. For batched or stacked tensors, the last two axes make the face; leading axes become depth (\ststack) or repeated panels — never a wider rectangle. -
Squares.
a×arenders as a square. Automatic when both arguments resolve to the same declared axis. -
Transpose. Draw
K^Tby physically swapping height and width:\ststack{KT}{...}{dh}{T}{3}against\ststack{K}{...}{T}{dh}{3}. Relabelling a faceK^Twhile leaving its shape alone is invalid — it is the single most common lie in attention figures. -
Contraction. In
(m×k)(k×n), both occurrences ofkget the same edge length. With the ledger this is free: pass the same axis name to the width of the left face and the height of the right one. The same applies toeinsumand attention axes. -
Partition. Explicit shards tile their parent exactly along the split axis, equal shards are equal in size, and concatenation reverses the split. Place shards from the previous face's edge so no gap can creep in:
\stface[role=r1]{W1a}{(...)}{d}{dffl} \stface[role=r2]{W1b}{($(W1a.east)+(2*\stunit,0)$)}{d}{dffl}The offset is
half-width of the next facein\stunit, so the two faces are exactly adjacent. Size concatenated parts from their declared shapes — Q/K/V are equal segments only when their output shapes are equal. -
Elision. If intermediate shards are omitted, draw an ellipsis. Never stretch the visible shards to impersonate the full parent.
-
Axis changes. Geometry may change only at an explicit reshape, flatten, transpose, split, or concat operator, and the axis identity must be stated — e.g.
(h/p)·d_h = d/p. Do not silently reuse one generic rectangle on both sides of an axis change. -
Illustrative counts. Cell counts need not equal real dimensions. Choosing
d=9to stand for 4096 is fine. It never waives rules 1–7: the ratios you draw are read as facts. Ifd = h·d_h, pick numbers where that arithmetic actually holds.
Sharded matmul
When a matmul is sharded, expand the block algebra in the top formula as well as in the middle row, otherwise the reader cannot check the figure:
XW = [XW^(1) | ... | XW^(p)] = [H^(1) | ... | H^(p)]
[H^(1) | ... | H^(p)] [W^(1); ...; W^(p)] = Σ_r H^(r) W^(r) = Σ_r P^(r)
Column sharding splits the output axis (shards sit side by side); row sharding splits the
contracted axis (shards stack vertically and require a reduction). examples/tp-ffn-allreduce.tex
draws both in one figure.
Distinguish global from local
Per-rank shapes and global shapes are different objects. Label them differently
(d_ff/p vs d_ff) and, when both appear, say in the Axes row which one the figure
is drawing.