Extracted from the tensor-formula-viz skill and rebuilt around the idea that the geometry rules should be enforced by construction rather than restated as prose an agent has to remember. - assets/supertensor.sty: faces, stacks, index faces, shared caption lanes, meaning box, signature. Macros take a declared axis and a declared role, so equal shapes get equal edges, a x a is square, a transpose swaps the face, and contracted axes share an edge length -- without any manual alignment. - scripts/preflight.sh: decide the TikZ/CJK path before drawing. - scripts/build.sh: compile and fail on silent corruption (missing CJK glyphs, overfull boxes, undeclared roles), then export pdf/svg/png/thumb. - scripts/test.sh: build every figure as a regression test for the package. - examples/: three golden figures (TP-FFN, causal MHA, MoE top-k gather) plus an anti-pattern gallery of figures that compile cleanly and still lie. - SKILL.md + references/: lean entry point, details loaded on demand.
3.1 KiB
Symbol semantics
Shape is not meaning. Two tensors of shape T×k can be a score matrix, a list of token
positions, or a Boolean support, and drawing all three the same way is the fastest way to
mislead a reader who is trying to follow the mechanism.
Classify every non-obvious symbol
For each symbol record: semantic kind, dtype/domain, what one entry means, and its range when meaningful. Kinds worth separating:
| kind | example | domain |
|---|---|---|
| value / activation | X, H |
ℝ |
| score / logit | S = QKᵀ/√d_h |
ℝ |
| probability | A = softmax(S) |
[0,1], rows sum to 1 |
| index / coordinate | I = TopKIndices(G) |
{0,…,E−1} |
| rank / order | selected-slot axis r |
{1,…,k} |
| count | n_e tokens per expert |
ℕ |
| id | token id, device id | opaque |
| mask / support | D, causal M |
{0,1} or {0,−∞} |
| permutation | gather order | bijection |
| shape parameter | p, h |
ℕ, not drawn as a tensor |
One block, one object
Never merge a score, an index list and a mask under a label like M/S. Each conversion
gets an explicit operator and arrow. For selection or routing, close the entire chain:
continuous scores → discrete indices/ids → gather / scatter / mask / route → selected values
TopKValues and TopKIndices are different tensors; if both are used, show both.
Three grammars, deliberately different
| object | grammar | package |
|---|---|---|
| value / score / probability | magnitude — three separated lightness levels | \stface[pattern=dense] |
| index / id | discrete symbols in outlined cells, no lightness ramp | \stindexface{...}{entries} |
| mask / support | one flat level, exact structure, zeros unfilled | \stface[pattern=causal, level=3] or pattern=data |
examples/moe-topk-gather.tex puts all three in one figure on purpose. The reason indices
get no ramp: a ramp invites the reader to compare expert 3 > expert 0 as if the number
were a size.
Notation duties
- Define index notation and range at first use:
S_t = (s_{t,1},…,s_{t,k}),s_{t,r} ∈ {0,…,t}. - Distinguish the source-position axis
sfrom the selected-slot axisr, and say whether ordering, duplicates, padding or variable cardinality matter. - Show the address mapping once:
G[b,t,r,:] = X[b, S[b,t,r], :]. - If a mask is shown alongside the index tensor, state
M[b,t,s] = 1[s ∈ S_{b,t}]— do not let the figure imply they are the same object. - When one selector is shared across heads, ranks or branches, draw it once and mark
the broadcast/reuse axis. A per-head copy of a shared mask is a false claim about memory
and about the computation (
examples/mha-causal.tex:Mis a single face whileAis a three-sheet stack).
Colors carry semantics too
One tensor role keeps one hue for the whole figure — that is what \stsetrole is for. A
gathered, resharded or regrouped view of the same data keeps the same role color; a new
hue means a new object. Derived tensors may reuse their parent's family rather than
spending a hue (V → O → Y in the MHA example are all violet).