# Symbol semantics Shape is not meaning. Two tensors of shape `T×k` can be a score matrix, a list of token positions, or a Boolean support, and drawing all three the same way is the fastest way to mislead a reader who is trying to follow the mechanism. ## Classify every non-obvious symbol For each symbol record: **semantic kind**, **dtype/domain**, **what one entry means**, and its **range** when meaningful. Kinds worth separating: | kind | example | domain | |---|---|---| | value / activation | `X`, `H` | ℝ | | score / logit | `S = QKᵀ/√d_h` | ℝ | | probability | `A = softmax(S)` | [0,1], rows sum to 1 | | index / coordinate | `I = TopKIndices(G)` | {0,…,E−1} | | rank / order | selected-slot axis `r` | {1,…,k} | | count | `n_e` tokens per expert | ℕ | | id | token id, device id | opaque | | mask / support | `D`, causal `M` | {0,1} or {0,−∞} | | permutation | gather order | bijection | | shape parameter | `p`, `h` | ℕ, not drawn as a tensor | ## One block, one object Never merge a score, an index list and a mask under a label like `M/S`. Each conversion gets an explicit operator and arrow. For selection or routing, close the entire chain: ``` continuous scores → discrete indices/ids → gather / scatter / mask / route → selected values ``` `TopKValues` and `TopKIndices` are different tensors; if both are used, show both. ## Three grammars, deliberately different | object | grammar | package | |---|---|---| | value / score / probability | magnitude — three separated lightness levels | `\stface[pattern=dense]` | | index / id | discrete symbols in outlined cells, **no** lightness ramp | `\stindexface{...}{entries}` | | mask / support | one flat level, exact structure, zeros unfilled | `\stface[pattern=causal, level=3]` or `pattern=data` | `examples/moe-topk-gather.tex` puts all three in one figure on purpose. The reason indices get no ramp: a ramp invites the reader to compare `expert 3 > expert 0` as if the number were a size. ## Notation duties - Define index notation and range at first use: `S_t = (s_{t,1},…,s_{t,k})`, `s_{t,r} ∈ {0,…,t}`. - Distinguish the *source-position* axis `s` from the *selected-slot* axis `r`, and say whether ordering, duplicates, padding or variable cardinality matter. - Show the address mapping once: `G[b,t,r,:] = X[b, S[b,t,r], :]`. - If a mask is shown alongside the index tensor, state `M[b,t,s] = 1[s ∈ S_{b,t}]` — do not let the figure imply they are the same object. - When one selector is shared across heads, ranks or branches, draw it **once** and mark the broadcast/reuse axis. A per-head copy of a shared mask is a false claim about memory and about the computation (`examples/mha-causal.tex`: `M` is a single face while `A` is a three-sheet stack). ## Colors carry semantics too One tensor role keeps one hue for the whole figure — that is what `\stsetrole` is for. A gathered, resharded or regrouped view of the same data keeps the *same* role color; a new hue means a new object. Derived tensors may reuse their parent's family rather than spending a hue (`V → O → Y` in the MHA example are all violet). A `\stgroup` outline follows the same rule, because a group *is* a regrouped view: give it the role of the objects it wraps — the three-sheet `q` stack and the outline that names it as one composite are the same object, and a new hue there would claim a new tensor exists. Only when the members genuinely differ in role does the group take `role=neutral`; that is also the honest signal that the box is naming an arrangement rather than an object.