- scripts/lint.py: reject raw rectangles, absolute coordinates, hue-budget and callout/group/formula-order violations at the source level - tests/invalid/ + tests/lint-invalid/: negative fixtures proving the package and linter reject bad input; test.sh now runs both directions - references/fallback.md: degraded path when no LaTeX is available - tests/group-callout.tex: exercise \stgroup and \stcallout - agents/openai.yaml: agent config - Docs and .sty updated to match
3.5 KiB
Symbol semantics
Shape is not meaning. Two tensors of shape T×k can be a score matrix, a list of token
positions, or a Boolean support, and drawing all three the same way is the fastest way to
mislead a reader who is trying to follow the mechanism.
Classify every non-obvious symbol
For each symbol record: semantic kind, dtype/domain, what one entry means, and its range when meaningful. Kinds worth separating:
| kind | example | domain |
|---|---|---|
| value / activation | X, H |
ℝ |
| score / logit | S = QKᵀ/√d_h |
ℝ |
| probability | A = softmax(S) |
[0,1], rows sum to 1 |
| index / coordinate | I = TopKIndices(G) |
{0,…,E−1} |
| rank / order | selected-slot axis r |
{1,…,k} |
| count | n_e tokens per expert |
ℕ |
| id | token id, device id | opaque |
| mask / support | D, causal M |
{0,1} or {0,−∞} |
| permutation | gather order | bijection |
| shape parameter | p, h |
ℕ, not drawn as a tensor |
One block, one object
Never merge a score, an index list and a mask under a label like M/S. Each conversion
gets an explicit operator and arrow. For selection or routing, close the entire chain:
continuous scores → discrete indices/ids → gather / scatter / mask / route → selected values
TopKValues and TopKIndices are different tensors; if both are used, show both.
Three grammars, deliberately different
| object | grammar | package |
|---|---|---|
| value / score / probability | magnitude — three separated lightness levels | \stface[pattern=dense] |
| index / id | discrete symbols in outlined cells, no lightness ramp | \stindexface{...}{entries} |
| mask / support | one flat level, exact structure, zeros unfilled | \stface[pattern=causal, level=3] or pattern=data |
examples/moe-topk-gather.tex puts all three in one figure on purpose. The reason indices
get no ramp: a ramp invites the reader to compare expert 3 > expert 0 as if the number
were a size.
Notation duties
- Define index notation and range at first use:
S_t = (s_{t,1},…,s_{t,k}),s_{t,r} ∈ {0,…,t}. - Distinguish the source-position axis
sfrom the selected-slot axisr, and say whether ordering, duplicates, padding or variable cardinality matter. - Show the address mapping once:
G[b,t,r,:] = X[b, S[b,t,r], :]. - If a mask is shown alongside the index tensor, state
M[b,t,s] = 1[s ∈ S_{b,t}]— do not let the figure imply they are the same object. - When one selector is shared across heads, ranks or branches, draw it once and mark
the broadcast/reuse axis. A per-head copy of a shared mask is a false claim about memory
and about the computation (
examples/mha-causal.tex:Mis a single face whileAis a three-sheet stack).
Colors carry semantics too
One tensor role keeps one hue for the whole figure — that is what \stsetrole is for. A
gathered, resharded or regrouped view of the same data keeps the same role color; a new
hue means a new object. Derived tensors may reuse their parent's family rather than
spending a hue (V → O → Y in the MHA example are all violet).
A \stgroup outline follows the same rule, because a group is a regrouped view: give it
the role of the objects it wraps — the three-sheet q stack and the outline that names it
as one composite are the same object, and a new hue there would claim a new tensor exists.
Only when the members genuinely differ in role does the group take role=neutral; that is
also the honest signal that the box is naming an arrangement rather than an object.