supertensor: shape-aware tensor figure toolkit

Extracted from the tensor-formula-viz skill and rebuilt around the idea that
the geometry rules should be enforced by construction rather than restated as
prose an agent has to remember.

- assets/supertensor.sty: faces, stacks, index faces, shared caption lanes,
  meaning box, signature. Macros take a declared axis and a declared role, so
  equal shapes get equal edges, a x a is square, a transpose swaps the face,
  and contracted axes share an edge length -- without any manual alignment.
- scripts/preflight.sh: decide the TikZ/CJK path before drawing.
- scripts/build.sh: compile and fail on silent corruption (missing CJK glyphs,
  overfull boxes, undeclared roles), then export pdf/svg/png/thumb.
- scripts/test.sh: build every figure as a regression test for the package.
- examples/: three golden figures (TP-FFN, causal MHA, MoE top-k gather) plus
  an anti-pattern gallery of figures that compile cleanly and still lie.
- SKILL.md + references/: lean entry point, details loaded on demand.
This commit is contained in:
dela
2026-08-05 12:17:33 +08:00
commit 7a22bef9e3
20 changed files with 1661 additions and 0 deletions
+67
View File
@@ -0,0 +1,67 @@
# Symbol semantics
Shape is not meaning. Two tensors of shape `T×k` can be a score matrix, a list of token
positions, or a Boolean support, and drawing all three the same way is the fastest way to
mislead a reader who is trying to follow the mechanism.
## Classify every non-obvious symbol
For each symbol record: **semantic kind**, **dtype/domain**, **what one entry means**, and
its **range** when meaningful. Kinds worth separating:
| kind | example | domain |
|---|---|---|
| value / activation | `X`, `H` | ℝ |
| score / logit | `S = QKᵀ/√d_h` | ℝ |
| probability | `A = softmax(S)` | [0,1], rows sum to 1 |
| index / coordinate | `I = TopKIndices(G)` | {0,…,E−1} |
| rank / order | selected-slot axis `r` | {1,…,k} |
| count | `n_e` tokens per expert | ℕ |
| id | token id, device id | opaque |
| mask / support | `D`, causal `M` | {0,1} or {0,−∞} |
| permutation | gather order | bijection |
| shape parameter | `p`, `h` | ℕ, not drawn as a tensor |
## One block, one object
Never merge a score, an index list and a mask under a label like `M/S`. Each conversion
gets an explicit operator and arrow. For selection or routing, close the entire chain:
```
continuous scores → discrete indices/ids → gather / scatter / mask / route → selected values
```
`TopKValues` and `TopKIndices` are different tensors; if both are used, show both.
## Three grammars, deliberately different
| object | grammar | package |
|---|---|---|
| value / score / probability | magnitude — three separated lightness levels | `\stface[pattern=dense]` |
| index / id | discrete symbols in outlined cells, **no** lightness ramp | `\stindexface{...}{entries}` |
| mask / support | one flat level, exact structure, zeros unfilled | `\stface[pattern=causal, level=3]` or `pattern=data` |
`examples/moe-topk-gather.tex` puts all three in one figure on purpose. The reason indices
get no ramp: a ramp invites the reader to compare `expert 3 > expert 0` as if the number
were a size.
## Notation duties
- Define index notation and range at first use:
`S_t = (s_{t,1},…,s_{t,k})`, `s_{t,r} ∈ {0,…,t}`.
- Distinguish the *source-position* axis `s` from the *selected-slot* axis `r`, and say
whether ordering, duplicates, padding or variable cardinality matter.
- Show the address mapping once: `G[b,t,r,:] = X[b, S[b,t,r], :]`.
- If a mask is shown alongside the index tensor, state `M[b,t,s] = 1[s ∈ S_{b,t}]` — do not
let the figure imply they are the same object.
- When one selector is shared across heads, ranks or branches, draw it **once** and mark
the broadcast/reuse axis. A per-head copy of a shared mask is a false claim about memory
and about the computation (`examples/mha-causal.tex`: `M` is a single face while `A` is a
three-sheet stack).
## Colors carry semantics too
One tensor role keeps one hue for the whole figure — that is what `\stsetrole` is for. A
gathered, resharded or regrouped view of the same data keeps the *same* role color; a new
hue means a new object. Derived tensors may reuse their parent's family rather than
spending a hue (`V → O → Y` in the MHA example are all violet).