Files
SuperTensor/references/semantics.md
T
dela 7a22bef9e3 supertensor: shape-aware tensor figure toolkit
Extracted from the tensor-formula-viz skill and rebuilt around the idea that
the geometry rules should be enforced by construction rather than restated as
prose an agent has to remember.

- assets/supertensor.sty: faces, stacks, index faces, shared caption lanes,
  meaning box, signature. Macros take a declared axis and a declared role, so
  equal shapes get equal edges, a x a is square, a transpose swaps the face,
  and contracted axes share an edge length -- without any manual alignment.
- scripts/preflight.sh: decide the TikZ/CJK path before drawing.
- scripts/build.sh: compile and fail on silent corruption (missing CJK glyphs,
  overfull boxes, undeclared roles), then export pdf/svg/png/thumb.
- scripts/test.sh: build every figure as a regression test for the package.
- examples/: three golden figures (TP-FFN, causal MHA, MoE top-k gather) plus
  an anti-pattern gallery of figures that compile cleanly and still lie.
- SKILL.md + references/: lean entry point, details loaded on demand.
2026-08-05 12:17:33 +08:00

3.1 KiB
Raw Blame History

Symbol semantics

Shape is not meaning. Two tensors of shape T×k can be a score matrix, a list of token positions, or a Boolean support, and drawing all three the same way is the fastest way to mislead a reader who is trying to follow the mechanism.

Classify every non-obvious symbol

For each symbol record: semantic kind, dtype/domain, what one entry means, and its range when meaningful. Kinds worth separating:

kind example domain
value / activation X, H ℝ
score / logit S = QKᵀ/√d_h ℝ
probability A = softmax(S) [0,1], rows sum to 1
index / coordinate I = TopKIndices(G) {0,…,E−1}
rank / order selected-slot axis r {1,…,k}
count n_e tokens per expert ℕ
id token id, device id opaque
mask / support D, causal M {0,1} or {0,−∞}
permutation gather order bijection
shape parameter p, h ℕ, not drawn as a tensor

One block, one object

Never merge a score, an index list and a mask under a label like M/S. Each conversion gets an explicit operator and arrow. For selection or routing, close the entire chain:

continuous scores → discrete indices/ids → gather / scatter / mask / route → selected values

TopKValues and TopKIndices are different tensors; if both are used, show both.

Three grammars, deliberately different

object grammar package
value / score / probability magnitude — three separated lightness levels \stface[pattern=dense]
index / id discrete symbols in outlined cells, no lightness ramp \stindexface{...}{entries}
mask / support one flat level, exact structure, zeros unfilled \stface[pattern=causal, level=3] or pattern=data

examples/moe-topk-gather.tex puts all three in one figure on purpose. The reason indices get no ramp: a ramp invites the reader to compare expert 3 > expert 0 as if the number were a size.

Notation duties

  • Define index notation and range at first use: S_t = (s_{t,1},…,s_{t,k}), s_{t,r} ∈ {0,…,t}.
  • Distinguish the source-position axis s from the selected-slot axis r, and say whether ordering, duplicates, padding or variable cardinality matter.
  • Show the address mapping once: G[b,t,r,:] = X[b, S[b,t,r], :].
  • If a mask is shown alongside the index tensor, state M[b,t,s] = 1[s ∈ S_{b,t}] — do not let the figure imply they are the same object.
  • When one selector is shared across heads, ranks or branches, draw it once and mark the broadcast/reuse axis. A per-head copy of a shared mask is a false claim about memory and about the computation (examples/mha-causal.tex: M is a single face while A is a three-sheet stack).

Colors carry semantics too

One tensor role keeps one hue for the whole figure — that is what \stsetrole is for. A gathered, resharded or regrouped view of the same data keeps the same role color; a new hue means a new object. Derived tensors may reuse their parent's family rather than spending a hue (V → O → Y in the MHA example are all violet).