Files
SuperTensor/references/semantics.md
T
dela 7a22bef9e3 supertensor: shape-aware tensor figure toolkit
Extracted from the tensor-formula-viz skill and rebuilt around the idea that
the geometry rules should be enforced by construction rather than restated as
prose an agent has to remember.

- assets/supertensor.sty: faces, stacks, index faces, shared caption lanes,
  meaning box, signature. Macros take a declared axis and a declared role, so
  equal shapes get equal edges, a x a is square, a transpose swaps the face,
  and contracted axes share an edge length -- without any manual alignment.
- scripts/preflight.sh: decide the TikZ/CJK path before drawing.
- scripts/build.sh: compile and fail on silent corruption (missing CJK glyphs,
  overfull boxes, undeclared roles), then export pdf/svg/png/thumb.
- scripts/test.sh: build every figure as a regression test for the package.
- examples/: three golden figures (TP-FFN, causal MHA, MoE top-k gather) plus
  an anti-pattern gallery of figures that compile cleanly and still lie.
- SKILL.md + references/: lean entry point, details loaded on demand.
2026-08-05 12:17:33 +08:00

68 lines
3.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Symbol semantics
Shape is not meaning. Two tensors of shape `T×k` can be a score matrix, a list of token
positions, or a Boolean support, and drawing all three the same way is the fastest way to
mislead a reader who is trying to follow the mechanism.
## Classify every non-obvious symbol
For each symbol record: **semantic kind**, **dtype/domain**, **what one entry means**, and
its **range** when meaningful. Kinds worth separating:
| kind | example | domain |
|---|---|---|
| value / activation | `X`, `H` | ℝ |
| score / logit | `S = QKᵀ/√d_h` | ℝ |
| probability | `A = softmax(S)` | [0,1], rows sum to 1 |
| index / coordinate | `I = TopKIndices(G)` | {0,…,E−1} |
| rank / order | selected-slot axis `r` | {1,…,k} |
| count | `n_e` tokens per expert | ℕ |
| id | token id, device id | opaque |
| mask / support | `D`, causal `M` | {0,1} or {0,−∞} |
| permutation | gather order | bijection |
| shape parameter | `p`, `h` | ℕ, not drawn as a tensor |
## One block, one object
Never merge a score, an index list and a mask under a label like `M/S`. Each conversion
gets an explicit operator and arrow. For selection or routing, close the entire chain:
```
continuous scores → discrete indices/ids → gather / scatter / mask / route → selected values
```
`TopKValues` and `TopKIndices` are different tensors; if both are used, show both.
## Three grammars, deliberately different
| object | grammar | package |
|---|---|---|
| value / score / probability | magnitude — three separated lightness levels | `\stface[pattern=dense]` |
| index / id | discrete symbols in outlined cells, **no** lightness ramp | `\stindexface{...}{entries}` |
| mask / support | one flat level, exact structure, zeros unfilled | `\stface[pattern=causal, level=3]` or `pattern=data` |
`examples/moe-topk-gather.tex` puts all three in one figure on purpose. The reason indices
get no ramp: a ramp invites the reader to compare `expert 3 > expert 0` as if the number
were a size.
## Notation duties
- Define index notation and range at first use:
`S_t = (s_{t,1},…,s_{t,k})`, `s_{t,r} ∈ {0,…,t}`.
- Distinguish the *source-position* axis `s` from the *selected-slot* axis `r`, and say
whether ordering, duplicates, padding or variable cardinality matter.
- Show the address mapping once: `G[b,t,r,:] = X[b, S[b,t,r], :]`.
- If a mask is shown alongside the index tensor, state `M[b,t,s] = 1[s ∈ S_{b,t}]` — do not
let the figure imply they are the same object.
- When one selector is shared across heads, ranks or branches, draw it **once** and mark
the broadcast/reuse axis. A per-head copy of a shared mask is a false claim about memory
and about the computation (`examples/mha-causal.tex`: `M` is a single face while `A` is a
three-sheet stack).
## Colors carry semantics too
One tensor role keeps one hue for the whole figure — that is what `\stsetrole` is for. A
gathered, resharded or regrouped view of the same data keeps the *same* role color; a new
hue means a new object. Derived tensors may reuse their parent's family rather than
spending a hue (`V → O → Y` in the MHA example are all violet).