Files
SuperTensor/references/geometry.md
T
dela 7a22bef9e3 supertensor: shape-aware tensor figure toolkit
Extracted from the tensor-formula-viz skill and rebuilt around the idea that
the geometry rules should be enforced by construction rather than restated as
prose an agent has to remember.

- assets/supertensor.sty: faces, stacks, index faces, shared caption lanes,
  meaning box, signature. Macros take a declared axis and a declared role, so
  equal shapes get equal edges, a x a is square, a transpose swaps the face,
  and contracted axes share an edge length -- without any manual alignment.
- scripts/preflight.sh: decide the TikZ/CJK path before drawing.
- scripts/build.sh: compile and fail on silent corruption (missing CJK glyphs,
  overfull boxes, undeclared roles), then export pdf/svg/png/thumb.
- scripts/test.sh: build every figure as a regression test for the package.
- examples/: three golden figures (TP-FFN, causal MHA, MoE top-k gather) plus
  an anti-pattern gallery of figures that compile cleanly and still lie.
- SKILL.md + references/: lean entry point, details loaded on demand.
2026-08-05 12:17:33 +08:00

78 lines
3.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Geometry invariants
A figure that is drawn to the wrong geometry is not a stylistic problem — it teaches the
reader a false fact about the computation. These invariants are mandatory. Most of them
are automatic if you declare the geometry ledger and never pass a raw number.
## The ledger
```tex
\stdim{T}{6} % sequence length -> 6 cells, everywhere in the figure
\stdim{d}{9} % model width -> 9 cells, everywhere
\stdim{dh}{3} % head width -> 3 cells, and 3*dh = d holds visually
```
`\stface{...}{T}{dh}` resolves the names through the ledger, so **one symbolic axis maps
to exactly one physical edge length across the whole figure**. Equal shapes therefore form
an equivalence class automatically: `Q` and `V` at `T×d_h` come out identical without you
lining anything up by hand.
Raw integers are accepted (`\stface{A}{(0,0)}{4}{4}`) but they opt out of the guarantee.
Use them only for a face whose axis appears nowhere else.
## The rules
1. **Face orientation.** A matrix face `a×b` is height `a`, width `b`. Always. For batched
or stacked tensors, the *last two* axes make the face; leading axes become depth
(`\ststack`) or repeated panels — never a wider rectangle.
2. **Squares.** `a×a` renders as a square. Automatic when both arguments resolve to the
same declared axis.
3. **Transpose.** Draw `K^T` by physically swapping height and width:
`\ststack{KT}{...}{dh}{T}{3}` against `\ststack{K}{...}{T}{dh}{3}`. Relabelling a face
`K^T` while leaving its shape alone is invalid — it is the single most common lie in
attention figures.
4. **Contraction.** In `(m×k)(k×n)`, both occurrences of `k` get the same edge length.
With the ledger this is free: pass the same axis name to the width of the left face and
the height of the right one. The same applies to `einsum` and attention axes.
5. **Partition.** Explicit shards tile their parent exactly along the split axis, equal
shards are equal in size, and concatenation reverses the split. Place shards from the
previous face's edge so no gap can creep in:
```tex
\stface[role=r1]{W1a}{(...)}{d}{dffl}
\stface[role=r2]{W1b}{($(W1a.east)+(2*\stunit,0)$)}{d}{dffl}
```
The offset is `half-width of the next face` in `\stunit`, so the two faces are exactly
adjacent. Size concatenated parts from their *declared* shapes — Q/K/V are equal
segments only when their output shapes are equal.
6. **Elision.** If intermediate shards are omitted, draw an ellipsis. Never stretch the
visible shards to impersonate the full parent.
7. **Axis changes.** Geometry may change only at an explicit reshape, flatten, transpose,
split, or concat operator, and the axis identity must be stated — e.g. `(h/p)·d_h = d/p`.
Do not silently reuse one generic rectangle on both sides of an axis change.
8. **Illustrative counts.** Cell counts need not equal real dimensions. Choosing `d=9`
to stand for 4096 is fine. It never waives rules 1–7: the *ratios* you draw are read as
facts. If `d = h·d_h`, pick numbers where that arithmetic actually holds.
## Sharded matmul
When a matmul is sharded, expand the block algebra in the top formula as well as in the
middle row, otherwise the reader cannot check the figure:
```
XW = [XW^(1) | ... | XW^(p)] = [H^(1) | ... | H^(p)]
[H^(1) | ... | H^(p)] [W^(1); ...; W^(p)] = Σ_r H^(r) W^(r) = Σ_r P^(r)
```
Column sharding splits the *output* axis (shards sit side by side); row sharding splits the
*contracted* axis (shards stack vertically and require a reduction). `examples/tp-ffn-allreduce.tex`
draws both in one figure.
## Distinguish global from local
Per-rank shapes and global shapes are different objects. Label them differently
(`d_ff/p` vs `d_ff`) and, when both appear, say in the **Axes** row which one the figure
is drawing.