Files
SuperTensor/references/geometry.md
T
dela 7a22bef9e3 supertensor: shape-aware tensor figure toolkit
Extracted from the tensor-formula-viz skill and rebuilt around the idea that
the geometry rules should be enforced by construction rather than restated as
prose an agent has to remember.

- assets/supertensor.sty: faces, stacks, index faces, shared caption lanes,
  meaning box, signature. Macros take a declared axis and a declared role, so
  equal shapes get equal edges, a x a is square, a transpose swaps the face,
  and contracted axes share an edge length -- without any manual alignment.
- scripts/preflight.sh: decide the TikZ/CJK path before drawing.
- scripts/build.sh: compile and fail on silent corruption (missing CJK glyphs,
  overfull boxes, undeclared roles), then export pdf/svg/png/thumb.
- scripts/test.sh: build every figure as a regression test for the package.
- examples/: three golden figures (TP-FFN, causal MHA, MoE top-k gather) plus
  an anti-pattern gallery of figures that compile cleanly and still lie.
- SKILL.md + references/: lean entry point, details loaded on demand.
2026-08-05 12:17:33 +08:00

3.8 KiB
Raw Blame History

Geometry invariants

A figure that is drawn to the wrong geometry is not a stylistic problem — it teaches the reader a false fact about the computation. These invariants are mandatory. Most of them are automatic if you declare the geometry ledger and never pass a raw number.

The ledger

\stdim{T}{6}       % sequence length  -> 6 cells, everywhere in the figure
\stdim{d}{9}       % model width      -> 9 cells, everywhere
\stdim{dh}{3}      % head width       -> 3 cells, and 3*dh = d holds visually

\stface{...}{T}{dh} resolves the names through the ledger, so one symbolic axis maps to exactly one physical edge length across the whole figure. Equal shapes therefore form an equivalence class automatically: Q and V at T×d_h come out identical without you lining anything up by hand.

Raw integers are accepted (\stface{A}{(0,0)}{4}{4}) but they opt out of the guarantee. Use them only for a face whose axis appears nowhere else.

The rules

  1. Face orientation. A matrix face a×b is height a, width b. Always. For batched or stacked tensors, the last two axes make the face; leading axes become depth (\ststack) or repeated panels — never a wider rectangle.

  2. Squares. a×a renders as a square. Automatic when both arguments resolve to the same declared axis.

  3. Transpose. Draw K^T by physically swapping height and width: \ststack{KT}{...}{dh}{T}{3} against \ststack{K}{...}{T}{dh}{3}. Relabelling a face K^T while leaving its shape alone is invalid — it is the single most common lie in attention figures.

  4. Contraction. In (m×k)(k×n), both occurrences of k get the same edge length. With the ledger this is free: pass the same axis name to the width of the left face and the height of the right one. The same applies to einsum and attention axes.

  5. Partition. Explicit shards tile their parent exactly along the split axis, equal shards are equal in size, and concatenation reverses the split. Place shards from the previous face's edge so no gap can creep in:

    \stface[role=r1]{W1a}{(...)}{d}{dffl}
    \stface[role=r2]{W1b}{($(W1a.east)+(2*\stunit,0)$)}{d}{dffl}
    

    The offset is half-width of the next face in \stunit, so the two faces are exactly adjacent. Size concatenated parts from their declared shapes — Q/K/V are equal segments only when their output shapes are equal.

  6. Elision. If intermediate shards are omitted, draw an ellipsis. Never stretch the visible shards to impersonate the full parent.

  7. Axis changes. Geometry may change only at an explicit reshape, flatten, transpose, split, or concat operator, and the axis identity must be stated — e.g. (h/p)·d_h = d/p. Do not silently reuse one generic rectangle on both sides of an axis change.

  8. Illustrative counts. Cell counts need not equal real dimensions. Choosing d=9 to stand for 4096 is fine. It never waives rules 1–7: the ratios you draw are read as facts. If d = h·d_h, pick numbers where that arithmetic actually holds.

Sharded matmul

When a matmul is sharded, expand the block algebra in the top formula as well as in the middle row, otherwise the reader cannot check the figure:

XW = [XW^(1) | ... | XW^(p)] = [H^(1) | ... | H^(p)]

[H^(1) | ... | H^(p)] [W^(1); ...; W^(p)] = Σ_r H^(r) W^(r) = Σ_r P^(r)

Column sharding splits the output axis (shards sit side by side); row sharding splits the contracted axis (shards stack vertically and require a reduction). examples/tp-ffn-allreduce.tex draws both in one figure.

Distinguish global from local

Per-rank shapes and global shapes are different objects. Label them differently (d_ff/p vs d_ff) and, when both appear, say in the Axes row which one the figure is drawing.