Files
dela 866173a831 Add source linter, negative test fixtures, and fallback guidance
- scripts/lint.py: reject raw rectangles, absolute coordinates, hue-budget
  and callout/group/formula-order violations at the source level
- tests/invalid/ + tests/lint-invalid/: negative fixtures proving the
  package and linter reject bad input; test.sh now runs both directions
- references/fallback.md: degraded path when no LaTeX is available
- tests/group-callout.tex: exercise \stgroup and \stcallout
- agents/openai.yaml: agent config
- Docs and .sty updated to match
2026-08-05 16:41:22 +08:00

4.0 KiB
Raw Permalink Blame History

Geometry invariants

A figure that is drawn to the wrong geometry is not a stylistic problem — it teaches the reader a false fact about the computation. These invariants are mandatory. Most of them are automatic if you declare the geometry ledger and never pass a raw number.

The ledger

\stdim{T}{6}       % sequence length  -> 6 cells, everywhere in the figure
\stdim{d}{9}       % model width      -> 9 cells, everywhere
\stdim{dh}{3}      % head width       -> 3 cells, and 3*dh = d holds visually

\stface{...}{T}{dh} resolves the names through the ledger, so one symbolic axis maps to exactly one physical edge length across the whole figure. Equal shapes therefore form an equivalence class automatically: Q and V at T×d_h come out identical without you lining anything up by hand.

Declaring the same axis twice with the same value is allowed. Redeclaring it with a different value emits a package warning, preserves the first value, and fails build.sh.

Raw integers are accepted (\stface{A}{(0,0)}{4}{4}) but they opt out of the guarantee. Use them only for a face whose axis appears nowhere else. The source linter rejects a repeated raw dimension greater than one; give repeated dimensions a symbolic name.

The rules

  1. Face orientation. A matrix face a×b is height a, width b. Always. For batched or stacked tensors, the last two axes make the face; leading axes become depth (\ststack) or repeated panels — never a wider rectangle.

  2. Squares. a×a renders as a square. Automatic when both arguments resolve to the same declared axis.

  3. Transpose. Draw K^T by physically swapping height and width: \ststack{KT}{...}{dh}{T}{3} against \ststack{K}{...}{T}{dh}{3}. Relabelling a face K^T while leaving its shape alone is invalid — it is the single most common lie in attention figures.

  4. Contraction. In (m×k)(k×n), both occurrences of k get the same edge length. With the ledger this is free: pass the same axis name to the width of the left face and the height of the right one. The same applies to einsum and attention axes.

  5. Partition. Explicit shards tile their parent exactly along the split axis, equal shards are equal in size, and concatenation reverses the split. Place shards from the previous face's edge so no gap can creep in:

    \stface[role=r1]{W1a}{(...)}{d}{dffl}
    \stface[role=r2]{W1b}{($(W1a.east)+(2*\stunit,0)$)}{d}{dffl}
    

    The offset is half-width of the next face in \stunit, so the two faces are exactly adjacent. Size concatenated parts from their declared shapes — Q/K/V are equal segments only when their output shapes are equal.

  6. Elision. If intermediate shards are omitted, draw an ellipsis. Never stretch the visible shards to impersonate the full parent.

  7. Axis changes. Geometry may change only at an explicit reshape, flatten, transpose, split, or concat operator, and the axis identity must be stated — e.g. (h/p)·d_h = d/p. Do not silently reuse one generic rectangle on both sides of an axis change.

  8. Illustrative counts. Cell counts need not equal real dimensions. Choosing d=9 to stand for 4096 is fine. It never waives rules 1–7: the ratios you draw are read as facts. If d = h·d_h, pick numbers where that arithmetic actually holds.

Sharded matmul

When a matmul is sharded, expand the block algebra in the top formula as well as in the middle row, otherwise the reader cannot check the figure:

XW = [XW^(1) | ... | XW^(p)] = [H^(1) | ... | H^(p)]

[H^(1) | ... | H^(p)] [W^(1); ...; W^(p)] = Σ_r H^(r) W^(r) = Σ_r P^(r)

Column sharding splits the output axis (shards sit side by side); row sharding splits the contracted axis (shards stack vertically and require a reduction). examples/tp-ffn-allreduce.tex draws both in one figure.

Distinguish global from local

Per-rank shapes and global shapes are different objects. Label them differently (d_ff/p vs d_ff) and, when both appear, say in the Axes row which one the figure is drawing.