Files
SuperTensor/references/antipatterns.md
T
dela 7a22bef9e3 supertensor: shape-aware tensor figure toolkit
Extracted from the tensor-formula-viz skill and rebuilt around the idea that
the geometry rules should be enforced by construction rather than restated as
prose an agent has to remember.

- assets/supertensor.sty: faces, stacks, index faces, shared caption lanes,
  meaning box, signature. Macros take a declared axis and a declared role, so
  equal shapes get equal edges, a x a is square, a transpose swaps the face,
  and contracted axes share an edge length -- without any manual alignment.
- scripts/preflight.sh: decide the TikZ/CJK path before drawing.
- scripts/build.sh: compile and fail on silent corruption (missing CJK glyphs,
  overfull boxes, undeclared roles), then export pdf/svg/png/thumb.
- scripts/test.sh: build every figure as a regression test for the package.
- examples/: three golden figures (TP-FFN, causal MHA, MoE top-k gather) plus
  an anti-pattern gallery of figures that compile cleanly and still lie.
- SKILL.md + references/: lean entry point, details loaded on demand.
2026-08-05 12:17:33 +08:00

2.8 KiB
Raw Blame History

Anti-patterns

Every figure below compiles cleanly. build.sh is happy with all of them. They are still wrong, because the compiler checks TeX syntax and not whether the picture is true.

Render examples/antipatterns.tex and look at examples/build/antipatterns.png once before your first figure.

1. The transpose that only changed its label

A face captioned Kᵀ that is still T × d_h. The reader looks for the contracted axis, finds two faces of the same height, and concludes the contraction runs along the wrong dimension. Fix: swap the arguments — \ststack{KT}{...}{dh}{T}{3}. See geometry.md §3.

2. Shards that do not tile their parent

Two shards drawn with a gap, or stretched to fill a parent whose other shards were elided. Both assert a width that the tensor does not have. Fix: place each shard from the previous one's edge (($(W1a.east)+(2*\stunit,0)$)), and draw an ellipsis for anything omitted. See geometry.md §5–6.

3. An index drawn as a heatmap

Expert ids or token positions rendered with a lightness ramp. The ramp is a magnitude channel, so it says expert 3 > expert 0, which is meaningless. Fix: \stindexface. See semantics.md.

Same family: a Boolean mask drawn with graded cells (it has one level, not three), and a score matrix drawn as flat blocks (it has magnitude, and hiding it wastes the figure).

4. One pale level everywhere

A whole tensor in role!10. At full size it looks tasteful; at thumbnail size — which is how it will be seen on a slide — it is a blank rectangle. Fix: three separated levels, role!30 / role!55 / role!80. Contrast comes from lightness, not saturation. See style.md.

  • Periodic texture. A polynomial hash reduced mod 3 repeats every 3 rows, and the eye reads the resulting stripe as real structure. pattern=dense avoids it; if you write your own filler, check that rows 1, 2, 4, 5 of a tall face are not identical.
  • A label wider than its connector. The white underlay then covers the target tensor. Shorten the label or widen the gap — never let it sit on a face. See layout.md.
  • A new hue for a regrouped view of the same data. X and the per-expert buffers gathered out of X are the same object in a different order; a second hue claims they are different tensors.
  • Captions hanging at different depths because the faces in a row have different heights. Use \stlane.
  • A floating commentary card between two operands. If it is not a real operation, it belongs in the stage subtitle or the bottom box.
  • A meaning box that repeats the shapes. The shapes are already under every block. The box is for what the axes mean and what the operation does.
  • Solving crowding by shrinking type. The type hierarchy is a hard floor; move the stage to another row instead.