ADR 0079: CUBE and matrix operations
Context
ADR 0062 recorded a single repository-wide mnemonic audit. This record preserves the accepted decisions for this family as one decision-scoped owner. The former identifiers remain only in legacy_ids and the generated ADR index; current normative meaning is owned by the affected ASL/NDF clauses.
Decisions
Decision 049: TMATMUL dimensions are M, N, and K in LB order
BSTART.TMATMUL uses LB0=M, LB1=N, and LB2=K. Omission of any one of
these fields supplies the value one for that field. An explicitly encoded or
otherwise resolved value of zero is illegal. Each of M, N, and K MUST
be a nonzero power of two.
The left source has valid shape M x K, the right source has valid shape
K x N, and the destination has valid shape M x N. Each physical Tile row
and column extent MUST be a power of two, and the physical extent MUST contain
the complete valid rectangle. Shape, capacity, and compatibility checks MUST
complete before destination allocation or any other effect.
Decision 050: TMATMUL supports mixed same-class input types and fixed accumulator classes
The BSTART.TMATMUL DataType selects the left source type. An optional
B.DATR.DataType selects the right source type; if that field is absent, the
right source type defaults to the left source type. Absence is distinct from
an explicitly encoded zero, because encoded DataType zero denotes FP64,
which is not supported by TMATMUL. The block schema MUST therefore preserve
field presence when applying this default.
The supported floating input set is FP32, TF32, HF32, FP16, BF16,
HiF8, E4M3, E5M2, E3M2, E2M3, E2M1X2, and E1M2X2. The supported
signed input set is S16, S8, and S4X2. The supported unsigned input set
is U16, U8, and U4X2. The two inputs MAY use different types within the
same floating, signed, or unsigned class. A cross-class pair is illegal.
FP64, E8M0, HiF4X2, S64, S32, U64, U32, every globally reserved
DataType encoding, and every other type not listed above are reserved or
unsupported for ordinary TMATMUL and MUST be rejected before effects.
A floating pair produces an FP32 accumulator result, a signed pair produces
an S32 accumulator result, and an unsigned pair produces a U32 accumulator
result. This result uses the private CUBE output representation. TCVT, not
TMATMUL, performs conversion to a canonical public representation. A numeric
profile MAY define format-specific arithmetic details, but it MUST preserve
these operand classes, accumulator classes, and encoded numeric controls.
Decision 051: TMATMUL binding, participation, and supplementary fields are closed
The Local form binds a left Local source, a right Local source, and an explicit
new Local destination. The Shared forms bind either a Local left source plus a
Shared right source, or a Shared left source plus a Shared right source; the
destination remains an explicit new Local allocation. B.IOS is source-only
for TMATMUL and reads only fully published Shared values.
PE_MASK=0000 is a strict no-op before descriptor access, readiness checks,
allocation, lifetime consumption, faults, or effects. Every executing Local or
Shared binding MUST use PE_MASK=1111; a nonzero partial mask is illegal before
effects. All source payloads are snapshotted after complete preflight and before
destination allocation or commit, so source-destination aliasing observes the
old source values and the destination becomes visible only as one complete
result.
TMATMUL does not consume mathematical B.IOR operands. Its B.DATR fields
are closed to DataType, RMode, and Sat. An omitted DataType applies the
right-type default in Decision 050 in ADR-0079, omitted RMode selects RNE, and omitted Sat
selects disabled saturation. Layout, CMode, PadValueOrByteId, and
Canonicalize MUST be zero. The private CUBE result remains noncanonical until
an explicit TCVT operation.
The earlier statement that TMATMUL did not consume B.FPATR is superseded
by ADR 0064. Every Matrix CUBE bundle contains exactly one B.FPATR; its
all-zero form selects no post-processing, and nonzero modes determine the
additional scalar, Local source, and Local destination schema.
Decision 052: TMATMUL.BIAS adds one 1 x N right-side broadcast source
BSTART.TMATMUL.BIAS inherits the complete dimension, input-type, accumulator,
binding, participation, supplementary-field, preflight, and commit contract of
BSTART.TMATMUL. It adds one Bias source after the left and right matrix
sources.
The Bias source valid shape MUST be exactly 1 x N, its layout MUST be
row-major, and its DataType MUST equal the result accumulator class selected by
Decision 050 in ADR-0079: FP32, S32, or U32. Its payload uses the same private CUBE result
representation as that accumulator class. For every output row i and column
j, Bias[0,j] is added once to the complete dot product for output D[i,j].
No row broadcast, scalar broadcast, full-matrix Bias, or Bias addition inside
the K reduction is defined.
Bias remains a Local source in both the all-Local and Shared-matrix forms.
Shared bindings MAY supply the right matrix or both matrix operands exactly as
defined for TMATMUL; they do not bind the Bias or destination. The complete
matrix and Bias sources are snapshotted after preflight, and the explicit new
Local destination becomes visible only as one complete result.
Decision 053: TMATMUL.ACC uses explicit Local accumulator input and destination
BSTART.TMATMUL.ACC inherits the complete dimension, matrix input-type,
accumulator-class, binding, participation, supplementary-field, preflight, and
commit contract of BSTART.TMATMUL. It adds one explicit Local accumulator
source C before the left and right matrix sources and writes one explicit new
Local destination D.
C MUST have valid shape M x N, the same row-major layout and physical
capacity as D, and the same private CUBE accumulator representation and
DataType as the result: FP32, S32, or U32. The result is
D = C + A x B, with the encoded RMode and Sat controls applied according
to the selected numeric profile. There is no implicit ACC operand or implicit
destination.
The operation snapshots C, A, and B after complete preflight and before
writing D. C and D MAY resolve to the same Local Tile; this case has
read-old/write-new behavior and becomes visible only as one complete result.
In Shared-matrix forms, Shared bindings MAY supply the right matrix or both
matrix operands, but C and D remain explicit Local bindings.
Decision 054: TMATMULMX scales each matrix side independently
BSTART.TMATMULMX inherits TMATMUL dimension ordering, matrix and destination
shapes, explicit Local destination, Core4 participation, private FP32 result,
numeric controls, preflight, snapshot, and atomic commit rules. Its matrix
inputs are floating only. Each side independently uses one of FP16, BF16,
E4M3, E5M2, E2M1X2, or E1M2X2; the two sides MAY use different listed
types. Every other type, including HiF4X2, is unsupported or reserved for
this opcode and MUST reject before effects.
An FP16 or BF16 side is not microscaled and MUST omit its scale source. An
E4M3, E5M2, E2M1X2, or E1M2X2 side is microscaled and MUST provide one
E8M0 scale source. Providing a scale for an unscaled side or omitting a scale
for a scaled side is illegal. Consequently the canonical source sequence is
left matrix, optional left scale, right matrix, optional right scale, followed
by the explicit new Local destination; matrix DataTypes determine the sequence
unambiguously.
For a scaled left matrix, the scale valid shape is
M x ceil(K / 32). For a scaled right matrix, it is
ceil(K / 32) x N. Scale Tiles use row-major layout, their physical row and
column extents are powers of two containing the complete valid shape, and each
E8M0 element applies to the corresponding group of at most 32 K-dimension
matrix elements. An unscaled side behaves as if every scale factor were the
multiplicative identity; no implicit or materialized scale Tile exists.
Shared bindings remain source-only. They MAY supply the right matrix and its
required scale, or both matrices and whichever scales their DataTypes require.
An omitted scale has no Shared binding slot. The destination is always a new
Local FP32 private CUBE result, and TCVT remains the canonical conversion
boundary.
Decision 055: TGEMV is the Local-only TMATMUL specialization with M=1
BSTART.TGEMV is exactly the ordinary TMATMUL contract specialized to
M=1. LB0 is therefore fixed to one; canonical assembly omits it, while an
explicit LB0 is legal only when it resolves to one. LB1=N and LB2=K,
with the ordinary omission default one and nonzero power-of-two requirement.
The left source is a row vector with valid shape 1 x K, the right source is
a matrix with valid shape K x N, and the explicit new Local destination has
valid shape 1 x N. The vector and destination use row-major layout; the
right source follows the ordinary matrix layout and physical-capacity rules.
DataType selection, mixed same-class operands, FP32/S32/U32 private
result classes, B.DATR controls and defaults, full Core4 participation,
preflight, snapshots, and commit are otherwise unchanged from TMATMUL.
TGEMV is Local-only. B.IOS is illegal and no Shared operand form is
defined. PE_MASK=0000 is the strict no-op; every executing Local binding uses
1111. The canonical binding sequence is the 1 x K vector, the K x N
matrix, and the explicit new 1 x N Local destination.