Skip to main content

ADR 0079: CUBE and matrix operations

Context​

ADR 0062 recorded a single repository-wide mnemonic audit. This record preserves the accepted decisions for this family as one decision-scoped owner. The former identifiers remain only in legacy_ids and the generated ADR index; current normative meaning is owned by the affected ASL/NDF clauses.

Decisions​

Decision 049: TMATMUL dimensions are M, N, and K in LB order​

BSTART.TMATMUL uses LB0=M, LB1=N, and LB2=K. Omission of any one of these fields supplies the value one for that field. An explicitly encoded or otherwise resolved value of zero is illegal. Each of M, N, and K MUST be a nonzero power of two.

The left source has valid shape M x K, the right source has valid shape K x N, and the destination has valid shape M x N. Each physical Tile row and column extent MUST be a power of two, and the physical extent MUST contain the complete valid rectangle. Shape, capacity, and compatibility checks MUST complete before destination allocation or any other effect.

Decision 050: TMATMUL supports mixed same-class input types and fixed accumulator classes​

The BSTART.TMATMUL DataType selects the left source type. An optional B.DATR.DataType selects the right source type; if that field is absent, the right source type defaults to the left source type. Absence is distinct from an explicitly encoded zero, because encoded DataType zero denotes FP64, which is not supported by TMATMUL. The block schema MUST therefore preserve field presence when applying this default.

The supported floating input set is FP32, TF32, HF32, FP16, BF16, HiF8, E4M3, E5M2, E3M2, E2M3, E2M1X2, and E1M2X2. The supported signed input set is S16, S8, and S4X2. The supported unsigned input set is U16, U8, and U4X2. The two inputs MAY use different types within the same floating, signed, or unsigned class. A cross-class pair is illegal.

FP64, E8M0, HiF4X2, S64, S32, U64, U32, every globally reserved DataType encoding, and every other type not listed above are reserved or unsupported for ordinary TMATMUL and MUST be rejected before effects.

A floating pair produces an FP32 accumulator result, a signed pair produces an S32 accumulator result, and an unsigned pair produces a U32 accumulator result. This result uses the private CUBE output representation. TCVT, not TMATMUL, performs conversion to a canonical public representation. A numeric profile MAY define format-specific arithmetic details, but it MUST preserve these operand classes, accumulator classes, and encoded numeric controls.

Decision 051: TMATMUL binding, participation, and supplementary fields are closed​

The Local form binds a left Local source, a right Local source, and an explicit new Local destination. The Shared forms bind either a Local left source plus a Shared right source, or a Shared left source plus a Shared right source; the destination remains an explicit new Local allocation. B.IOS is source-only for TMATMUL and reads only fully published Shared values.

PE_MASK=0000 is a strict no-op before descriptor access, readiness checks, allocation, lifetime consumption, faults, or effects. Every executing Local or Shared binding MUST use PE_MASK=1111; a nonzero partial mask is illegal before effects. All source payloads are snapshotted after complete preflight and before destination allocation or commit, so source-destination aliasing observes the old source values and the destination becomes visible only as one complete result.

TMATMUL does not consume mathematical B.IOR operands. Its B.DATR fields are closed to DataType, RMode, and Sat. An omitted DataType applies the right-type default in Decision 050 in ADR-0079, omitted RMode selects RNE, and omitted Sat selects disabled saturation. Layout, CMode, PadValueOrByteId, and Canonicalize MUST be zero. The private CUBE result remains noncanonical until an explicit TCVT operation.

The earlier statement that TMATMUL did not consume B.FPATR is superseded by ADR 0064. Every Matrix CUBE bundle contains exactly one B.FPATR; its all-zero form selects no post-processing, and nonzero modes determine the additional scalar, Local source, and Local destination schema.

Decision 052: TMATMUL.BIAS adds one 1 x N right-side broadcast source​

BSTART.TMATMUL.BIAS inherits the complete dimension, input-type, accumulator, binding, participation, supplementary-field, preflight, and commit contract of BSTART.TMATMUL. It adds one Bias source after the left and right matrix sources.

The Bias source valid shape MUST be exactly 1 x N, its layout MUST be row-major, and its DataType MUST equal the result accumulator class selected by Decision 050 in ADR-0079: FP32, S32, or U32. Its payload uses the same private CUBE result representation as that accumulator class. For every output row i and column j, Bias[0,j] is added once to the complete dot product for output D[i,j]. No row broadcast, scalar broadcast, full-matrix Bias, or Bias addition inside the K reduction is defined.

Bias remains a Local source in both the all-Local and Shared-matrix forms. Shared bindings MAY supply the right matrix or both matrix operands exactly as defined for TMATMUL; they do not bind the Bias or destination. The complete matrix and Bias sources are snapshotted after preflight, and the explicit new Local destination becomes visible only as one complete result.

Decision 053: TMATMUL.ACC uses explicit Local accumulator input and destination​

BSTART.TMATMUL.ACC inherits the complete dimension, matrix input-type, accumulator-class, binding, participation, supplementary-field, preflight, and commit contract of BSTART.TMATMUL. It adds one explicit Local accumulator source C before the left and right matrix sources and writes one explicit new Local destination D.

C MUST have valid shape M x N, the same row-major layout and physical capacity as D, and the same private CUBE accumulator representation and DataType as the result: FP32, S32, or U32. The result is D = C + A x B, with the encoded RMode and Sat controls applied according to the selected numeric profile. There is no implicit ACC operand or implicit destination.

The operation snapshots C, A, and B after complete preflight and before writing D. C and D MAY resolve to the same Local Tile; this case has read-old/write-new behavior and becomes visible only as one complete result. In Shared-matrix forms, Shared bindings MAY supply the right matrix or both matrix operands, but C and D remain explicit Local bindings.

Decision 054: TMATMULMX scales each matrix side independently​

BSTART.TMATMULMX inherits TMATMUL dimension ordering, matrix and destination shapes, explicit Local destination, Core4 participation, private FP32 result, numeric controls, preflight, snapshot, and atomic commit rules. Its matrix inputs are floating only. Each side independently uses one of FP16, BF16, E4M3, E5M2, E2M1X2, or E1M2X2; the two sides MAY use different listed types. Every other type, including HiF4X2, is unsupported or reserved for this opcode and MUST reject before effects.

An FP16 or BF16 side is not microscaled and MUST omit its scale source. An E4M3, E5M2, E2M1X2, or E1M2X2 side is microscaled and MUST provide one E8M0 scale source. Providing a scale for an unscaled side or omitting a scale for a scaled side is illegal. Consequently the canonical source sequence is left matrix, optional left scale, right matrix, optional right scale, followed by the explicit new Local destination; matrix DataTypes determine the sequence unambiguously.

For a scaled left matrix, the scale valid shape is M x ceil(K / 32). For a scaled right matrix, it is ceil(K / 32) x N. Scale Tiles use row-major layout, their physical row and column extents are powers of two containing the complete valid shape, and each E8M0 element applies to the corresponding group of at most 32 K-dimension matrix elements. An unscaled side behaves as if every scale factor were the multiplicative identity; no implicit or materialized scale Tile exists.

Shared bindings remain source-only. They MAY supply the right matrix and its required scale, or both matrices and whichever scales their DataTypes require. An omitted scale has no Shared binding slot. The destination is always a new Local FP32 private CUBE result, and TCVT remains the canonical conversion boundary.

Decision 055: TGEMV is the Local-only TMATMUL specialization with M=1​

BSTART.TGEMV is exactly the ordinary TMATMUL contract specialized to M=1. LB0 is therefore fixed to one; canonical assembly omits it, while an explicit LB0 is legal only when it resolves to one. LB1=N and LB2=K, with the ordinary omission default one and nonzero power-of-two requirement.

The left source is a row vector with valid shape 1 x K, the right source is a matrix with valid shape K x N, and the explicit new Local destination has valid shape 1 x N. The vector and destination use row-major layout; the right source follows the ordinary matrix layout and physical-capacity rules. DataType selection, mixed same-class operands, FP32/S32/U32 private result classes, B.DATR controls and defaults, full Core4 participation, preflight, snapshots, and commit are otherwise unchanged from TMATMUL.

TGEMV is Local-only. B.IOS is illegal and no Shared operand form is defined. PE_MASK=0000 is the strict no-op; every executing Local binding uses 1111. The canonical binding sequence is the 1 x K vector, the K x N matrix, and the explicit new 1 x N Local destination.