Skip to main content

ADR 0085: Numeric post-process and format operations

Context​

ADR 0062 recorded a single repository-wide mnemonic audit. This record preserves the accepted decisions for this family as one decision-scoped owner. The former identifiers remain only in legacy_ids and the generated ADR index; current normative meaning is owned by the affected ASL/NDF clauses.

Decisions​

Decision 174: B.FPATR numeric post-processing is bit-exact architecture​

Every assigned nonzero B.FPATR.PreQuantMode, every assigned ReluMode, and every enabled RowMax or GroupMax result MUST have one bit-exact architectural result. Identical input encodings, parameters, dimensions, and numeric controls MUST produce identical destination and auxiliary-output encodings across all conforming implementations.

The matrix post-processing profile MUST define conversion order, parameter interpretation, rounding, saturation, overflow, exceptional values, output packing, activation, and auxiliary-reduction results. A nonzero assigned mode MUST NOT fall back to an identity transform or an implementation-selected numeric policy.

Decision 175: B.FPATR reductions precede destination conversion and activation​

Matrix RowMax and GroupMax consume the complete raw accumulator result before destination conversion or activation. MaxAbsEn and RowMaxInit affect this raw-accumulator reduction stage. PreQuant then converts only the primary D result, and the selected ReLU, scalar LReLU/PReLU, or vector PReLU operation selects the positive or negative path multiplier before the single destination conversion of each D element.

PreQuant and activation do not alter RowMaxOut or GroupMaxOut. Those auxiliary outputs retain the accumulator data type and format. The processed D and all enabled auxiliary outputs are prepared from the same pre-commit state and published as one atomic output group.

Decision 176: B.FPATR activation selects the pre-conversion multiplier​

For a nonnegative raw accumulator, the ordinary quantization scale is the selected multiplier. For a negative raw accumulator, no activation selects the ordinary quantization scale, ReluMode=1 selects zero, and scalar or per-column LReLU/PReLU selects its FP19 activation parameter. The selected multiplier is applied before the mode's intermediate and destination conversion; activation never decodes and re-encodes an already rounded destination value.

The scalar mode obtains one finite nonnegative FP19 parameter from the assigned dense B.IOR source. The vector mode obtains one such FP19 parameter per destination column from the assigned 1 x N Local parameter Tile. An activation parameter replaces, rather than multiplies, the negative-path quantization scale. NaN does not participate in an ordered comparison with zero.

Decision 177: B.FPATR scalar and vector quantization use multiplicative scales​

A scalar-parameter PreQuant mode multiplies every raw accumulator element by one scalar scale. A vector-parameter mode multiplies each element by the scale selected by its destination column from the assigned 1 x N parameter Tile. No assigned scale mode interprets the parameter as a divisor.

For an integer-output mode whose parameter carrier includes an offset, the scaled value is rounded and saturated to the assigned S5, S9, or S17 intermediate before adding that signed offset. Final destination encoding then uses the selected B.DATR.Sat clamp/wrap rule. The two S32-to-S16 shift modes perform the assigned signed arithmetic right shift and saturate the S16 result; they do not consume a multiplicative scale.

Decision 178: B.FPATR integer saturation control distinguishes clamp from wrap​

After every required intermediate saturation, B.DATR.Sat=1 clamps the result to the minimum or maximum representable destination value. B.DATR.Sat=0 does not clamp an ordinary finite overflow; it truncates the rounded integer to the destination element width, producing the corresponding modulo-2^N two's-complement or unsigned encoding.

Fixed shift modes already produce a saturated S16 result and reject an encoded Sat request. Scalar LReLU/PReLU and vector PReLU participate at the same pre-conversion intermediate point rather than performing a second destination encoding.

Decision 179: B.FPATR PreQuant codes retain one closed source, destination, and parameter table​

The assigned PreQuantMode table is:

CodeModeAccumulatorDestinationParameter
0NoQuantFP32, S32, or U32unchangednone
1F322F16FP32FP16none
2VREQ8S32S8per-column FP19 scale and signed 9-bit offset
3REQ8S32S8scalar FP19 scale and signed 9-bit offset
4VDEQF16S32FP16per-column FP19 scale
5DEQF16S32FP16scalar FP19 scale
12VSHIFTS322S16S32S16per-column shift code
13SHIFTS322S16S32S16scalar shift code
16F322BF16FP32BF16none
17REQ4S32S4X2scalar FP19 scale and signed 5-bit offset
18VREQ4S32S4X2per-column FP19 scale and signed 5-bit offset
19DEQS16S32S16scalar FP19 scale and signed 17-bit offset
20VDEQS16S32S16per-column FP19 scale and signed 17-bit offset
23VQF322B8_PREFP32S8per-column FP19 scale and signed 9-bit offset
24QF322B8_PREFP32S8scalar FP19 scale and signed 9-bit offset
25QF322HIF8_PREFP32HiF8scalar FP19 scale
26QF322FP8_PREFP32E4M3scalar FP19 scale
27QF322F32_PREFP32FP32scalar FP19 scale
28VQF322HIF8_PREFP32HiF8per-column FP19 scale
32QF322F16_PREFP32FP16scalar FP19 scale
33VQF322F16_PREFP32FP16per-column FP19 scale
34QF322BF16_PREFP32BF16scalar FP19 scale
35QS322BF16_PRES32BF16scalar FP19 scale
36VQF322BF16_PREFP32BF16per-column FP19 scale
37VQF322FP8_PREFP32E4M3per-column FP19 scale
38VQF322F32_PREFP32FP32per-column FP19 scale
39VQS322BF16_PRES32BF16per-column FP19 scale

Every other six-bit value is reserved. A nonzero mode used with a different accumulator class MUST reject before source snapshots, allocation, numeric status, or destination effects. The four-bit shift code represents an arithmetic right shift by one through sixteen bits.

Decision 180: B.FPATR parameters and special values are canonical​

FP19 uses one sign bit, an eight-bit exponent with bias 127, and a ten-bit fraction. It preserves signed zero and gradual subnormals and assigns IEEE-like infinity and NaN classes as values, but B.FPATR parameter legality is narrower. A quantization scale MUST be positive normal. An activation parameter MUST be positive zero or positive normal. Subnormal, infinite, NaN, negative, or nonzero-unused-bit carriers reject before effects. Scalar parameters use B.IOR; vector parameters use one row-major 1 x N U64 carrier Tile and select the element for the destination column.

Signed integer offsets are two's-complement values at their assigned 5-, 9-, or 17-bit width. For float-to-integer special values, Sat=0 uses the common destination indefinite encoding and records invalid status. Sat=1 converts NaN to zero and clamps positive or negative infinity to the corresponding destination endpoint; NaN records invalid and infinity records overflow plus inexact status. With Sat=1, floating NaN produces zero; with Sat=0, it produces the destination canonical quiet NaN. A signaling NaN additionally records invalid status in either case.

Floating finite results preserve subnormals with tininess detected after rounding. On floating overflow, Sat=1 returns the largest finite value with the input sign; Sat=0 returns signed infinity where the destination format has infinity, or the canonical quiet NaN for finite-only E4M3. Overflow and inexact status are recorded.

Decision 181: B.FPATR output carriers and numeric status publish atomically​

FP16, BF16, HiF8, E4M3, FP32, S16, and S8 use their architectural element encodings. Each S4X2 logical element occupies the low nibble of its model carrier and adjacent logical elements use the existing packed-memory nibble order. RowMaxOut and GroupMaxOut retain the raw accumulator data type, use row-major M x 1 and M x ceil(N/GroupN) shapes, and observe the fixed increasing-column reduction order.

All post-processing and reduction flags are accumulated before commit. The processed D payload, enabled auxiliary payloads, descriptors, and sticky numeric status are published together. A failed preflight, conversion, or allocation exposes none of them.

Decision 182: floating and scale formats expose exact finite decompositions​

Every assigned floating or scale Tile DataType has one exact descriptor for its carrier width, logical lane width, lanes per carrier, sign, exponent and fraction fields, exponent bias, constrained carrier bits, and supported special-value classes. Integer Tile DataTypes have no floating-format descriptor.

For every valid finite floating or scale encoding, the formal model returns an availability flag, sign, integer significand, and integer exponent whose exact value is (-1)^sign * UInt(significand) * 2^exponent. This decomposition uses only integers and bitvectors and performs no rounding. Invalid internal encodings, infinities, NaNs, and integer Tile DataTypes return unavailable.

TF32 and HF32 retain their required low-zero carrier constraints. E3M2 and E2M3 retain their required high-zero carrier constraints. Packed E2M1X2, E1M2X2, and HiF4X2 decompose one selected four-bit logical lane. E8M0 encodes 2^(raw-127) for raw values 0x00..0xFE; 0xFF is unavailable NaN, and a scale block contains 32 logical K elements.

Decision 183: TCVT to E8M0 rounds a positive base-two exponent​

The named hardware profile accepts exactly FP16, BF16, and FP32 as sources when TCVT selects an E8M0 destination. Other sources to E8M0 reject before destination allocation or payload effects. This restriction does not narrow other TCVT destination types.

For a positive finite source in the inclusive range 2^-127 through 2^127, TCVT rounds log2(source) to an integer exponent under the resolved RMode and writes exponent + 127. Exact powers of two set no status; other in-range values record inexact. RNE, RTM, RTP, RTZ, RNA, RTO, and RHB retain their architectural meanings in the exponent domain.

Positive finite values below 2^-127 underflow and values above 2^127 overflow. With Sat=1 they clamp to 0x00 or 0xFE; with Sat=0 they produce 0xFF. Underflow or overflow also records inexact. Positive infinity uses the overflow rule. Positive or negative zero, every negative value, and every NaN produce 0xFF and record invalid. Canonicalize keeps its existing private-CUBE-source representation role and does not change this value map.