Problem
FST stores an x87 value to memory in a smaller floating-point format. It must round the
value using the rounding mode in the x87 control word. This also applies to subnormal
results: nonzero values below the smallest normal value in the destination format.
The source helper shifts the significand, the bits holding the value's
significant digits, to fit the subnormal range. It then omits the rounding step at that
new boundary. This can store a value that is too small.
The example below should store 0x00400001 with rounding toward positive infinity, but
the source stores 0x00400000. It also omits the underflow flag and reports the wrong
C1 rounding flag. It correctly sets the precision flag for this input.
Reproducing case
The opcode notation D9 /2 means opcode byte D9 with the register-selection field of
the following ModR/M byte set to 2; the other fields select the memory address.
Use 64-bit mode with x87 enabled: CR0.EM=0 and CR0.TS=0 allow these instructions to
execute with x87 enabled. ST(0) is the top x87 register and
ST(1) is the next. The control word selects rounding and which exceptions are masked.
“Masked” means the instruction uses the defined fallback behavior for that exception.
Store raw extended value 3f808000010000000000 to a 32-bit floating-point memory
destination with FST m32fp (D9 /2) and control word 0x0b7f (round toward positive
infinity, exceptions masked). Start with ST(0) nonempty, no pending exceptions, and a
writable four-byte destination.
The source returns 0x00400000, says it did not round up, and omits the underflow flag.
Its existing discarded-bit check can still set precision. With round-to-nearest
(0x037f), the source returns 0x00400000 and still omits underflow.
Each 20-digit hexadecimal input specifies the exact 80-bit x87 register contents.
Source checked: Intel SDM executable specification revision
d307f89f742765865b87c5d4d23f552b3c72e871.
Separate hardware test
A separate hardware test on an AMD EPYC-Milan processor returned 0x00400001,
reported rounded-up, and set underflow plus precision (0x30) for the directed-rounding
case. With round-to-nearest (0x037f), it returned 0x00400000. The directed-rounding
result agrees with the Intel manual's destination-width rounding requirement.
Manual reference
References use Intel SDM 325462-089US, October 2025.
Intel SDM Volume 2A, FST/FSTP, page 3-379 (PDF page 1075) says
that memory stores round the significand to the destination width
using the x87 control word. Volume 1, sections 8.5.5–8.5.6,
pages 8-29–8-30 (PDF pages 237–238) describes tiny and inexact results and their
exception flags. Section 4.9.1.5, page 4-23 (PDF page 113)
requires masked underflow to be reported when the result is tiny and inexact.
Proposed fix
After shifting into the destination's subnormal range, recompute the bits needed for
rounding: the last retained bit, the first discarded bit (the round bit), and whether
any later discarded bit is nonzero (the sticky bit). Apply the selected rounding mode,
including a carry into the smallest normal result. Set underflow and precision from the
tiny and discarded bits according to the existing mask rules.
AI disclosure
Assisted-by: Codex
Codex assisted with source analysis, test review, and drafting this report.
Source links
Current helper pages:
FP87::Convert_To_Float.
Problem
FST stores an x87 value to memory in a smaller floating-point format. It must round the
value using the rounding mode in the x87 control word. This also applies to subnormal
results: nonzero values below the smallest normal value in the destination format.
The source helper shifts the significand, the bits holding the value's
significant digits, to fit the subnormal range. It then omits the rounding step at that
new boundary. This can store a value that is too small.
The example below should store
0x00400001with rounding toward positive infinity, butthe source stores
0x00400000. It also omits the underflow flag and reports the wrongC1 rounding flag. It correctly sets the precision flag for this input.
Reproducing case
The opcode notation
D9 /2means opcode byte D9 with the register-selection field ofthe following ModR/M byte set to 2; the other fields select the memory address.
Use 64-bit mode with x87 enabled: CR0.EM=0 and CR0.TS=0 allow these instructions to
execute with x87 enabled. ST(0) is the top x87 register and
ST(1) is the next. The control word selects rounding and which exceptions are masked.
“Masked” means the instruction uses the defined fallback behavior for that exception.
Store raw extended value
3f808000010000000000to a 32-bit floating-point memorydestination with FST m32fp (
D9 /2) and control word0x0b7f(round toward positiveinfinity, exceptions masked). Start with ST(0) nonempty, no pending exceptions, and a
writable four-byte destination.
The source returns
0x00400000, says it did not round up, and omits the underflow flag.Its existing discarded-bit check can still set precision. With round-to-nearest
(
0x037f), the source returns0x00400000and still omits underflow.Each 20-digit hexadecimal input specifies the exact 80-bit x87 register contents.
Source checked: Intel SDM executable specification revision
d307f89f742765865b87c5d4d23f552b3c72e871.Separate hardware test
A separate hardware test on an AMD EPYC-Milan processor returned
0x00400001,reported rounded-up, and set underflow plus precision (
0x30) for the directed-roundingcase. With round-to-nearest (
0x037f), it returned0x00400000. The directed-roundingresult agrees with the Intel manual's destination-width rounding requirement.
Manual reference
References use Intel SDM 325462-089US, October 2025.
Intel SDM Volume 2A, FST/FSTP, page 3-379 (PDF page 1075) says
that memory stores round the significand to the destination width
using the x87 control word. Volume 1, sections 8.5.5–8.5.6,
pages 8-29–8-30 (PDF pages 237–238) describes tiny and inexact results and their
exception flags. Section 4.9.1.5, page 4-23 (PDF page 113)
requires masked underflow to be reported when the result is tiny and inexact.
Proposed fix
After shifting into the destination's subnormal range, recompute the bits needed for
rounding: the last retained bit, the first discarded bit (the round bit), and whether
any later discarded bit is nonzero (the sticky bit). Apply the selected rounding mode,
including a carry into the smallest normal result. Set underflow and precision from the
tiny and discarded bits according to the existing mask rules.
AI disclosure
Assisted-by: Codex
Codex assisted with source analysis, test review, and drafting this report.
Source links
Current helper pages:
FP87::Convert_To_Float.