Skip to content

[General]: FST truncates subnormal results without rounding them #42

Description

@shuklaayush

Problem

FST stores an x87 value to memory in a smaller floating-point format. It must round the
value using the rounding mode in the x87 control word. This also applies to subnormal
results: nonzero values below the smallest normal value in the destination format.

The source helper shifts the significand, the bits holding the value's
significant digits, to fit the subnormal range. It then omits the rounding step at that
new boundary. This can store a value that is too small.

The example below should store 0x00400001 with rounding toward positive infinity, but
the source stores 0x00400000. It also omits the underflow flag and reports the wrong
C1 rounding flag. It correctly sets the precision flag for this input.

Reproducing case

The opcode notation D9 /2 means opcode byte D9 with the register-selection field of
the following ModR/M byte set to 2; the other fields select the memory address.

Use 64-bit mode with x87 enabled: CR0.EM=0 and CR0.TS=0 allow these instructions to
execute with x87 enabled. ST(0) is the top x87 register and
ST(1) is the next. The control word selects rounding and which exceptions are masked.
“Masked” means the instruction uses the defined fallback behavior for that exception.

Store raw extended value 3f808000010000000000 to a 32-bit floating-point memory
destination with FST m32fp (D9 /2) and control word 0x0b7f (round toward positive
infinity, exceptions masked). Start with ST(0) nonempty, no pending exceptions, and a
writable four-byte destination.

The source returns 0x00400000, says it did not round up, and omits the underflow flag.
Its existing discarded-bit check can still set precision. With round-to-nearest
(0x037f), the source returns 0x00400000 and still omits underflow.

Each 20-digit hexadecimal input specifies the exact 80-bit x87 register contents.

Source checked: Intel SDM executable specification revision
d307f89f742765865b87c5d4d23f552b3c72e871.

Separate hardware test

A separate hardware test on an AMD EPYC-Milan processor returned 0x00400001,
reported rounded-up, and set underflow plus precision (0x30) for the directed-rounding
case. With round-to-nearest (0x037f), it returned 0x00400000. The directed-rounding
result agrees with the Intel manual's destination-width rounding requirement.

Manual reference

References use Intel SDM 325462-089US, October 2025.

Intel SDM Volume 2A, FST/FSTP, page 3-379 (PDF page 1075) says
that memory stores round the significand to the destination width
using the x87 control word. Volume 1, sections 8.5.5–8.5.6,
pages 8-29–8-30 (PDF pages 237–238)
describes tiny and inexact results and their
exception flags. Section 4.9.1.5, page 4-23 (PDF page 113)
requires masked underflow to be reported when the result is tiny and inexact.

Proposed fix

After shifting into the destination's subnormal range, recompute the bits needed for
rounding: the last retained bit, the first discarded bit (the round bit), and whether
any later discarded bit is nonzero (the sticky bit). Apply the selected rounding mode,
including a carry into the smallest normal result. Set underflow and precision from the
tiny and discarded bits according to the existing mask rules.

AI disclosure

Assisted-by: Codex

Codex assisted with source analysis, test review, and drafting this report.

Source links

Current helper pages:
FP87::Convert_To_Float.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions