Skip to content

Avoid float16 overflow in inverse_softplus - #765

Open
Asterisk-Hunter wants to merge 2 commits into
mllam:mainfrom
Asterisk-Hunter:fix/inverse-softplus-float16
Open

Asterisk-Hunter wants to merge 2 commits into
mllam:mainfrom
Asterisk-Hunter:fix/inverse-softplus-float16

Conversation

@Asterisk-Hunter

@Asterisk-Hunter Asterisk-Hunter commented Oct 5, 2026 •

Copy link
Copy Markdown

Describe your changes

Finite float16 inputs such as 12, 15, and 20 overflow the positive expm1 intermediate in inverse_softplus, producing infinite outputs and NaN gradients below its default threshold.

Use the equivalent negative-exponential expression while preserving the existing input clamp, linear branch, and dtype. Add regression coverage for beta 0.5, 1, and 2 below, at, and above the threshold. No new dependencies.

Validation:

  • pytest -vv -s --doctest-modules: 307 passed, 62 warnings, in the locked CPU development environment (Python 3.12, PyTorch 2.12.0).
  • pytest -q tests/test_utils.py: 9 passed.
  • pre-commit run --all-files: all hooks passed.
  • Original main reproduces non-finite outputs and gradients for all three beta cases.

Issue Link

Closes #764.

Type of change

  • 🐛 Bug fix
  • ✨ New feature
  • 💥 Breaking change
  • 📖 Documentation

Checklist before requesting a review

  • Branch is up to date with main.
  • Performed a self-review.
  • Existing docstrings and type annotations remain applicable.
  • Added a comment explaining the stable expression.
  • Added regression tests.
  • PR title uses imperative form.
  • Full test suite passes locally.
  • Maintainer reviewer and assignee selected.

README changes are not needed for this internal numerical fix.

Author checklist after completed review

Asterisk-Hunter and others added 2 commits October 5, 2026 15:49
Co-authored-by: OpenAI Codex <noreply@openai.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
@Asterisk-Hunter
Asterisk-Hunter marked this pull request as ready for review October 5, 2026 10:29
@nikhil3495

Copy link
Copy Markdown
Contributor

I checked this branch out locally and compared the old and new inverse_softplus.

tests/test_utils.py -k softplus: 9 passed.
float16: the old version returns inf for finite inputs like 12 (beta=1) and gives NaN gradients; the new one returns 12 with a gradient of ~1.0 and no inf/nan for any finite input I tried.
float32/float64: outputs differ from main only by rounding, and float32 gradients agree to ~2e-7 relative over ~2000 values, including points around the threshold and the lower clamp.
inf, nan, and very large inputs behave identically to main.

One thing I couldn't tell is how often this path actually runs in float16 under autocast, so I can't say how many users hit it. The change itself looks correct to me.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

inverse_softplus overflows for float16 inputs below its threshold

2 participants