Reject unexpected spellings of NaN and Infinity in Python JSON parsing - #30443
Open
ManoharPaturi wants to merge 1 commit into
Open
ManoharPaturi wants to merge 1 commit into
ManoharPaturi wants to merge 1 commit into
Conversation
Python's float() accepts many spellings of NaN and Infinity beyond the three canonical ones the proto3 JSON mapping allows, so the pure Python json_format silently accepted inputs like "NAN", "+nan", "inf", "INF", "+Infinity" or " NaN " that every other implementation rejects. Reject them so a Python backend can not be fed values that would fail to parse elsewhere.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The proto3 JSON mapping allows exactly three spellings for non finite float values: "NaN", "Infinity" and "-Infinity". The pure Python json_format delegated quoted values to float(), which accepts a much wider family of spellings, so inputs like "NAN", "+nan", "inf", "INF", "+Infinity" or " NaN " were silently accepted here while every other implementation rejects them. A backend that validates JSON in Python therefore accepts payloads that a C++, Java or upb frontend would refuse, which is a quiet validation gap rather than a crash.
This adds an explicit check in _ConvertFloat: when the stripped and lowercased value is one of the nan or inf word family, it must be exactly one of the three canonical spellings, otherwise ParseError is raised. Unquoted JSON numbers are unaffected because they arrive as float instances, and the float and double paths share this function, so both are covered.
Behavior before the change, observed with the pure Python implementation:
After the change all three raise ParseError, while "NaN", "Infinity" and "-Infinity" continue to round trip through MessageToJson and Parse.
Testing: extended testInvalidFloatValue in python/google/protobuf/internal/json_format_test.py with fourteen rejected spellings and updated the expectation for "nan", whose message now names all three canonical forms. The full json_format_test.py suite passes under PROTOCOL_BUFFERS_PYTHON_IMPLEMENTATION=python (112 tests locally).