FIX: escape literal digits in DigitBijectionConverter to avoid token collision - #2741
Open
Mallika Kalangi (bhargavikvmpl-2001) wants to merge 2 commits into
Conversation
…collision DigitBijectionConverter maps each letter to a num_digits-digit token and joins adjacent tokens without separators, so a bare run of digits in the encoded text is read back as letters. A digit that was literal in the plaintext is passed through unchanged, which makes it indistinguishable from an encoded letter: decode() takes num_digits characters at a time and consumes whatever prefix of the literal number happens to be in the mapping. With the default num_digits=2 and seed=1, "abc 123 xyz" encodes to "278218 123 781191" and decodes back to "abc w3 xyz" -- "12" is w's token, so the literal 123 loses its first two digits and gains a letter. Across 40 seeds and 8 digit-bearing prompts, 44% of round-trips are corrupted at num_digits=2 and 4% at num_digits=3. Prompts that carry numbers are ordinary in practice: "CVE-2021-44228" becomes "CVE-2021-4rt", "pi is 3.14159" becomes "pi is 3.14c9". This is the same collision microsoft#2696 fixed for the literal apostrophe against _CASE_MARKER, in the other direction, and it reaches the target as well as PyRIT's own decoder: get_teaching_instructions said only to preserve "all other punctuation" and gave the model no rule for literal digits, so the target could not resolve the ambiguity either. Literal digits are now escaped with a tilde, which can never appear inside a token, so every unescaped digit run in the encoded text is a pure concatenation of tokens. A literal tilde is doubled, matching the existing treatment of a literal apostrophe. The escape binds to exactly one following character and is resolved before any token lookup, so no escaped digit can be re-read as part of a neighbouring token. Digit-free prompts encode exactly as before. get_teaching_instructions states both new rules and gains a worked example carrying a number, so the target applies the same scheme. Randomized round-trip fuzzing over an alphabet of letters, digits, apostrophes and tildes: 0 mismatches in 9000 cases across num_digits 2, 3 and 4. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Roman Lutz (romanlutz)
requested changes
Sep 22, 2026
| # before any digit run is considered for token lookup. | ||
| if encoded_text[i] == self._LITERAL_MARKER: | ||
| escaped = encoded_text[i + 1 : i + 2] | ||
| if escaped == self._LITERAL_MARKER or escaped in string.digits: |
Contributor
There was a problem hiding this comment.
Could we check that escaped is non-empty before consuming it? If a model response ends with a lone ~, the slice is "", and Python evaluates "" in string.digits as True. The decoder then silently drops that marker instead of preserving it through the existing fallback:
converter = DigitBijectionConverter(seed=42)
converter.decode("~") # Returns "" instead of "~"
converter.decode(converter.encode(prompt="top") + "~") # Returns "top" instead of "top~"This can happen with malformed or truncated responses. A guard would let an incomplete escape reach the fallback:
if escaped and (escaped == self._LITERAL_MARKER or escaped in string.digits):Please also add direct decoder cases for "~" and "~~~". The round-trip tests cannot catch this because the encoder always doubles literal tildes.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
DigitBijectionConvertercorrupts any prompt containing a literal digit. This is the same collision #2696 fixed for the literal apostrophe against_CASE_MARKER, in the other direction.Each letter maps to a
num_digits-digit token and adjacent tokens are joined without separators, so a bare digit run in the encoded text is read back as letters. A digit that was literal in the plaintext is passed through unchanged, which makes it indistinguishable from an encoded letter:decodetakesnum_digitscharacters at a time and consumes whatever prefix of the literal number happens to be in the mapping.With the default
num_digits=2andseed=1:"12"isw's token, so the literal123loses its first two digits and gains a letter.Across 40 seeds and 8 digit-bearing prompts:
num_digitsPrompts carrying numbers are ordinary in practice:
"CVE-2021-44228"decodes to"CVE-2021-4rt","pi is 3.14159"to"pi is 3.14c9".It reached the target as well as PyRIT's own decoder.
get_teaching_instructionssaid only to preserve "all other punctuation" and gave the model no rule for literal digits, so the target could not resolve the ambiguity either.Fix
Literal digits are escaped with a tilde, which can never appear inside a token, so every unescaped digit run in the encoded text is a pure concatenation of tokens. A literal tilde is doubled, matching the existing treatment of a literal apostrophe. The escape binds to exactly one following character and is resolved before any token lookup, so an escaped digit can never be re-read as part of a neighbouring token.
get_teaching_instructionsstates both new rules and gains a worked example carrying a number, so the target applies the same scheme.Digit-free prompts encode byte-identically to before — the three existing exact-string tests are unchanged.
Two choices worth a reviewer's opinion
~as rare in prompts and free of the markdown/JSON escaping hazards that`or\carry. The choice is arbitrary and easy to change.AsciiSmugglerConverterfix for BUG: AsciiSmugglerConverter silently drops non-ASCII characters #2539 chose to raise loudly on input it could not represent. I went the other way here because numbers are common in red-team objectives and FIX: escape literal apostrophes in DigitBijectionConverter to avoid case-marker collision #2696 set the precedent of escaping a colliding literal rather than refusing it. Happy to switch to a loud failure if that is preferred.LetterBijectionConverterandTokenBijectionConverterare unaffected — letters map to single letters there, so no multi-character token can absorb a neighbour.Tests and Documentation
Five tests added to
tests/unit/converter/test_bijection_converter.py:"abc 123 xyz"->"101112 ~1~2~3 333435")num_digits2/3/4 x 12 seeds, so the guarantee does not rest on one lucky mappingget_teaching_instructionscovers the literal-digit ruleBeyond the suite, randomized round-trip fuzzing over an alphabet of letters, digits, apostrophes and tildes found 0 mismatches in 9000 cases across
num_digits2, 3 and 4. The same sweep before the fix reported 44% corruption at the default width.Full runs:
tests/unit/converter/1682 passed.ruff checkandruff format --checkclean.No documentation or notebook changes — the converter's behaviour on digit-free input is unchanged and no doc sample encodes a prompt with digits, so JupyText was not run.
🤖 Generated with Claude Code