Skip to content

FIX: escape literal digits in DigitBijectionConverter to avoid token collision - #2741

Open
Mallika Kalangi (bhargavikvmpl-2001) wants to merge 2 commits into
microsoft:mainfrom
bhargavikvmpl-2001:fix/digit-bijection-literal-digit-collision
Open

Mallika Kalangi (bhargavikvmpl-2001) wants to merge 2 commits into
microsoft:mainfrom
bhargavikvmpl-2001:fix/digit-bijection-literal-digit-collision

Conversation

@bhargavikvmpl-2001

Copy link
Copy Markdown
Contributor

Description

DigitBijectionConverter corrupts any prompt containing a literal digit. This is the same collision #2696 fixed for the literal apostrophe against _CASE_MARKER, in the other direction.

Each letter maps to a num_digits-digit token and adjacent tokens are joined without separators, so a bare digit run in the encoded text is read back as letters. A digit that was literal in the plaintext is passed through unchanged, which makes it indistinguishable from an encoded letter: decode takes num_digits characters at a time and consumes whatever prefix of the literal number happens to be in the mapping.

With the default num_digits=2 and seed=1:

"abc 123 xyz"  ->  encode  ->  "278218 123 781191"  ->  decode  ->  "abc w3 xyz"

"12" is w's token, so the literal 123 loses its first two digits and gains a letter.

Across 40 seeds and 8 digit-bearing prompts:

num_digits round-trips corrupted
2 (default) 44%
3 4%
4 0% in sample

Prompts carrying numbers are ordinary in practice: "CVE-2021-44228" decodes to "CVE-2021-4rt", "pi is 3.14159" to "pi is 3.14c9".

It reached the target as well as PyRIT's own decoder. get_teaching_instructions said only to preserve "all other punctuation" and gave the model no rule for literal digits, so the target could not resolve the ambiguity either.

Fix

Literal digits are escaped with a tilde, which can never appear inside a token, so every unescaped digit run in the encoded text is a pure concatenation of tokens. A literal tilde is doubled, matching the existing treatment of a literal apostrophe. The escape binds to exactly one following character and is resolved before any token lookup, so an escaped digit can never be re-read as part of a neighbouring token.

get_teaching_instructions states both new rules and gains a worked example carrying a number, so the target applies the same scheme.

Digit-free prompts encode byte-identically to before — the three existing exact-string tests are unchanged.

Two choices worth a reviewer's opinion

LetterBijectionConverter and TokenBijectionConverter are unaffected — letters map to single letters there, so no multi-character token can absorb a neighbour.

Tests and Documentation

Five tests added to tests/unit/converter/test_bijection_converter.py:

  • literal number round-trips, asserting the exact encoded form ("abc 123 xyz" -> "101112 ~1~2~3 333435")
  • literal tilde round-trips as a doubled marker
  • parametrized round-trip over 6 digit-bearing prompts x num_digits 2/3/4 x 12 seeds, so the guarantee does not rest on one lucky mapping
  • get_teaching_instructions covers the literal-digit rule

Beyond the suite, randomized round-trip fuzzing over an alphabet of letters, digits, apostrophes and tildes found 0 mismatches in 9000 cases across num_digits 2, 3 and 4. The same sweep before the fix reported 44% corruption at the default width.

Full runs: tests/unit/converter/ 1682 passed. ruff check and ruff format --check clean.

No documentation or notebook changes — the converter's behaviour on digit-free input is unchanged and no doc sample encodes a prompt with digits, so JupyText was not run.

🤖 Generated with Claude Code

…collision

DigitBijectionConverter maps each letter to a num_digits-digit token and joins
adjacent tokens without separators, so a bare run of digits in the encoded text
is read back as letters. A digit that was literal in the plaintext is passed
through unchanged, which makes it indistinguishable from an encoded letter:
decode() takes num_digits characters at a time and consumes whatever prefix of
the literal number happens to be in the mapping.

With the default num_digits=2 and seed=1, "abc 123 xyz" encodes to
"278218 123 781191" and decodes back to "abc w3 xyz" -- "12" is w's token, so
the literal 123 loses its first two digits and gains a letter. Across 40 seeds
and 8 digit-bearing prompts, 44% of round-trips are corrupted at num_digits=2
and 4% at num_digits=3. Prompts that carry numbers are ordinary in practice:
"CVE-2021-44228" becomes "CVE-2021-4rt", "pi is 3.14159" becomes "pi is 3.14c9".

This is the same collision microsoft#2696 fixed for the literal apostrophe against
_CASE_MARKER, in the other direction, and it reaches the target as well as
PyRIT's own decoder: get_teaching_instructions said only to preserve "all other
punctuation" and gave the model no rule for literal digits, so the target could
not resolve the ambiguity either.

Literal digits are now escaped with a tilde, which can never appear inside a
token, so every unescaped digit run in the encoded text is a pure concatenation
of tokens. A literal tilde is doubled, matching the existing treatment of a
literal apostrophe. The escape binds to exactly one following character and is
resolved before any token lookup, so no escaped digit can be re-read as part of
a neighbouring token. Digit-free prompts encode exactly as before.

get_teaching_instructions states both new rules and gains a worked example
carrying a number, so the target applies the same scheme.

Randomized round-trip fuzzing over an alphabet of letters, digits, apostrophes
and tildes: 0 mismatches in 9000 cases across num_digits 2, 3 and 4.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
# before any digit run is considered for token lookup.
if encoded_text[i] == self._LITERAL_MARKER:
escaped = encoded_text[i + 1 : i + 2]
if escaped == self._LITERAL_MARKER or escaped in string.digits:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we check that escaped is non-empty before consuming it? If a model response ends with a lone ~, the slice is "", and Python evaluates "" in string.digits as True. The decoder then silently drops that marker instead of preserving it through the existing fallback:

converter = DigitBijectionConverter(seed=42)
converter.decode("~")  # Returns "" instead of "~"
converter.decode(converter.encode(prompt="top") + "~")  # Returns "top" instead of "top~"

This can happen with malformed or truncated responses. A guard would let an incomplete escape reach the fallback:

if escaped and (escaped == self._LITERAL_MARKER or escaped in string.digits):

Please also add direct decoder cases for "~" and "~~~". The round-trip tests cannot catch this because the encoder always doubles literal tildes.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants