Skip to content

FIX: validate n-gram size and reference text in PlagiarismScorer - #2972

Open
RohithPariki wants to merge 2 commits into
microsoft:mainfrom
RohithPariki:fix/validate-plagiarism-scorer-inputs
Open

RohithPariki wants to merge 2 commits into
microsoft:mainfrom
RohithPariki:fix/validate-plagiarism-scorer-inputs

Conversation

@RohithPariki

Copy link
Copy Markdown

Fixes #2971

Problem

In pyrit/score/float_scale/plagiarism_scorer.py:

  1. PlagiarismScorer accepted non-positive integer values for n (such as n=0, n=-1, or floats like 2.5, booleans like False/True).
    • When n=0, _ngram_set returned {()} because range(len(tokens) + 1) iterates over empty slices tokens[i:i]. For PlagiarismMetric.JACCARD, the intersection and union of {()} with itself evaluate to identical sets, yielding 1.0 (100% plagiarism) for completely unrelated text.
    • When n < 0 (e.g. n=-1), Python negative slicing produced spurious sets causing unintended overlap scores.
    • Non-integer or boolean types for n bypassed initialization and failed during scoring or generated invalid calculations (False acts as 0).
  2. reference_text was not validated during initialization. If initialized with an empty string "", whitespace-only " ", or strings without word tokens (e.g. "!@#$%"), self._tokenize(reference_text) produces an empty list (reference_len=0), silently returning 0.0 for all candidate responses even if they are identical copies of the reference text.
  3. metric was not validated at construction time to be an instance of PlagiarismMetric.

Root Cause

  • __init__ in PlagiarismScorer did not validate reference_text, metric, or n before assigning them.
  • _ngram_set(tokens, n) assumed n >= 1.
  • _plagiarism_score relied on reference_len == 0 early return instead of enforcing valid reference tokens upon scorer instantiation.

Solution

  1. Input Validation in PlagiarismScorer.__init__:
    • Validates reference_text is a non-empty string and strips to ensure non-whitespace.
    • Tokenizes reference_text with self._tokenize and ensures it contains at least one word token. Raises ValueError with clear guidance otherwise.
    • Validates metric is an instance of PlagiarismMetric.
    • Validates n is an integer (isinstance(n, int) and not isinstance(n, bool)) and >= 1. Raises ValueError otherwise.
  2. Defensive Validation in _plagiarism_score:
    • Enforces n >= 1 and not a boolean.
    • Enforces valid metric object or value.
  3. Comprehensive Unit Tests:
    • Added parameterized tests covering invalid reference_text (empty, whitespace, non-string, punctuation-only / no word tokens).
    • Added parameterized tests covering invalid n (0, negative numbers, floats, booleans, non-integer types).
    • Added boundary tests for valid integer n values (1, 2, 5, 10).
    • Added tests ensuring invalid metrics are rejected with informative error messages.

Testing

  • Ran pytest on all PlagiarismScorer test suites:
    python -m pytest tests/unit/score/test_plagiarism_scorer.py -q
    # 79 passed
  • Linting and formatting checks:
    python -m ruff check pyrit/score/float_scale/plagiarism_scorer.py tests/unit/score/test_plagiarism_scorer.py
    python -m ruff format --check pyrit/score/float_scale/plagiarism_scorer.py tests/unit/score/test_plagiarism_scorer.py
    All checks passed.

Risk

Very low. This strictly hardens input validation on scorer construction and helper boundaries against silent incorrect metrics or false-positive 1.0 scores without modifying the core similarity logic for valid inputs.

@RohithPariki
RohithPariki force-pushed the fix/validate-plagiarism-scorer-inputs branch from 0b1d474 to 1cc3ddc Compare October 3, 2026 13:52

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

BUG: PlagiarismScorer accepts invalid n-gram size and blank reference text

1 participant