Skip to content

Commit ef12e33

Browse files
committed
FEAT: Add LayaRefusalScorer, a local CPU refusal scorer
Adds a true/false refusal scorer that runs on the machine running PyRIT, with no API call and no per-response cost. The scorer puts the refusal question to Laya (Convai Innovations, Apache 2.0), an encoder that answers typed questions in one forward pass, takes the representation Laya forms of each answer option rather than its own verdict, and reads that with a logistic head trained on PyRIT's own human-labeled refusal rows. The head is trained on first use from the pinned datasets, so no opaque weights file ships and the whole model is reproducible from the repository at a pinned dataset and checkpoint state. Trained on refusal.csv and evaluated on refusal_extra.csv it is right on 96.9 percent of rows, and 91.4 percent the other way round. Excluding rows whose response text also appears in training gives 97.5 percent of 40 rows and 91.2 percent of 68 rows. SelfAskRefusalScorer with GPT-4o reaches 97 to 98 percent on the same rows. Laya's own verdict without the trained head agrees with the labels on 53 to 71 percent of rows depending on the order its answer options are presented in, so both orders are averaged and the trained head is what makes it usable. Scoring is two forward passes, about a second per response on CPU.
1 parent 58f4883 commit ef12e33

5 files changed

Lines changed: 842 additions & 1 deletion

File tree

‎doc/code/scoring/1_true_false_scorers.ipynb‎

Lines changed: 27 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -508,7 +508,33 @@
508508
"Do not apply a second chat wrapper to the bundled prompt. If using a chat server that\n",
509509
"supplies its own formatting, pass a matching `prompt_template` explicitly.\n",
510510
"\n",
511-
"All five need their respective endpoints/credentials even though they are not \"self-ask\"."
511+
"All five need their respective endpoints/credentials even though they are not \"self-ask\".\n",
512+
"\n",
513+
"## Local classifier scorers\n",
514+
"\n",
515+
"- **`LayaRefusalScorer`**: detects refusals on the machine running PyRIT, with no API call.\n",
516+
" It puts the refusal question to [Laya](https://huggingface.co/convaiinnovations/laya), an\n",
517+
" Apache 2.0 encoder that answers typed questions in one forward pass, takes the representation\n",
518+
" Laya forms of each answer option, and reads it with a logistic head trained on PyRIT's own\n",
519+
" human-labeled refusal rows. Scoring is two forward passes, roughly a second per response on\n",
520+
" CPU, with no per-response cost.\n",
521+
"\n",
522+
"Trained on one of PyRIT's two refusal datasets and evaluated on the other, it is right on\n",
523+
"96.9% and 91.4% of rows respectively (97.5% and 91.2% once rows whose response text also\n",
524+
"appears in training are excluded). `SelfAskRefusalScorer` with GPT-4o reaches 97-98% on the\n",
525+
"same rows. Laya's own verdict, used as it comes without the trained head, agrees with the\n",
526+
"labels on 53-71% of rows, so the training step is what makes it usable.\n",
527+
"\n",
528+
"The head trains itself on first use from the pinned datasets, which takes a few minutes of CPU\n",
529+
"once, and `abstain_band` returns an undetermined score for the uncertain tail so it can be\n",
530+
"routed to an LLM judge:\n",
531+
"\n",
532+
"```python\n",
533+
"from pyrit.score import LayaRefusalScorer\n",
534+
"\n",
535+
"scorer = LayaRefusalScorer() # pip install laya\n",
536+
"scores = await scorer.score_text_async(\"I'm sorry, I can't help with that.\")\n",
537+
"```"
512538
]
513539
},
514540
{

‎doc/code/scoring/1_true_false_scorers.py‎

Lines changed: 26 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -290,6 +290,32 @@
290290
# supplies its own formatting, pass a matching `prompt_template` explicitly.
291291
#
292292
# All five need their respective endpoints/credentials even though they are not "self-ask".
293+
#
294+
# ## Local classifier scorers
295+
#
296+
# - **`LayaRefusalScorer`**: detects refusals on the machine running PyRIT, with no API call.
297+
# It puts the refusal question to [Laya](https://huggingface.co/convaiinnovations/laya), an
298+
# Apache 2.0 encoder that answers typed questions in one forward pass, takes the representation
299+
# Laya forms of each answer option, and reads it with a logistic head trained on PyRIT's own
300+
# human-labeled refusal rows. Scoring is two forward passes, roughly a second per response on
301+
# CPU, with no per-response cost.
302+
#
303+
# Trained on one of PyRIT's two refusal datasets and evaluated on the other, it is right on
304+
# 96.9% and 91.4% of rows respectively (97.5% and 91.2% once rows whose response text also
305+
# appears in training are excluded). `SelfAskRefusalScorer` with GPT-4o reaches 97-98% on the
306+
# same rows. Laya's own verdict, used as it comes without the trained head, agrees with the
307+
# labels on 53-71% of rows, so the training step is what makes it usable.
308+
#
309+
# The head trains itself on first use from the pinned datasets, which takes a few minutes of CPU
310+
# once, and `abstain_band` returns an undetermined score for the uncertain tail so it can be
311+
# routed to an LLM judge:
312+
#
313+
# ```python
314+
# from pyrit.score import LayaRefusalScorer
315+
#
316+
# scorer = LayaRefusalScorer() # pip install laya
317+
# scores = await scorer.score_text_async("I'm sorry, I can't help with that.")
318+
# ```
293319
# %% [markdown]
294320
# ## Multimodal scorers
295321
#

‎pyrit/score/__init__.py‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -78,6 +78,7 @@
7878
from pyrit.score.true_false.decoding_scorer import DecodingScorer
7979
from pyrit.score.true_false.float_scale_threshold_scorer import FloatScaleThresholdScorer
8080
from pyrit.score.true_false.gandalf_scorer import GandalfScorer
81+
from pyrit.score.true_false.laya_refusal_scorer import LayaRefusalScorer
8182
from pyrit.score.true_false.llamaguard_parser import LLAMAGUARD_3_CATEGORY_CODES, parse_llamaguard_response
8283
from pyrit.score.true_false.llamaguard_policy import LlamaGuardCategory, LlamaGuardPolicy
8384
from pyrit.score.true_false.llamaguard_scorer import (
@@ -184,6 +185,7 @@
184185
"TraceAcquisitionError": "pyrit.score.observation.trace_client",
185186
"TraceClient": "pyrit.score.observation.trace_client",
186187
"JsonSchemaResponseHandler": "pyrit.score.response_handler",
188+
"LayaRefusalScorer": "pyrit.score.true_false.laya_refusal_scorer",
187189
"LDAPInjectionOutputScorer": "pyrit.score.true_false.regex.ldap_injection_output_scorer",
188190
"LikertScaleEvalFiles": "pyrit.score.float_scale.self_ask_likert_scorer",
189191
"LikertScale": "pyrit.score.float_scale.likert_scale",

0 commit comments

Comments
 (0)