microsoft / microsoft/PyRIT

Proposal: DecayDeltaScorer — a turn-over-turn rate-of-change scorer for multi-turn attacks

Open
#2,619 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
4.5k
Forks
893
Avg merge
3d 50m
Merged PRs (30d)
165

Description

## Motivation

PyRIT has strong multi-turn attack orchestration (including Crescendo) and a `ConversationScorer` that evaluates conversation context by re-scoring the accumulated history. However, I haven't found a scorer that explicitly captures the **rate of risk change between turns** without repeatedly evaluating the full conversation.

I also searched the repository for related concepts such as `delta`, `rate_of_change`, `turning_point`, and `decay` across `pyrit/score` and `pyrit/executor`, but didn't find an existing equivalent. Happy to be pointed to an existing implementation if I've missed one.

## Proposed approach

I'd like to propose a `DecayDeltaScorer`, implementing `FloatScaleScorer` / `MessageFloatScaleScorer` and following the wrapping pattern used by `FloatScaleThresholdScorer`.

Rather than thresholding the wrapped scorer's output, it would calculate the change in score between consecutive turns:

```text
Δ(t) = S(t) − S(t−1)
```

where `S(t)` is the wrapped scorer's score for the current turn and `S(t−1)` is its score for the previous turn in the same conversation.

The goal is for this to be a **deterministic, lightweight complement to `ConversationScorer`**, rather than a replacement:

* `ConversationScorer`: "How risky is the conversation at this point?"
* `DecayDeltaScorer`: "How quickly is the risk changing?"

This could be particularly useful for gradual, multi-turn escalation such as Crescendo attacks, where individual turns may remain below a detection threshold while the overall trajectory is consistently increasing.

## Implementation question

This is the main reason I'm opening an issue before submitting a PR.

I traced `get_scores()` in `memory_interface.py`. It supports filtering by score type/category/timestamp/scorer identifier, but I don't see a direct conversation-level filter for retrieving the wrapped scorer's previous score.

Two possible approaches I see are:

1. Re-score previous turns retrieved through `get_conversation_messages()`.
2. Retrieve persisted scores using the message-piece IDs associated with the conversation.

I'd appreciate guidance on which approach better fits PyRIT's existing architecture and performance expectations before I proceed with an implementation.

## Background

This proposal is motivated by my research into trajectory-level detection of gradual multi-turn attacks. In particular, I've observed cases where per-turn monitoring provides little or no pre-critical signal, while monitoring the trajectory reveals a consistent increase in risk.

I'm happy to share additional methodology or benchmark details if they would be useful for evaluating the proposed scoring approach.

If the approach fits PyRIT's scoring architecture, I'd be happy to follow up with a PR.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading memory_interface.py, especially get_scores() and get_conversation_messages(), then compare the wrapping patterns of FloatScaleScorer, MessageFloatScaleScorer, and FloatScaleThresholdScorer. Resolve whether prior turns should be re-scored or persisted scores retrieved by message-piece IDs, and establish the expected behavior for consecutive-turn deltas before implementation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, security
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.