FEAT explore default settings on targets used for generating red teaming prompts
- Dominant language
- Python
- Stars
- 4.5k
- Forks
- 893
- Avg merge
- 3d 50m
- Merged PRs (30d)
- 165
Description
#### Is your feature request related to a problem? Please describe.
We have default values (e.g., for temperature) on targets that most people just use as is without thinking twice about it. Are those values set reasonably? Do we get vastly better performance from the red teaming LLM if we change them? Note that this item primarily considers the red teaming targets, NOT the attack target (`prompt_target`), converter target, or the scoring target. Those could be spin-off tasks.
#### Describe the solution you'd like
These things need to be explored by defining a test set, a set of attacks, and a thorough evaluation (with retries due to randomness of responses).
Finally, the results should be captured, perhaps in a little post for the #362 (not yet built). If the default values need adjusting that should also be done.
#### Describe alternatives you've considered, if relevant
#### Additional context
https://x.com/simonw/status/1847514672490016800?s=46
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by locating the default settings for the red teaming targets, excluding the attack target (prompt_target), converter target, and scoring target. Define a test set and attacks, run a thorough evaluation with retries for response randomness, and capture the results in a post related to issue #362; adjust defaults only if the findings support it.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, security
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100