microsoft / microsoft/agent-learning

Conduct Red Team and Adversarial Testing of Agent Learning

Open
#28 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
10
Forks
9
Avg merge
1d 19h
Merged PRs (30d)
4

Description

Summary

Perform a structured security evaluation of Agent Learning to identify vulnerabilities, abuse paths, and learning manipulation risks.

Areas of Focus
Reward hacking
Policy poisoning
Episode tampering
Prompt injection
Tool abuse
Data poisoning
Adversarial evaluator behavior
Deliverables
Threat model
Attack catalog
Findings report
Mitigation recommendations
Acceptance Criteria
Threat model completed
Adversarial test suite created
Security findings documented
Mitigation backlog generated

Contributor guide

No contributing guide indexed for this repository

Research direction

No files, tests, or entry points are identified in the issue, so first survey the repository to locate the Agent Learning implementation and existing test structure. Define the threat model and adversarial test scope around reward hacking, policy poisoning, episode tampering, prompt injection, tool abuse, data poisoning, and evaluator behavior; done means the listed acceptance criteria are delivered.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, security, testing
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.