microsoft / microsoft/agent-learning
Conduct Red Team and Adversarial Testing of Agent Learning
- Dominant language
- Python
- Stars
- 10
- Forks
- 9
- Avg merge
- 1d 19h
- Merged PRs (30d)
- 4
Description
Summary
Perform a structured security evaluation of Agent Learning to identify vulnerabilities, abuse paths, and learning manipulation risks.
Areas of Focus
Reward hacking
Policy poisoning
Episode tampering
Prompt injection
Tool abuse
Data poisoning
Adversarial evaluator behavior
Deliverables
Threat model
Attack catalog
Findings report
Mitigation recommendations
Acceptance Criteria
Threat model completed
Adversarial test suite created
Security findings documented
Mitigation backlog generated
Contributor guide
No contributing guide indexed for this repository
Research direction
No files, tests, or entry points are identified in the issue, so first survey the repository to locate the Agent Learning implementation and existing test structure. Define the threat model and adversarial test scope around reward hacking, policy poisoning, episode tampering, prompt injection, tool abuse, data poisoning, and evaluator behavior; done means the listed acceptance criteria are delivered.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, security, testing
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100