aws-samples / aws-samples/multiagent-collab-scenario-benchmark
How to Evaluate Custom Multi-Agent Systems with Your Benchmarks?
Open
- Dominant language
- Python
- Stars
- 18
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Description
Hi! Thanks a lot for the awesome repo.
I'm interested in using your benchmarks to test some multi-agent frameworks, and I was wondering if you could provide some examples or guidance on how to build a multi-agent system and generate conversations that can be evaluated and fairly compared with the results in your paper https://arxiv.org/pdf/[2412.05449v1](https://arxiv.org/pdf/2412.05449v1) .
For example, how should I initiate system messages for the user and for different agents?
Thanks in advance for your help!
Contributor guide
Assessment
This issue has not been assessed yet.