mlcommons / mlcommons/inference
Better support from Loadgen to get optimal server QPS for the submitters
Open
@arav-agarwal2 is already working on this.
Since Mar 9, 2026.
postmortem 5.0
- Dominant language
- Python
- Stars
- 1.6k
- Forks
- 650
- Avg merge
- 1d 22h
- Merged PRs (30d)
- 6
Description
Options:
- Testing LoadGen's find peak performance mode
- Log additional advice related to possible increase of QPS
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.