mlcommons / mlcommons/inference
Selection criteria of models for benchmark
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 1.6k
- Forks
- 650
- Avg merge
- 1d 22h
- Merged PRs (30d)
- 6
Description
Hello,
I have a genuinely curious question. How did you select the models for the inference benchmark? There are many models available, and the MLPerf Datacenter inference benchmark is conducted on the following models:
- resnet50-v1.5
- retinanet 800x800
- bert
- dlrm-v2
- 3d-unet
- gpt-j
- stable-diffusion-xl
- llama2-70b
- llama3.1-405b
- mixtral-8x7b
- rgat
- pointpainting
Is there a particular reason why these models were chosen? Thank you for any pointers or guidance.
Thanks.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no files, tests, or entry points. Start by reviewing the benchmark model list and the MLPerf Datacenter inference models cited in the issue, then look for existing documentation on selection criteria. Done means documenting a clear rationale for the selected models or pointing to the authoritative guidance.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- machine-learning
- Domain
- machine-learning, testing-qa
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100