mlcommons / mlcommons/modelbench
Just getting started with this framework, not sure what to do next....
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 134
- Forks
- 36
- Avg merge
- 1d 11h
- Merged PRs (30d)
- 17
Description
Congrats on release v1.0.0 last week!
I ran through the modelbench README.md. I cloned the repo, set everything up, and executed the "Running Your First Benchmark" without issue. I then followed the step under "Using the Journal" and saw the results for prompt id airr_practice_1_0_41321, and the scores for the various models. I think I understand what the sample prompt is, and what it is trying to accomplish. Is there any further documentation on airr_practice_1_0_41321 ?
Now I am not sure what to do next. Are there other prompts or prompt suites I can test in a similar matter? I did checkout the ML Common website, and the white paper, but I am not sure what else I can execute besides the practice prompt.
Would be willing to contribute to the README.md for next steps if someone can guide me on what my options are to proceed with executing other prompts.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with README.md, especially the “Running Your First Benchmark” and “Using the Journal” sections, and review the references to prompt id airr_practice_1_0_41321. A completed contribution would define the available next steps and document how to proceed beyond the practice prompt, once those options are established.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- documentation
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100