The accuracy issue of MT bench
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 39.5k
- Forks
- 4.8k
- PR merge metrics
- No merged PRs in 30d
Description
I used the latest code to test the mt bench score of llama-2-chat, and the test result was only about 5.86. However, the official data provided was as high as around 6.3. For my own model, using the same response, the average difference between the two GPT4 scores was surprisingly about 0.2. Additionally, the issue in # 2659 seems to have not been resolved yet, and I am not sure if this is the cause of the error
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No file, test, or entry point is named. Start by reproducing the reported MT bench scores with the latest code and compare them with the official llama-2-chat result and the GPT-4 score differences; check issue #2659 for related context. Done means identifying the source of the discrepancy or confirming whether #2659 explains it.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100