open-compass / open-compass/opencompass

[Feature] 请问在使用VLLM测评模型humaneval时,batch_size 不同导致 测评结果有区别是为什么?

Open
#1,097 0 comments 0 reactions 1 assignee View on GitHub

@tonysy is already working on this.

Since Apr 26, 2024.

Dominant language
Python
Stars
7.5k
Forks
869
Avg merge
17h 52m
Merged PRs (30d)
13

Description

Describe the feature

image
image
在batch_size 分别为128,64,16的情况下,deepseek 1.3B 的P@1 分别是31.71、30.49、29.27
请问这是为什么?

Will you implement it?
  • I would like to implement this feature and create a PR!

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.