ByteDance-Seed / ByteDance-Seed/decoupleQ

batchsize不同情况的推理

Open
#18 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Cuda
Stars
153
Forks
10
PR merge metrics
No merged PRs in 30d

Description

1.想问一下在使用run_inference_llama.sh时,是否可以设置不同batchsize
2.paper里的结果,batchsize都设置的多少呢?在batchsize较大情况下(eg.batchsize=16、32、64)的throughput您是否测试过呢

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with run_inference_llama.sh to determine whether batch size can be configured and how inference settings are passed. Check the paper's reported batch size and whether throughput results for batch sizes 16, 32, and 64 are available; done means providing those answers or documenting the limits of the available evidence.

Written by the indexing model from the issue text.

Assessment

Tech stack
shell
Domain
machine-learning
Issue type
Documentation
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.