lm-sys / lm-sys/FastChat

Inquiry on GPU Performance Benchmarks for faschat Models

Open
#2,699 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
39.5k
Forks
4.8k
PR merge metrics
No merged PRs in 30d

Description

I'm currently using an RTX 3090 for the OpenChat 7B model and am considering upgrading to a more powerful GPU. My goal is to use the same model but achieve higher generation speeds. Before making this investment, I seek more information about the potential performance improvements.

Would you be able to provide, or direct me to, any benchmarking results for generation speed across different GPUs when using large models in the faschat repository? Even if you have only conducted simple tests or have approximate information about the extent to which a more powerful GPU could enhance the speed of the process, I would be very interested in that information.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Review the FastChat repository for any existing benchmark material covering the OpenChat 7B model, large models, GPU types, or generation speed. A complete result would provide or link to comparable generation-speed measurements across GPUs, including the RTX 3090 context raised in the issue.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
documentation, machine-learning, performance
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.