LAION-AI / LAION-AI/Open-Assistant

Load test inference-server on different hardware

Open
#1,629 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

inference testing
Dominant language
Python
Stars
37.4k
Forks
3.3k
PR merge metrics
No merged PRs in 30d

Description

We want to test how many users the inference-server can serve and with what response times on setups with different numbers / types of GPUs & CPUs devices.

On the Stability AI cluster we can perform the load tests with up to 8, 16, 32, 128 pre-emptable GPUs.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The issue names the inference-server but no files, tests, or entry points. Start by locating the inference-server and any existing deployment or benchmarking entry points, then establish how hardware configurations and response-time results should be recorded. Done means load-test results cover the proposed GPU and CPU setups and report capacity and response times.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, backend, machine-learning, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.