alibaba / alibaba/InferSim

Future plans and the project general idea

Open
#7 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
100
Forks
30
Avg merge
3d 14h
Merged PRs (30d)
1

Description

Hello, thank you for your great work. I really liked this project and would gladly contribute to it's development. Looks like that now the project is mainly focused on multi-gpu setups (correct me if I'm wrong). Thus the question is: is it a priority to support more lite setups. It would be cool to add the support of cases where one card is being used to launch several models (and therefore not 100% of memory is available for this model), because as I can see KV Cache now occupies all the remaining memory. And also I'd like to see more graphic cards available (I'm personally interested in A100). What do you think about all of this, let me know if it's not a priority or if you have another view on this project.

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no files, tests, or entry points. Start by reviewing the project’s current GPU and KV-cache support to determine how multi-model single-card use and additional cards such as A100 fit the simulator. The scope and acceptance criteria need maintainer agreement before implementation can begin.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.