MegEngine / MegEngine/InferLLM

是否有计划优化GPU上的推理加速

Open
#14 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
752
Forks
94
PR merge metrics
No merged PRs in 30d

Description

目前社区LLM采用主流GPTQ量化之后,量化层的kernel实现基本是负向优化,是否有计划支持GPU上量化后的模型推理加速。

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing the existing GPTQ quantization-layer kernel implementation and GPU inference path; the issue does not name files or tests. Benchmark quantized inference against the current implementation and define done as a measurable GPU speedup rather than a slowdown.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
machine-learning, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.