kvcache-ai / kvcache-ai/ktransformers

🚨FAQs | 常见问题🚨

Open
#1,608 20 comments 2 reactions 0 assignees View on GitHub
good first issue
Dominant language
Python
Stars
19.5k
Forks
1.6k
Avg merge
19h 32m
Merged PRs (30d)
27

Description

> [!NOTE]
> Please **avoid** creating issues regarding the following questions, as they might be closed without a response.
> 请**避免**创建与下述问题有关的 issues,这些 issues 可能不会被回复。

> [!TIP]
> Documentation: https://kvcache-ai.github.io/ktransformers/
> 中文文档:https://github.com/kvcache-ai/ktransformers/tree/main/doc/zh

----
### Most of problems / 大多数问题
#### How to install kt-kernel/ 怎么安装 kt-kernel

Please update repository and install again following the link below:
请参考文档进行安装,以及常见问题指导:
https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/README.md
#### Inference with kt-kernel+Sglang/ 在 kt-kernel + Sglang 上推理
##### 1. Install SGLang

```bash
git clone https://github.com/sgl-project/sglang.git
cd sglang
pip install -e "python[all]"
```

##### 2. Prepare Weights

You need both GPU weights and CPU weights for heterogeneous inference:

**GPU Weights:** Use the original / quantized model weights.

**CPU Weights:** Quantize to AMX-optimized format using the conversion script:

```bash
python scripts/convert_cpu_weights.py \
--input-path /path/to/model \
--input-type bf16 \ # Depends on your GPU weights type: fp8, fp16, or bf16
--output /path/to/cpu-weights \
--quant-method int8 # or int4
```

**Supported input formats:** FP8, FP16, BF16 → INT4/INT8.

For more details, see:
- [CPU Weights conversion](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/README.md#cpu-weights-for-cold-experts-on-cpu)
- [GPU Weights quantization](https://github.com/kvcache-ai/ktransformers/blob/main/kt-kernel/README.md#gpu-weights-for-hot-experts-on-gpu)

**Note:** LLAMAFILE backend supports GGUF format directly, but this feature is still in preview.

##### 3. Launch SGLang Server

Start the SGLang server with your normal SGLang parameters, and add the following KT-Kernel specific parameters to enable CPU-GPU heterogeneous inference:

**KT-Kernel Parameters to Add:**
- `--kt-method`: Backend method (AMXINT4, AMXINT8, or LLAMAFILE)
- `--kt-weight-path`: Path to the converted CPU weights
- `--kt-cpuinfer`: Number of CPU inference threads (set to physical cores)
- `--kt-threadpool-count`: Number of thread pools (set to NUMA node count)
- `--kt-num-gpu-experts`: Number of experts to keep on GPU
- `--kt-max-deferred-experts-per-token`: Deferred experts for pipelined execution

Example:
```bash
python -m sglang.launch_server \
[your normal SGLang parameters...] \
--kt-method AMXINT8 \
--kt-weight-path /path/to/cpu-weights \
--kt-cpuinfer 64 \
--kt-threadpool-count 2 \
--kt-num-gpu-experts 32 \
--kt-max-deferred-experts-per-token 2
```
----

### Support models/ 支持的模型列表

| Model_name/ 模型名称 | have_tested/ 是否测试过 |
|----------------------------------------|------------|
| GLM-4.5-Air | :white_check_mark: |
| GLM-4.5 | :white_check_mark: |
| Qwen3-30B-A3B | :white_check_mark: |
| Qwen3-235B-A22B-Thinking-2507 | :white_check_mark: |
| Qwen3-235B-A22B-Instruct-2507 | :white_check_mark: |
| Qwen3-Next-80B-A3B-Thinking | :white_check_mark: |
| DeepSeek-R1-0528 | :white_check_mark: |

>[!NOTE]
In principle, we can support any model that SGLang supports. If a model cannot run, then SGLang running purely on GPU likely won't be able to either. So if you test additional models that work, feel free to share them.

----

### Known Issues & Tips / 已知问题与提示

#### "KT fallback" log messages are normal / "KT fallback" 日志是正常的

If you see logs like `KT fallback: layer X prepare weight = XXms, compute = XXms`, this is **expected behavior** indicating that Layer-wise Prefill is working correctly. The "fallback" naming is misleading and will be updated in a future release.

如果看到类似 `KT fallback: layer X prepare weight = XXms` 的日志,这是**预期行为**,说明 Layer-wise Prefill 已正常启动。"fallback" 这个名称有误导性,后续版本会优化。

#### Low GPU utilization during prefill / Prefill 阶段 GPU 利用率低

If your GPU utilization is low during prefill, enable Layer-wise Prefill by adding:

```bash
--kt-gpu-prefill-token-threshold 2048
```

This allows the GPU to participate in prefill computation, significantly improving prefill throughput.

如果 prefill 阶段 GPU 利用率很低,请添加 `--kt-gpu-prefill-token-threshold 2048` 参数开启 Layer-wise Prefill。

#### GPTQ_INT4 + tensor-parallel-size > 1 / GPTQ_INT4 与多卡 TP 不兼容

GPTQ-Int4 quantized MoE models may fail with `IndexError: start out of range` when using `tensor-parallel-size > 1` with GPU experts. This is an SGLang-side weight loader issue. Workarounds:
- Use `--tensor-parallel-size 1`
- Use FP8 weights instead of GPTQ-Int4

GPTQ-Int4 量化的 MoE 模型在 `tensor-parallel-size > 1` 且使用 GPU experts 时可能报错。这是 SGLang 侧的问题。建议降低 TP 或使用 FP8 权重。

----

> [!TIP]
> If the problems still exist with the latest code, please create an issue.
若使用最新的代码仍然无法解决问题,请创建一个 issue。

Contributor guide

Open the contributing guide

Research direction

Start with the FAQ content in issue #1608 and compare it with kt-kernel/README.md and the linked documentation, including doc/zh. Clarify which documentation change is requested, then update the relevant FAQ material and verify that the installation, SGLang, model-support, and known-issue links and instructions are accurate.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
documentation
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.