microsoft / microsoft/BitNet

[Bug]: I2_S GEMM fast path produces garbage for multi-token prompts on AVX-only CPUs

Open
#617 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
40.3k
Forks
3.7k
PR merge metrics
No merged PRs in 30d

Description

Description

On CPUs without AVX2 (e.g., Intel Xeon E5-2690 v2, Ivy Bridge, AVX-only), the I2_S GEMM fast path in ggml-cpu.c:1492 produces corrupt output when processing prompts with more than ~3 tokens. Single-token generation (GEMV path) works correctly.

Symptoms

  • 1-3 token prompts → coherent output
  • 5+ token prompts → ?????? or garbled output
  • The corruption affects the KV cache: even after the prompt is processed, subsequent generation tokens are garbled

Root Cause

The GEMM fast path (ggml_gemm_i2_i8_s) is called for multi-token prompt evaluation (when src1 has multiple columns). The scalar fallback implementation has a bug in how it indexes the I2_S weight matrix and/or activation matrix for column strides > 1.

Workaround

Disabling the GEMM fast path forces I2_S through the dequantize-then-float-matmul path:

// ggml-cpu.c:1492 — change from:
if (src0->type == GGML_TYPE_I2_S && ggml_n_dims(src0) == 2) {
// to:
if (false && src0->type == GGML_TYPE_I2_S && ggml_n_dims(src0) == 2) {

This produces correct results but is slower (~0.6 tok/s prompt eval vs ~26 tok/s for F16 on the same hardware).

Key Distinction from #547

This is distinct from #547/PR #580 which covers the empty-body scalar fallback for ggml_vec_dot_i2_i8_s_* kernels. After applying those fixes (or equivalent scalar implementations), the GEMM path still produces wrong results for multi-token prompts.

The GEMV path (line ~1217) works correctly for single-token generation. The issue is specifically in ggml_gemm_i2_i8_s when called with nr > 1 (multiple activation columns).

Reproduction

# Build with AVX2 disabled (forces scalar fallback)
cmake -B build -DBITNET_ARM_TL1=OFF -DBITNET_X86_TL2=OFF
cmake --build build --target llama-server -j8

# Short prompt works
curl http://localhost:8081/v1/completions \
  -d '{"model":"bitnet","prompt":"Hello","max_tokens":20}'
# → " there! I'm happy to help" ✓

# Longer prompt fails
curl http://localhost:8081/v1/completions \
  -d '{"model":"bitnet","prompt":"The weather today is","max_tokens":20}'
# → "?????" ✗

Also requires fixes #588 (ReLU²) and PR #616 (weight scale direction) for coherent F16 output.

Environment

  • CPU: Intel Xeon E5-2690 v2 (Ivy Bridge, AVX only, no AVX2)
  • OS: Ubuntu 24.04 LTS
  • Compiler: Clang 18.1.3
  • CMake flags: -DBITNET_ARM_TL1=OFF -DBITNET_X86_TL2=OFF
  • Model: BitNet-b1.58-2B-4T (I2_S format)
  • BitNet commit: 390c30775

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in ggml-cpu.c at the I2_S GEMM fast path around line 1492 and inspect ggml_gemm_i2_i8_s for column-stride handling when nr > 1. Build with AVX2 disabled using the documented CMake flags, then reproduce with short and longer prompts and compare the GEMM path with the working GEMV path around line 1217. Done means multi-token I2_S prompts produce coherent output on AVX-only CPUs without disabling the fast path.

Written by the indexing model from the issue text.

Assessment

Tech stack
cmake, cpp
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
68/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.