microsoft / microsoft/BitNet

Apple Silicon Metal + BLAS segfault for I2_S when ubatch >= 32 routes generic MUL_MAT

Open Beginner friendly
#512 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
40.3k
Forks
3.7k
PR merge metrics
No merged PRs in 30d

Description

Environment

  • Platform: Apple Silicon Mac
  • Host: Apple M4 Max
  • OS: macOS
  • Compiler: Homebrew clang 18.1.8
  • BitNet / submodule state: BitNet using vendored 3rdparty/llama.cpp at Eddie-Wang1120/llama.cpp commit 1f86f058de0c3f4098dedae2ae8653c335c868a1
  • Model: microsoft/BitNet-b1.58-2B-4T-gguf / ggml-model-i2_s.gguf
  • Build flags:
    • GGML_METAL=ON
    • GGML_ACCELERATE=ON
    • GGML_BLAS=ON
    • GGML_BLAS_VENDOR=Apple
    • BITNET_ARM_TL1=OFF

Problem

On Apple Silicon with Metal enabled, i2_s inference can segfault when BLAS is enabled and the physical micro-batch crosses the BLAS routing threshold.

The crash is tied to physical ubatch, not logical batch:

  • -b 2048 -ub 31 -> stable
  • -b 32 -ub 31 -> stable
  • -b 2048 -ub 32 -> segfault
  • -b 2048 -ub 512 -> segfault

This means the failure starts exactly when the BLAS backend begins claiming the generic MUL_MAT path for larger batches.

Control Experiment

The same Metal runtime is stable when BLAS is disabled:

  • BLAS ON + -b 2048 -ub 512 -> segfault
  • BLAS OFF + -b 2048 -ub 512 -> stable

This strongly suggests the crash is in the BLAS-side handling of GGML_TYPE_I2_S, not in Metal itself and not in the outer chat request schema.

Root Cause

ggml-blas.cpp allows the generic BLAS MUL_MAT path to accept quantized source tensors when ggml_get_type_traits(src0->type)->to_float != NULL.

For GGML_TYPE_I2_S, that is not safe:

  • I2_S stores an external scale outside the per-row payload
  • the generic BLAS dequantize-to-float path assumes self-contained per-row data
  • once ubatch >= 32, BLAS starts claiming MUL_MAT
  • that eventually crashes in the i2_s dequant / BLAS matmul path

In crash reports, the top frames consistently land in:

  • dequantize_row_i2_s
  • ggml_backend_blas_mul_mat

Proposed Fix

Reject GGML_TYPE_I2_S in the generic BLAS MUL_MAT support check so that I2_S continues using its specialized non-BLAS path:

return src0->type != GGML_TYPE_I2_S &&
       ggml_is_contiguous(src0) &&
       ggml_is_contiguous(src1) &&
       src1->type == GGML_TYPE_F32 &&
       (ne0 >= min_batch && ne1 >= min_batch && ne10 >= min_batch) &&
       (src0->type == GGML_TYPE_F32 || ggml_get_type_traits(src0->type)->to_float != NULL);

Result After Patch

After applying the BLAS guard above:

  • BLAS ON + Metal + -b 2048 -ub 512 is stable
  • managed broker end-to-end requests no longer segfault under the same settings

This does not solve all i2_s quality issues, but it does remove the native crash path.

Related Issues

  • #468
  • #470
  • #195
  • #411

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in ggml-blas.cpp at the generic BLAS MUL_MAT support check, then reproduce the issue with Metal and Apple BLAS using i2_s and ubatch values of 31 and 32 or higher. Confirm that the I2_S guard preserves the specialized path and that the larger-ubatch configuration no longer segfaults while BLAS remains enabled.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
backend
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
72/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.