microsoft / microsoft/BitNet

ggml-cpu.c: `src1_cont` undeclared → build fails on ARM CPUs with i8mm (Apple M2/M3/M4, Graviton3, ...)

Open Beginner friendly
#618 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
40.3k
Forks
3.7k
PR merge metrics
No merged PRs in 30d

Description

Summary

Compilation of 3rdparty/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c fails on any ARM CPU that has the i8mm extension when GGML_NATIVE=ON (the default used by setup_env.py). The I2_S fast path uses a variable that only exists when llamafile is enabled, but the file explicitly disables llamafile on i8mm/SVE targets. Independent of the model and of the OS (found on macOS; a Linux ARM host with i8mm, e.g. Graviton3 or Neoverse V1/N2, should hit the same line).

Environment
  • MacBook Pro, Apple M2 Pro, macOS 26.5.2, Apple clang 21.0.0 (clang-2100.1.1.101), CMake 4.4.3
  • microsoft/BitNet at 0b341e5 (current main), submodule 3rdparty/llama.cpp at 390c3077 (branch release-bitnet-embedding-0.6b-270m)
  • Default setup_env.py configuration (-DBITNET_ARM_TL1=OFF), plus -DCMAKE_SHARED_LINKER_FLAGS=-Wl,-undefined,dynamic_lookup to get past the macOS link error of #611 / #595 (otherwise the build stops earlier and this error is never reached)
Error
3rdparty/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c:1495:40: error: use of undeclared identifier 'src1_cont'
 1495 |         const size_t src1_col_stride = src1_cont || src1->type != vec_dot_type ? i2s_row_size : nb11;
      |                                        ^~~~~~~~~
1 error generated.

Compile flags for that translation unit (from compile_commands.json): -mcpu=native+dotprod+i8mm+nosve+nosme ... -DGGML_USE_LLAMAFILE.

Root cause

ggml-cpu.c line 48:

#if defined(__ARM_FEATURE_SVE) || defined(__ARM_FEATURE_MATMUL_INT8)
#undef GGML_USE_LLAMAFILE
#endif

With i8mm enabled (-mcpu=native+...+i8mm, so __ARM_FEATURE_MATMUL_INT8 is defined), the llamafile block in ggml_compute_forward_mul_mat is compiled out. src1_cont is declared only inside that block (line 1361, #if GGML_USE_LLAMAFILE). The I2_S fast path added at lines 1491-1495 uses src1_cont unconditionally, so it only compiles when llamafile happens to be on: x86, and ARM without i8mm (e.g. Apple M1). Every ARM CPU with i8mm fails.

Suggested fix

Compute the flag locally where it is needed, e.g. in the I2_S block:

const bool src1_cont = ggml_is_contiguous(src1);

or hoist the existing declaration out of the #if GGML_USE_LLAMAFILE region.

Workarounds
  • -DCMAKE_C_FLAGS=-U__ARM_FEATURE_MATMUL_INT8 -DCMAKE_CXX_FLAGS=-U__ARM_FEATURE_MATMUL_INT8 (same mechanism ggml itself uses when a feature check fails), or
  • -DGGML_NATIVE=OFF -DGGML_CPU_ARM_ARCH=armv8.2-a+dotprod+fp16 (build without i8mm).

Both make the build complete. Note that I2_S inference on the resulting arm64 binary is still wrong (garbage output even with the official BitNet-b1.58-2B-4T-gguf), which is the separate, already reported #585 / #600.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in 3rdparty/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c, reading the feature guards near line 48 and the I2_S path in ggml_compute_forward_mul_mat around lines 1361 and 1491-1495. Build with the default setup_env.py configuration and GGML_NATIVE=ON on an i8mm ARM target; done means compilation completes without the src1_cont error.

Written by the indexing model from the issue text.

Assessment

Tech stack
c, cmake
Domain
backend, build-system
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
86/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.