ggml-cpu.c: `src1_cont` undeclared → build fails on ARM CPUs with i8mm (Apple M2/M3/M4, Graviton3, ...)
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 40.3k
- Forks
- 3.7k
- PR merge metrics
- No merged PRs in 30d
Description
Summary
Compilation of 3rdparty/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c fails on any ARM CPU that has the i8mm extension when GGML_NATIVE=ON (the default used by setup_env.py). The I2_S fast path uses a variable that only exists when llamafile is enabled, but the file explicitly disables llamafile on i8mm/SVE targets. Independent of the model and of the OS (found on macOS; a Linux ARM host with i8mm, e.g. Graviton3 or Neoverse V1/N2, should hit the same line).
Environment
- MacBook Pro, Apple M2 Pro, macOS 26.5.2, Apple clang 21.0.0 (clang-2100.1.1.101), CMake 4.4.3
microsoft/BitNetat0b341e5(currentmain), submodule3rdparty/llama.cppat390c3077(branchrelease-bitnet-embedding-0.6b-270m)- Default
setup_env.pyconfiguration (-DBITNET_ARM_TL1=OFF), plus-DCMAKE_SHARED_LINKER_FLAGS=-Wl,-undefined,dynamic_lookupto get past the macOS link error of #611 / #595 (otherwise the build stops earlier and this error is never reached)
Error
3rdparty/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c:1495:40: error: use of undeclared identifier 'src1_cont'
1495 | const size_t src1_col_stride = src1_cont || src1->type != vec_dot_type ? i2s_row_size : nb11;
| ^~~~~~~~~
1 error generated.
Compile flags for that translation unit (from compile_commands.json): -mcpu=native+dotprod+i8mm+nosve+nosme ... -DGGML_USE_LLAMAFILE.
Root cause
ggml-cpu.c line 48:
#if defined(__ARM_FEATURE_SVE) || defined(__ARM_FEATURE_MATMUL_INT8)
#undef GGML_USE_LLAMAFILE
#endif
With i8mm enabled (-mcpu=native+...+i8mm, so __ARM_FEATURE_MATMUL_INT8 is defined), the llamafile block in ggml_compute_forward_mul_mat is compiled out. src1_cont is declared only inside that block (line 1361, #if GGML_USE_LLAMAFILE). The I2_S fast path added at lines 1491-1495 uses src1_cont unconditionally, so it only compiles when llamafile happens to be on: x86, and ARM without i8mm (e.g. Apple M1). Every ARM CPU with i8mm fails.
Suggested fix
Compute the flag locally where it is needed, e.g. in the I2_S block:
const bool src1_cont = ggml_is_contiguous(src1);
or hoist the existing declaration out of the #if GGML_USE_LLAMAFILE region.
Workarounds
-DCMAKE_C_FLAGS=-U__ARM_FEATURE_MATMUL_INT8 -DCMAKE_CXX_FLAGS=-U__ARM_FEATURE_MATMUL_INT8(same mechanism ggml itself uses when a feature check fails), or-DGGML_NATIVE=OFF -DGGML_CPU_ARM_ARCH=armv8.2-a+dotprod+fp16(build without i8mm).
Both make the build complete. Note that I2_S inference on the resulting arm64 binary is still wrong (garbage output even with the official BitNet-b1.58-2B-4T-gguf), which is the separate, already reported #585 / #600.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start in 3rdparty/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c, reading the feature guards near line 48 and the I2_S path in ggml_compute_forward_mul_mat around lines 1361 and 1491-1495. Build with the default setup_env.py configuration and GGML_NATIVE=ON on an i8mm ARM target; done means compilation completes without the src1_cont error.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- c, cmake
- Domain
- backend, build-system
- Issue type
- Bug
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Activity status
- Active
- Clarity
- Clearly specified
- Newbie friendliness
- 86/100