microsoft / microsoft/BitNet

convert-hf-to-gguf-bitnet.py --outtype i2_s silently writes F16 (not I2_S) for LlamaForCausalLM BitNet checkpoints (Falcon3 / Falcon-E 1.58bit)

Open
#621 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
40.3k
Forks
3.7k
PR merge metrics
No merged PRs in 30d

Description

Summary

utils/convert-hf-to-gguf-bitnet.py accepts --outtype i2_s for every architecture, but the I2_S packing code path exists only in BitnetModel (architecture BitNetForCausalLM). For the LlamaForCausalLM BitNet checkpoints listed as supported in the README (tiiuae/Falcon3-*-1.58bit, tiiuae/Falcon-E-*), LlamaModel.write_tensors unpacks the offline-quantized ternary weights and stores them as F16, without any warning. The result is a valid but full-size GGUF (14.9 GB for Falcon3-7B instead of ~2.7 GB) with no I2_S tensors, while the metadata still claims an I2_S file type.

Since llama-quantize in this tree has no I2_S ftype either (#619), there is currently no way to obtain an I2_S GGUF for the Falcon 1.58-bit models from their HF checkpoints.

Environment
  • microsoft/BitNet at 0b341e5 (current main), submodule 3rdparty/llama.cpp at 390c3077
  • macOS 26.5.2, Apple M2 Pro, Python 3.14.7, gguf installed from 3rdparty/llama.cpp/gguf-py (as done by setup_env.py)
Steps to reproduce
hf download tiiuae/Falcon3-7B-Instruct-1.58bit --local-dir models/Falcon3-7B-Instruct-1.58bit
python utils/convert-hf-to-gguf-bitnet.py models/Falcon3-7B-Instruct-1.58bit --outtype i2_s

Converter log (every linear weight):

INFO:hf-to-gguf:blk.0.ffn_down.weight,       torch.uint8 --> F16, shape = {23040, 3072}
INFO:hf-to-gguf:blk.0.attn_q.weight,         torch.uint8 --> F16, shape = {3072, 3072}
...
INFO:gguf.gguf_writer:models/Falcon3-7B-Instruct-1.58bit/ggml-model-i2_s.gguf: n_tensors = 255, total_size = 14.9G

Resulting file (read with gguf.GGUFReader): tensor types {F16: 198, F32: 57}, zero I2_S tensors, general.file_type = 40. The same converter run on microsoft/bitnet-b1.58-2B-4T-bf16 correctly produces 210 I2_S tensors, so the difference is the architecture, not the input format (the Falcon checkpoint is offline-quantized: uint8 packed weights + weight_scale tensors, 196 of them, which LlamaModel does unpack correctly).

Root cause
  • LlamaModel.write_tensors (utils/convert-hf-to-gguf-bitnet.py, from line 776) handles the offline-quantized weights (unpack at ~line 800, scale_map), but its quantization dispatch only has TL1 and TL2 branches (lines ~869-878) followed by else: # default to float16 for quantized tensors (~line 880).
  • The I2_S branch (quantize_to_i2_s(data, override_scale=...)) exists only in BitnetModel.write_tensors (line 1164).
  • ftype_map / --outtype (lines 1216, 1237) accept i2_s regardless of the model class, so the request is silently downgraded to F16.
Suggested fix

Port the I2_S branch from BitnetModel.write_tensors to LlamaModel.write_tensors (the ternary values and scale_map are already available there, so quantize_to_i2_s(data, override_scale=scale) is a small change), or make the converter fail loudly when --outtype i2_s is requested for a class that cannot produce it. Related: #619 (conversion flow / missing I2_S in llama-quantize), #550 (Falcon3 TL2 support in setup_env.py), #616 (LlamaModel dequantization), #620 (general.file_type value).

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in utils/convert-hf-to-gguf-bitnet.py, comparing LlamaModel.write_tensors with BitnetModel.write_tensors, especially the quantization dispatch and scale_map handling. Reproduce the Falcon3 conversion with --outtype i2_s, then verify the GGUF tensor types and metadata with gguf.GGUFReader. Done means the requested output is I2_S or the converter clearly rejects unsupported architectures instead of writing F16 silently.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, tooling
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
76/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.