microsoft / microsoft/BitNet

Format conversion issue after downstream SFT

Open
#236 9 comments 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
40.3k
Forks
3.7k
PR merge metrics
No merged PRs in 30d

Description

I used my own dataset to finetune the model bitnet-b1.58-2B-4T-bf16 for downstream task. The saved checkpoint directory is as follows:

path/to/my/ckpt  
    ├── chat_template.jinja  
    ├── config.json  
    ├── generation_config.json  
    ├── model.safetensors  
    ├── optimizer.pt  
    ├── rng_state.pth  
    ├── scheduler.pt  
    ├── special_tokens_map.jso  
    ├── tokenizer.json  
    ├── tokenizer_config.json  
    ├── trainer_state.json  
    └── training_args.bin  

Now I'm trying to convert this model to the gguf format, as what it is in readme file:

python setup_env.py -md models/BitNet-b1.58-2B-4T -q i2_s

(replaced models/BitNet-b1.58-2B-4T with my actual model path)

And I modified the setup_env.py file to support my local model path, but I'm encountering an error when running the command, as follows:

INFO:hf-to-gguf:Loading model: checkpoint-1800
INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
INFO:hf-to-gguf:Set model parameters
INFO:hf-to-gguf:gguf: context length = 4096
INFO:hf-to-gguf:gguf: embedding length = 2560
INFO:hf-to-gguf:gguf: feed forward length = 6912
INFO:hf-to-gguf:gguf: head count = 20
INFO:hf-to-gguf:gguf: key-value head count = 5
INFO:hf-to-gguf:gguf: rope theta = 500000.0
INFO:hf-to-gguf:gguf: rms norm epsilon = 1e-05
INFO:hf-to-gguf:gguf: file type = 0
INFO:hf-to-gguf:Set model tokenizer
Traceback (most recent call last):
  File "/data/personal/liuzhi/projects/LLaMA-Factory/BitNet/utils/convert-hf-to-gguf-bitnet.py", line 1165, in <module>
    main()
  File "/data/personal/liuzhi/projects/LLaMA-Factory/BitNet/utils/convert-hf-to-gguf-bitnet.py", line 1150, in main
    model_instance.set_vocab()
  File "/data/personal/liuzhi/projects/LLaMA-Factory/BitNet/utils/convert-hf-to-gguf-bitnet.py", line 957, in set_vocab
    self._set_vocab_sentencepiece()
  File "/data/personal/liuzhi/projects/LLaMA-Factory/BitNet/utils/convert-hf-to-gguf-bitnet.py", line 383, in _set_vocab_sentencepiece
    raise FileNotFoundError(f"File not found: {tokenizer_path}")
FileNotFoundError: File not found: /data/personal/liuzhi/projects/LLaMA-Factory/ckpt/bitnet-b1.58-2B-4T-bf16/20250428_ele_v2/checkpoint-1800/tokenizer.model

It seems that the tokenizer.model file is missing but this file doesn't exist in the offical code and model files.

Could you please provide some guidance on how to convert my model to the gguf format and inference, like the official code did? Appreciate!

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Read setup_env.py and utils/convert-hf-to-gguf-bitnet.py, starting at main(), set_vocab(), and _set_vocab_sentencepiece(). Compare the checkpoint contents with the tokenizer files expected by the conversion entry point, then verify the converted GGUF can be used for inference as in the official workflow.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, tooling
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.