deepspeedai / deepspeedai/DeepSpeed
[BUG] is_zero_init_model is always False when I'm using zero_init!
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
Describe the bug
When I'm fine tuning llama2 with deepspeed zero3, I set "zero3_init_flag: true" in my accelerate config. The "is_deepspeed_zero3_enabled()" in transformers/integrations/deepspeed.py is also judged to True. But the "is_zero_init_model" is judged to False in _configure_distributed_model of deepspeed/runtime/engine.py. I'm not sure if it abnormal?
To Reproduce
Here is my code:
from datasets import load_dataset
import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig, AutoTokenizer, TrainingArguments
import bitsandbytes as bnb
from peft import LoraConfig
from trl import SFTTrainer
base_model_name ="/home/yangtong/data/llama2-hf/llama2-13b-chat_hf"
dataset = load_dataset("json",data_files="Belle_open_source_0.5M_changed.json",split="train")
result_dir = "tmp"
training_args = TrainingArguments(
report_to="none",
output_dir=result_dir,
# per_device_train_batch_size * gradient_accumulation_steps = batch_size
per_device_train_batch_size=1,
gradient_accumulation_steps=16,
learning_rate=2e-4,
logging_steps=10,
# max_steps=520,
num_train_epochs=0.016,
save_steps=500,
bf16 = True, # set bf16 to True with an A100
# optim='paged_adamw_32bit',
gradient_checkpointing=True
)
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
)
base_model = LlamaForCausalLM.from_pretrained(
base_model_name,
quantization_config=bnb_config,
torch_dtype=torch.bfloat16,
trust_remote_code=True
)
base_model.config.use_cache = False
base_model.config.pretraining_tp = 1
def find_all_linear_names(model):
cls = bnb.nn.Linear4bit
lora_module_names = set()
for name, module in model.named_modules():
if isinstance(module, cls):
names = name.split('.')
lora_module_names.add(names[0] if len(names) == 1 else names[-1])
if 'lm_head' in lora_module_names: # needed for 16-bit
lora_module_names.remove('lm_head')
return list(lora_module_names)
models=find_all_linear_names(base_model)
peft_config = LoraConfig(
lora_alpha=16,
lora_dropout=0.1,
r=64,
bias="none",
task_type="CAUSAL_LM",
target_modules=models
)
tokenizer = AutoTokenizer.from_pretrained(base_model_name, trust_remote_code=True)
tokenizer.deprecation_warnings["Asking-to-pad-a-fast-tokenizer"] = True
tokenizer.pad_token = tokenizer.eos_token
max_seq_length = 512
trainer = SFTTrainer(
model=base_model,
train_dataset=dataset,
peft_config=peft_config,
dataset_text_field="text",
max_seq_length=max_seq_length,
tokenizer=tokenizer,
args=training_args
)
trainer.train()
output_dir = os.path.join(result_dir, "final_checkpoint")
trainer.model.save_pretrained(output_dir)
Here is my accelerate config:
compute_environment: LOCAL_MACHINE
debug: false
deepspeed_config:
deepspeed_config_file: /home/yangtong/ft_dis/ds_config/3.json
zero3_init_flag: true
distributed_type: DEEPSPEED
downcast_bf16: 'no'
enable_cpu_affinity: false
machine_rank: 0
main_training_function: main
num_machines: 1
num_processes: 4
rdzv_backend: 'c10d'
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false
Here is my deepspeed config:
{
"optimizer": {
"type": "AdamW",
"params": {
"lr": 2e-4,
"betas": [
0.9,
0.999
],
"eps": "auto",
"weight_decay": "auto",
"adam_w_mode": true,
"torch_adam": true
}
},
"scheduler": {
"type": "WarmupDecayLR",
"params": {
"warmup_min_lr": "auto",
"warmup_max_lr": "auto",
"warmup_num_steps": "auto",
"total_num_steps": "auto"
}
},
"zero_optimization": {
"stage": 3,
"allgather_partitions": true,
"allgather_bucket_size": 2e8,
"reduce_scatter": true,
"reduce_bucket_size": 2e8,
"contiguous_gradients": true,
"overlap_comm": true,
"offload_optimizer": {
"device": "none",
"pin_memory": true
},
"offload_param": {
"device": "cpu",
"pin_memory": true
},
"sub_group_size": 1e9,
"stage3_prefetch_bucket_size": "auto",
"stage3_param_persistence_threshold": "auto",
"stage3_max_live_parameters": 1e9,
"stage3_max_reuse_distance": 1e9,
"stage3_gather_16bit_weights_on_model_save": true
},
"bf16": {
"enabled": true
},
"gradient_clipping": "auto",
"train_batch_size": "auto",
"train_micro_batch_size_per_gpu": 1,
"gradient_accumulation_steps": 16,
"wall_clock_breakdown": false
}
Expected behavior
Parameters first partition and then load to GPUs.
System info (please complete the following information):
- OS: Ubuntu 22.04.4 LTS (Linux 5.15.0-106-generic)
- GPU count and types 2 x Intel(R) Xeon(R) Gold 6326 CPU @ 2.90GHz
- Python version 3.10.13
- Pytorch version 2.2.2
- CUDA version 11.8.0
- bitsandbytes==0.43.0
- huggingface_hub==0.23.2
- accelerate==0.30.1
- transformers==4.41.1
- peft==0.9.0
- deepspeed==0.14.0
Launcher context
accelerate launch \
--config_file "config/z3_3.yaml" \
--num_processes 1 \
ft_acc.py
Here is engine.py:
I will truly appreciate if anyone can help me solve it ! @loadams @tjruwase @deepcharm
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading transformers/integrations/deepspeed.py and the _configure_distributed_model entry point in deepspeed/runtime/engine.py, then reproduce the behavior with the provided Accelerate and DeepSpeed configurations. Trace whether zero3_init_flag reaches is_zero_init_model and establish whether parameter partitioning matches the stated expectation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100