deepspeedai / deepspeedai/DeepSpeed
lora error in stage 1
Open
@samadejacobs is already working on this.
Since Apr 21, 2023.
bug
deepspeed-chat
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
use lora param in stage1
here is context:
deepspeed main.py \
--lora_dim 8 --only_optimize_lora \
--data_path wangrui6/Zhihu-KOL Cohere/miracl-zh-queries-22-12 Hello-SimpleAI/HC3-Chinese mkqa-Chinese \
--data_split 10,0,0 \
--model_name_or_path $MODEL \
--per_device_train_batch_size 2 \
--per_device_eval_batch_size 2 \
--learning_rate 9.65e-6 \
--num_train_epochs 16 \
--deepspeed --seed 1234 --num_warmup_steps 0 \
--lr_scheduler_type cosine \
--output_dir $OUTPUT_PATH \
&> $OUTPUT_PATH/training.log
add about lora
--lora_dim 8 --only_optimize_lora \
get error:
RuntimeError: element 0 of tensors does not require grad and does not have a grad_fn
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.