deepseek-ai / deepseek-ai/DeepSeek-VL2

How does training and inference efficiency compare to other MLLMs?

Open
#20 0 comments 2 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
5.4k
Forks
1.8k
PR merge metrics
No merged PRs in 30d

Description

The model size of DeepSeek-VL2 is 27.5B, and the trainable parameters of lora fine-tuning are 212M, which is about 9 times that of lora fine-tuning InternVL2-8B (23M)。

Using the same data, the training time is about 20 times that of InternVL2-8B. Is this in line with expectations?

I use the swift framework for fine-tuning, and the training script is as follows
```
nproc_per_node=2

CUDA_VISIBLE_DEVICES=0,1 \
NPROC_PER_NODE=$nproc_per_node \
swift sft \
--model /mnt/dolphinfs/hdd_pool/docker/user/rushzy/deepseek-ai/deepseek-vl2 \
--train_type lora \
--dataset /mnt/dolphinfs/hdd_pool/docker/user/rushzy/data/train.jsonl \
--num_train_epochs 3 \
--learning_rate 8e-5 \
--lora_rank 8 \
--lora_alpha 12 \
--max_length 4096 \
--save_only_model True \
--eval_steps 2000 \
--save_steps 2000 \
--save_total_limit -1 \
--output_dir /mnt/dolphinfs/hdd_pool/docker/user/rushzy/output/test \
--deepspeed ./ds_configs/ds_zero3_cosine.json \
--lazy_tokenize True \
--per_device_train_batch_size 2 \
--gradient_accumulation_steps 4
```

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.