AnswerDotAI / AnswerDotAI/fsdp_qlora
Training using FSDP, qLoRa on multinode
Open
- Dominant language
- Jupyter Notebook
- Stars
- 1.6k
- Forks
- 201
- PR merge metrics
- No merged PRs in 30d
Description
How to train a TinyLlama-like model using FSDP, qLoRa on two (or more) nodes if each node has one (or more) GPU using train.py (main branch) ?
I am grateful in advance for any help!
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading train.py on the main branch and identifying how the existing FSDP and qLoRA training flow is launched. Determine what is needed for two or more nodes with one or more GPUs per node. Done would be a reproducible multinode training procedure or documented support for the requested setup.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- distributed-systems, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100