Lightning-AI / Lightning-AI/litgpt

LoRA with quantization: `micro_batch_size` effect on memory footprint

Open
#501 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug help wanted quantization
Dominant language
Python
Stars
13.7k
Forks
1.5k
Avg merge
15h 37m
Merged PRs (30d)
1

Description

> [!Important]
> These are just quick tests with a single model on a single graphics card, so take it with a grain of salt.
Nevertheless, this issue is worth discussing in my opinion.

Hi there 👋

I'm still not confident that the quantization works properly, so I decided to do quick tests with a small model just to see how much we can gain in memory footprint with and without quantization.
During experiments, I noticed a somewhat weird behavior: with a smaller `micro_batch_size` (1 or 2) the gain is bigger than with the default `micro_batch_size` of 4. In my head, I can explain it with the size of activations that outweighs quantization effect (though still doubtful). What I cannot explain is that with the `micro_batch_size` larger than the default value of 4 the memory footprint might be even larger than without using quantization.

All ran with `Pythia-70m`, precision `16-mixed`, quantization `bnb.nf4` and default parameters apart from `micro_batch_size`.

| Micro BatchSize | Cuda allocated | Cuda_allocated $_{quantized}$ | nvidia-smi | nvidia-smi $_{quantized}$ | Card |
|-----------------|----------------|-------------------------------|------------|---------------------------|------|
| 1 | 1.37 | 0.80 | 1.63 | 0.96 | T4 |
| 2 | 1.92 | 1.42 | 2.04 | 1.73 | T4 |
| 4 | 3.04 | 2.69 | 3.93 | 3.26 | T4 |
| 8 | 5.29 | 5.23 | 7.93 | 7.78 | T4 |
| 16 | 9.79 | 10.32 | 11.304 | 10.407 | T4 |
| 32 | 18.77 | 20.49 | 19.217 | 20.62 | A10G |

----

Don't know who will be assigned to this task, so here is a list of steps that I would do:
- [ ] Sanity check with different models/precisions/graphic cards
- [ ] Comparison with Higgingface implementation of QLoRA
- [ ] Memory profiling with PyTorch Profiler

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the reported Pythia-70m measurements with 16-mixed precision, bnb.nf4 quantization, and the listed micro_batch_size values. Compare different models, precisions, and graphics cards, then use the Hugging Face QLoRA implementation and PyTorch Profiler to determine whether the memory behavior is expected or indicates a quantization issue.

Written by the indexing model from the issue text.

Assessment

Tech stack
huggingface, python, pytorch
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.