Lightning-AI / Lightning-AI/lightning-thunder

FSDP2 & Thunder looks memory hungrier than `thunder.distributed.fsdp` for certain models

Open
#1,176 3 comments 1 reaction 1 assignee View on GitHub

@IvanYashchuk is already working on this.

Since Sep 30, 2024.

distributed memory use thunderfx
Dominant language
Python
Stars
1.5k
Forks
121
PR merge metrics
No merged PRs in 30d

Description

Note: If you have a model or program that is not supported yet but should be, please use the program coverage template.

🐛 Bug

Let's take stablecode-completion-alpha-3b whose sequence length (Config.block_size) is 16384,

torchrun --standalone --role rank --tee 3 --local-ranks-filter 0 --nproc-per-node 8 thunder/benchmarks/benchmark_litgpt.py --model_name stablecode-completion-alpha-3b --distributed_mode fsdp2 --shard_mode zero2 --compile thunder_inductor_cat_cudnn_dynamo

This command goes OOM while the same config (= model and compile) works with a single H100 with the memory usage of 77.02 GB.

For the sequence length of 16384

FSDP Impl Thunder Torch Compile Diff
thunder fsdp 67.74 #N/A #N/A
FSDP2 OOM 62.7 #N/A
FSDP1 #N/A 62.97 #N/A

For the sequence length of 8192

FSDP Impl Thunder Torch Compile Diff
thunder fsdp 37.58 #N/A #N/A
FSDP2 56.69 35.21 21.48
FSDP1 #N/A 35.24 #N/A

When --distributed_mode is "fsdp", then the benchmark script chooses thunder.distributed.fsdp for --compile of thunder w/o dynamo keyword, and FSDP1 for the others.

Clearly, FSDP2 & Thunder uses too much memory even compared to thunder's fsdp, while thunder's fsdp itslef seems to use more memory than Eager and Torch Compile.
When I was on #940, I didn't see this trend of memory usage. Also, for Llama-3-8B,
thunder still uses more memory but the gap is not that huge.

FSDP Impl Thunder Torch Compile Diff
thunder fsdp 75.79 #N/A #N/A
FSDP2 74.84 72.61 2.23
FSDP1 #N/A 73.41 #N/A
To Reproduce

Apply a diff like this and run commands like

torchrun --standalone --role rank --tee 3 --local-ranks-filter 0 thunder/benchmarks/benchmark_litgpt.py --model_name stablecode-completion-alpha-3b --warmup_iters 0 --max_iters 3 --compile eager --dump_memory_snapshot false --block_size 2048
@@ -227,6 +269,7 @@ class Benchmark_litGPT:
         fsdp_bucket_params: float | None = None,
         checkpoint_activations: bool = False,
         n_layers: int | None = None,
+        block_size: int | None = None,
         profiler_start: int = 15,
         profiler_stop: int = 15,
         skip_data_sync: bool = False,
@@ -360,6 +403,8 @@ class Benchmark_litGPT:

         if n_layers is not None:
             self.config.n_layer = n_layers
+        if block_size is not None:
+            self.config.block_size = block_size

         # Initialize the model
         t0 = time.perf_counter()
Code sample
Expected behavior
Environment

pjnl-20240919

Additional context

related to #1175

cc @carmocca @crcrpar

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.