deepspeedai / deepspeedai/DeepSpeed

[BUG] DeepCompile: MemoryProfiling error /pytorch/build/aten/src/ATen/RegisterCUDA.cpp:7280: SymIntArrayRef expected to contain only concrete integers

Open
#7,311 7 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug training
Dominant language
Python
Stars
43.1k
Forks
5k
Avg merge
4d 15h
Merged PRs (30d)
112

Description

Describe the bug
@tohtana Thanks for your great work about DeepCompile.
I am trying to enable DeepCompile for a 2-nodes training job, while hit below exception:

MemoryProfiling error /pytorch/build/aten/src/ATen/RegisterCUDA.cpp:7280: SymIntArrayRef expected to contain only concrete integers

Have you met similar issue before? Appreciate for any clue!

Launching compile passes: global_steps=0 passes=[<function add_z1_reduce at 0x7fe9857581f0>]
MemoryProfiling error /pytorch/build/aten/src/ATen/RegisterCUDA.cpp:7280: SymIntArrayRef expected to contain only concrete integers

While executing %iota : [num_users=3] = call_function[target=torch.ops.prims.iota.default](args = (%primals_2,), kwargs = {start: 0, step: 1, dtype: torch.int64, device: cuda:0, requires_grad: False})
Original traceback:
  File "/scratch/azureml/cr/j/89c766bec23f4cbb85212c424847a0fd/exe/wd/bgm/modeling/load_train.py", line 361, in causal_lm_forward
    outputs = model.__original_forward__(input_ids, attention_mask=attention_mask)
  File "/opt/conda/envs/ptca/lib/python3.10/site-packages/transformers/models/qwen2/modeling_qwen2.py", line 1165, in forward
    outputs = self.model(
  File "/opt/conda/envs/ptca/lib/python3.10/site-packages/transformers/models/qwen2/modeling_qwen2.py", line 858, in forward
    cache_position = torch.arange(

[rank0]: Traceback (most recent call last):
[rank0]:   File "/scratch/azureml/cr/j/89c766bec23f4cbb85212c424847a0fd/exe/wd/train.py", line 350, in <module>
[rank0]:     run_app()
[rank0]:   File "/scratch/azureml/cr/j/89c766bec23f4cbb85212c424847a0fd/exe/wd/train.py", line 334, in run_app
[rank0]:     train(
[rank0]:   File "/scratch/azureml/cr/j/89c766bec23f4cbb85212c424847a0fd/exe/wd/train.py", line 164, in train
[rank0]:     outputs, _ = _forward(model, batch, conf)
[rank0]:   File "/scratch/azureml/cr/j/89c766bec23f4cbb85212c424847a0fd/exe/wd/train.py", line 112, in _forward
[rank0]:     outputs = model(**inputs, use_cache=False)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1739, in _wrapped_call_impl
[rank0]:     return self._call_impl(*args, **kwargs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1750, in _call_impl
[rank0]:     return forward_call(*args, **kwargs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/deepspeed/utils/nvtx.py", line 20, in wrapped_fn
[rank0]:     ret_val = func(*args, **kwargs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/deepspeed/runtime/engine.py", line 2054, in forward
[rank0]:     loss = self.module(*inputs, **kwargs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1737, in _wrapped_call_impl
[rank0]:     return self._compiled_call_impl(*args, **kwargs)  # type: ignore[misc]
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/eval_frame.py", line 574, in _fn
[rank0]:     return fn(*args, **kwargs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1750, in _call_impl
[rank0]:     return forward_call(*args, **kwargs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/convert_frame.py", line 1380, in __call__
[rank0]:     return self._torchdynamo_orig_callable(
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/convert_frame.py", line 1164, in __call__
[rank0]:     result = self._inner_convert(
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/convert_frame.py", line 547, in __call__
[rank0]:     return _compile(
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/convert_frame.py", line 986, in _compile
[rank0]:     guarded_code = compile_inner(code, one_graph, hooks, transform)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/convert_frame.py", line 715, in compile_inner
[rank0]:     return _compile_inner(code, one_graph, hooks, transform)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_utils_internal.py", line 95, in wrapper_function
[rank0]:     return function(*args, **kwargs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/convert_frame.py", line 750, in _compile_inner
[rank0]:     out_code = transform_code_object(code, transform)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/bytecode_transformation.py", line 1361, in transform_code_object
[rank0]:     transformations(instructions, code_options)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/convert_frame.py", line 231, in _fn
[rank0]:     return fn(*args, **kwargs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/convert_frame.py", line 662, in transform
[rank0]:     tracer.run()
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/symbolic_convert.py", line 2868, in run
[rank0]:     super().run()
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/symbolic_convert.py", line 1052, in run
[rank0]:     while self.step():
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/symbolic_convert.py", line 962, in step
[rank0]:     self.dispatch_table[inst.opcode](self, inst)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/symbolic_convert.py", line 657, in wrapper
[rank0]:     return handle_graph_break(self, inst, speculation.reason)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/symbolic_convert.py", line 698, in handle_graph_break
[rank0]:     self.output.compile_subgraph(self, reason=reason)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/output_graph.py", line 1136, in compile_subgraph
[rank0]:     self.compile_and_call_fx_graph(
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/output_graph.py", line 1382, in compile_and_call_fx_graph
[rank0]:     compiled_fn = self.call_user_compiler(gm)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/output_graph.py", line 1432, in call_user_compiler
[rank0]:     return self._call_user_compiler(gm)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/output_graph.py", line 1483, in _call_user_compiler
[rank0]:     raise BackendCompilerFailed(self.compiler_fn, e).with_traceback(
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/output_graph.py", line 1462, in _call_user_compiler
[rank0]:     compiled_fn = compiler_fn(gm, self.example_inputs())
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/repro/after_dynamo.py", line 130, in __call__
[rank0]:     compiled_gm = compiler_fn(gm, example_inputs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/__init__.py", line 2385, in __call__
[rank0]:     return self.compiler_fn(model_, inputs_, **self.kwargs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/deepspeed/compile/backend.py", line 275, in backend_fn
[rank0]:     return torch._inductor.compile(gm, real_inputs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_inductor/__init__.py", line 48, in compile
[rank0]:     return compile_fx(gm, example_inputs, config_patches=options)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_inductor/compile_fx.py", line 1863, in compile_fx
[rank0]:     return aot_autograd(
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_dynamo/backends/common.py", line 83, in __call__
[rank0]:     cg = aot_module_simplified(gm, example_inputs, **self.kwargs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_functorch/aot_autograd.py", line 1155, in aot_module_simplified
[rank0]:     compiled_fn = dispatch_and_compile()
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_functorch/aot_autograd.py", line 1131, in dispatch_and_compile
[rank0]:     compiled_fn, _ = create_aot_dispatcher_function(
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_functorch/aot_autograd.py", line 580, in create_aot_dispatcher_function
[rank0]:     return _create_aot_dispatcher_function(
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_functorch/aot_autograd.py", line 830, in _create_aot_dispatcher_function
[rank0]:     compiled_fn, fw_metadata = compiler_fn(
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/torch/_functorch/_aot_autograd/jit_compile_runtime_wrappers.py", line 678, in aot_dispatch_autograd
[rank0]:     compiled_fw_func = aot_config.fw_compiler(fw_module, adjusted_flat_args)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/deepspeed/compile/inductor.py", line 27, in wrapped_compiler
[rank0]:     mod_graph = dc_compiler(gm, fake_inputs)
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/deepspeed/compile/backend.py", line 178, in make_fw_graph
[rank0]:     run_opt_passes(
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/deepspeed/compile/backend.py", line 129, in run_opt_passes
[rank0]:     mem_prof.run(*create_inputs_fn())
[rank0]:   File "/opt/conda/envs/ptca/lib/python3.10/site-packages/deepspeed/compile/profilers/graph_profile.py", line 258, in run
[rank0]:     return return_val
[rank0]: torch._dynamo.exc.BackendCompilerFailed: backend='backend_fn' raised:
[rank0]: UnboundLocalError: local variable 'return_val' referenced before assignment

To Reproduce
Steps to reproduce the behavior:

  1. Go to '...'
  2. Click on '....'
  3. Scroll down to '....'
  4. See error

Expected behavior
A clear and concise description of what you expected to happen.

ds_report output
Please run ds_report to give us details about your setup.

Screenshots
If applicable, add screenshots to help explain your problem.

System info (please complete the following information):

  • OS: [e.g. Ubuntu 18.04]
  • GPU count and types: two machines with x8 A100s each]
  • Interconnects (if applicable) [e.g., two machines connected with 100 Gbps IB]
  • Python version
  • Any other relevant info about your setup

Launcher context
Are you launching your experiment with the deepspeed launcher, MPI, or something else?

Docker context
Are you using a specific docker image that you can share?

Additional context
Add any other context about the problem here.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with deepspeed/compile/profilers/graph_profile.py:258 and trace the MemoryProfiling path from deepspeed/compile/backend.py:129 and :178. Reproduce with the reported Qwen2, DeepCompile, two-node A100 setup, then verify that the SymIntArrayRef error and follow-on UnboundLocalError no longer occur during compilation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
distributed-systems, machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.