NVIDIA / NVIDIA/Model-Optimizer

calling mtq.quantize causes AttributeError: 'tuple' object has no attribute 'dim'

Open
#2,424 1 comment 0 reactions 1 assignee View on GitHub

@jingyu-ml is already working on this.

Since Sep 16, 2026.

question
Dominant language
Python
Stars
3.8k
Forks
604
Avg merge
2d 8h
Merged PRs (30d)
142

Description

Make sure you already checked the examples and documentation before submitting an issue.

How would you like to use ModelOpt

running python Model-Optimizer/examples/diffusers/quantization/quantize.py --model sdxl-turbo --format fp8 --n-steps 4 --calib-size 128

results in

File "c:\python.venv\lib\site-packages\torch\nn\functional.py", line 3040, in group_norm
if input.dim() < 2:
AttributeError: 'tuple' object has no attribute 'dim'

the full output:

c:\python.venv\lib\site-packages\torch\jit_script.py:1491: FutureWarning: torch.jit.script is deprecated. Please switch to torch.compile or torch.export.
warnings.warn(
c:\python.venv\lib\site-packages\modelopt\torch_init_.py:55: UserWarning: transformers 4.56.0 is not tested with current version of modelopt and may cause issues. Please install recommended version with pip install -U nvidia-modelopt[hf] if working with HF models.
_warnings.warn(
2026-09-13 18:24:43 | INFO | main | Starting Enhanced Diffusion Model Quantization
2026-09-13 18:24:43 | INFO | main | Validating configurations...
2026-09-13 18:24:43 | INFO | main | Creating pipeline for sdxl-turbo
2026-09-13 18:24:43 | INFO | main | Model path: stabilityai/sdxl-turbo
2026-09-13 18:24:43 | INFO | main | Data type: {'default': torch.float16}
Loading pipeline components...: 57%|█████████████████████████████▋ | 4/7 [00:02<00:01, 1.70it/s]torch_dtype is deprecated! Use dtype instead!
Loading pipeline components...: 100%|████████████████████████████████████████████████████| 7/7 [00:05<00:00, 1.35it/s]
2026-09-13 18:24:56 | INFO | main | Pipeline created successfully
2026-09-13 18:24:56 | INFO | main | Moving pipeline to CUDA
2026-09-13 18:24:57 | INFO | main | Initializing calibration...
2026-09-13 18:24:57 | INFO | main | Loading calibration prompts from {'name': 'Gustavosta/Stable-Diffusion-Prompts', 'split': 'train', 'column': 'Prompt'}
2026-09-13 18:25:00 | INFO | main | Quantizing backbone: unet
2026-09-13 18:25:00 | INFO | main | Building quantization config for fp8
2026-09-13 18:25:00 | INFO | main | Quant config {'quant_cfg': [{'quantizer_name': '*', 'enable': False}, {'quantizer_name': '*weight_quantizer', 'cfg': {'num_bits': (4, 3), 'axis': None, 'trt_high_precision_dtype': 'Half'}}, {'quantizer_name': '*input_quantizer', 'cfg': {'num_bits': (4, 3), 'axis': None, 'trt_high_precision_dtype': 'Half'}}, {'quantizer_name': '*output_quantizer', 'enable': False}, {'quantizer_name': '*softmax_quantizer', 'cfg': {'num_bits': (4, 3), 'axis': None, 'trt_high_precision_dtype': 'Half'}}], 'algorithm': {'method': 'max'}}
2026-09-13 18:25:00 | INFO | main | Checking for LoRA layers...
2026-09-13 18:25:00 | INFO | main | Starting model quantization for unet...
Inserted 3502 quantizers
2026-09-13 18:25:01 | INFO | main | Starting calibration with 64 batches
Calibration: 0%| | 0/64 [00:00<?, ?batch/s]Token indices sequence length is longer than the specified maximum sequence length for this model (101 > 77). Running this sequence through the model will result in indexing errors
The following part of your input was truncated because CLIP can only handle sequences up to 77 tokens: ['incev, in style of lee souder, in plastic, dark atmosphere, tilt shift, depth of field,', 'trending on art station <|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|><|endoftext|>']
Token indices sequence length is longer than the specified maximum sequence length for this model (101 > 77). Running this sequence through the model will result in indexing errors
The following part of your input was truncated because CLIP can only handle sequences up to 77 tokens: ['incev, in style of lee souder, in plastic, dark atmosphere, tilt shift, depth of field,', 'trending on art station <|endoftext|>!!!!!!!!!!!!!!!!!!!']
Calibration: 0%| | 0/64 [00:03<?, ?batch/s]
2026-09-13 18:25:05 | ERROR | main | Quantization failed: 'tuple' object has no attribute 'dim'
Traceback (most recent call last):
File "C:\python\Model-Optimizer\examples\diffusers\quantization\quantize.py", line 706, in main
quantizer.quantize_model(
File "C:\python\Model-Optimizer\examples\diffusers\quantization\quantize.py", line 239, in quantize_model
mtq.quantize(backbone, quant_config, forward_loop)
File "c:\python.venv\lib\site-packages\modelopt\torch\quantization\model_quant.py", line 247, in quantize
return calibrate(model, config.get("algorithm"), forward_loop=forward_loop)
File "c:\python.venv\lib\site-packages\modelopt\torch\quantization\model_quant.py", line 110, in calibrate
apply_mode(
File "c:\python.venv\lib\site-packages\modelopt\torch\opt\conversion.py", line 419, in apply_mode
model, metadata = get_mode(m).convert(model, config, **kwargs) # type: ignore [call-arg]
File "c:\python.venv\lib\site-packages\modelopt\torch\quantization\mode.py", line 351, in wrapped_func
return wrapped_calib_func(
File "c:\python.venv\lib\site-packages\modelopt\torch\quantization\mode.py", line 273, in wrapped_calib_func
func(model, forward_loop=forward_loop, **kwargs)
File "c:\python.venv\lib\site-packages\torch\utils_contextlib.py", line 124, in decorate_context
return func(*args, **kwargs)
File "c:\python.venv\lib\site-packages\modelopt\torch\quantization\model_calib.py", line 354, in max_calibrate
forward_loop(model)
File "C:\python\Model-Optimizer\examples\diffusers\quantization\quantize.py", line 704, in forward_loop
calibrator.run_calibration(batched_prompts)
File "C:\python\Model-Optimizer\examples\diffusers\quantization\calibration.py", line 103, in run_calibration
self.pipe(**common_args, **extra_args).images
File "c:\python.venv\lib\site-packages\torch\utils_contextlib.py", line 124, in decorate_context
return func(*args, **kwargs)
File "c:\python.venv\lib\site-packages\diffusers\pipelines\stable_diffusion_xl\pipeline_stable_diffusion_xl.py", line 1292, in call
image = self.vae.decode(latents, return_dict=False)[0]
File "c:\python.venv\lib\site-packages\diffusers\utils\accelerate_utils.py", line 46, in wrapper
return method(self, *args, **kwargs)
File "c:\python.venv\lib\site-packages\diffusers\models\autoencoders\autoencoder_kl.py", line 294, in decode
decoded = self._decode(z).sample
File "c:\python.venv\lib\site-packages\diffusers\models\autoencoders\autoencoder_kl.py", line 265, in _decode
dec = self.decoder(z)
File "c:\python.venv\lib\site-packages\torch\nn\modules\module.py", line 1783, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
File "c:\python.venv\lib\site-packages\torch\nn\modules\module.py", line 1794, in _call_impl
return forward_call(*args, **kwargs)
File "c:\python.venv\lib\site-packages\diffusers\models\autoencoders\vae.py", line 298, in forward
sample = self.mid_block(sample, latent_embeds)
File "c:\python.venv\lib\site-packages\torch\nn\modules\module.py", line 1783, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
File "c:\python.venv\lib\site-packages\torch\nn\modules\module.py", line 1794, in _call_impl
return forward_call(*args, **kwargs)
File "c:\python.venv\lib\site-packages\diffusers\models\unets\unet_2d_blocks.py", line 746, in forward
hidden_states = resnet(hidden_states, temb)
File "c:\python.venv\lib\site-packages\torch\nn\modules\module.py", line 1783, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
File "c:\python.venv\lib\site-packages\torch\nn\modules\module.py", line 1794, in _call_impl
return forward_call(*args, **kwargs)
File "c:\python.venv\lib\site-packages\diffusers\models\resnet.py", line 327, in forward
hidden_states = self.norm1(hidden_states)
File "c:\python.venv\lib\site-packages\torch\nn\modules\module.py", line 1783, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
File "c:\python.venv\lib\site-packages\torch\nn\modules\module.py", line 1794, in _call_impl
return forward_call(*args, **kwargs)
File "c:\python.venv\lib\site-packages\torch\nn\modules\normalization.py", line 334, in forward
return F.group_norm(input, self.num_groups, self.weight, self.bias, self.eps)
File "c:\python.venv\lib\site-packages\torch\nn\functional.py", line 3040, in group_norm
if input.dim() < 2:
AttributeError: 'tuple' object has no attribute 'dim'

Who can help?
  • ?

System information

  • Container used (if applicable): ?
  • OS (e.g., Ubuntu 22.04, CentOS 7, Windows 10): ? Windows 11
  • CPU architecture (x86_64, aarch64): Core i7 13700K
  • GPU name (e.g. H100, A100, L40S): RTX5080
  • GPU memory size: 32gb
  • Number of GPUs: 1
  • Library versions (if applicable):
    • Python: 3.10.14
    • ModelOpt version or commit hash: 0.46.1
    • CUDA: 13.2
    • PyTorch: 2.14.0+cu132
    • Transformers: 4.56.0
    • TensorRT-LLM: ?
    • ONNXRuntime: 1.23.2
    • TensorRT: 10.16.1.11
  • Any other details that may help: ?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.