🐛 [Bug] BERT failure with cuda 12.4 (Schema not found for node add)
Open
Nobody has claimed this yet.
bug
Story: Build & Install & Packaging
- Dominant language
- Python
- Stars
- 3k
- Forks
- 410
- Avg merge
- 3d 18h
- Merged PRs (30d)
- 78
Description
Bug Description
Using nightly, the CI produces this issue with cuda 12.4
=========================== short test summary info ============================
FAILED models/test_models.py::TestModels::test_bert_base_uncased - RuntimeError:
Schema not found for node. File a bug report.
Node: %6932 : int = aten::add(%5840, %6931, %26)
To Reproduce
Steps to reproduce the behavior:
Expected behavior
Environment
Build information about Torch-TensorRT can be found by turning on debug messages
- Torch-TensorRT Version (e.g. 1.0.0):
- PyTorch Version (e.g. 1.0):
- CPU Architecture:
- OS (e.g., Linux):
- How you installed PyTorch (
conda,pip,libtorch, source): - Build command you used (if compiling from source):
- Are you using local sources or building from archives:
- Python version:
- CUDA version:
- GPU models and configuration:
- Any other relevant information:
Additional context
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with models/test_models.py::TestModels::test_bert_base_uncased and reproduce the nightly CI failure, recording the environment details listed in the issue. Trace the reported aten::add schema failure and verify that the test passes with CUDA 12.4 after the compatibility issue is addressed.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- compilers, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100