Not able to observe any speedup on a Nvidia T4 (Turing arch)
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 9k
- Forks
- 1.5k
- Avg merge
- 2d 4h
- Merged PRs (30d)
- 3
Description
I trained a model with the fast mixed precision using amp, pytorch on a RTX 2070. I am now trying to run inference on the network on a NVIDIA T4 and I am not able to observe any speedup.
On checking the state_dict of the checkpoint it looks like they are all Torch.Tensors and not Torch.HalfTensor. I can see that the size of the model is smaller in Mixed Precision version than FP32. In order to run inference I just loaded the amp state dict and did an amp initialize. Is there anything else I had to do in order for it to run in FP16?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the reported inference setup on an NVIDIA T4: load the AMP checkpoint, call AMP initialization, and compare it with FP32 inference. Inspect how the state_dict and AMP initialization affect inference precision; done means determining whether additional setup is required and explaining the absence of speedup.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- machine-learning, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100