Fine-tuning bert-base model using mixed precision does not reduce model size
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 25/100
- Issue type
- Bug
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- python
- Domain
- machine-learning
Research direction
No repository file, test, or entry point is named in the report. Start by reproducing fine-tuning of bert-base with Apex O1, O2, and without Apex using the stated PyTorch 1.4 and CUDA 10.2 versions. Done means explaining whether identical saved model sizes are expected and identifying any confirmed Apex behavior or missing warning.
Written by the indexing model from the issue text.
Description
Hello all,
I have fine-tuned bert model with different apex levels i.e. O1, O2, and without apex. At the end, I am having same model size for all the fine-tuned models. Could anyone explain why the model sizes are same for all the apex levels and without apex too?
PS: I do not get any warning related to apex while fine-tuning
Version:
pytorch : 1.4
torch.version.cuda : 10.2
Nvidia cuda version: 10.2
- Dominant language
- Python
- Stars
- 9k
- Forks
- 1.5k
- Avg merge
- 2d 4h
- Merged PRs (30d)
- 3
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/apex
-
Difficulty 4/5 3-5 days Newbie friendliness 64/100
-
Difficulty 3/5 1-2 days Newbie friendliness 45/100
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 48/100
-
Difficulty 3/5 1-2 days Newbie friendliness 55/100
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 35/100
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
bancolombia/sentinel#23 ·
-
test md OpenCI
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
-
integration:quickjs org:external priority:backlog topic:code-interpreter topic:middleware type:feature
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
langchain-ai/deepagents#6450 ·
-
bug client
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100