google-research / google-research/t5x
Why scale the expert gradient during training, due to mixed precision training?
Open
- Dominant language
- Python
- Stars
- 3k
- Forks
- 338
- PR merge metrics
- No merged PRs in 30d
Description
This issue has no description.
Contributor guide
Research direction
The issue has no body, so begin by reading its comment thread and searching the repository for “expert gradient” and its scaling logic. Trace the relevant training entry point and document why the scaling is used, including whether mixed-precision training is the reason; done means the rationale is clear and supported by the implementation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100