backward gradient is vanished
Open
- Dominant language
- Lua
- Stars
- 872
- Forks
- 235
- PR merge metrics
- No merged PRs in 30d
Description
Hi,thanks for your excellent work, and I'm focused on the work for a period. I think the core is Gradient optimization. But still then I haven't reproduce your experiment. Could you provide a little advice to me?
I build the network with the block that mentioned in your paper (B->A->C->P). and Backward is using full precision data (weights&gradient) for ||r|| < 1.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.