backward gradient is vanished
オープン
- 主要言語
- Lua
- スター
- 872
- フォーク
- 235
- PR マージ指標
- 30日以内にマージされた PR はありません
説明
Hi,thanks for your excellent work, and I'm focused on the work for a period. I think the core is Gradient optimization. But still then I haven't reproduce your experiment. Could you provide a little advice to me?
I build the network with the block that mentioned in your paper (B->A->C->P). and Backward is using full precision data (weights&gradient) for ||r|| < 1.
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
評価
この issue はまだ評価されていません。