Euler2.0分布式训练大规模Embedding,从checkpoint恢复模型失败
- Dominant language
- C++
- Stars
- 2.9k
- Forks
- 553
- PR merge metrics
- No merged PRs in 30d
Description
如下图


应该是tensorflow 1.12.0的Bug,Euler集成的tensorflow可以升级吗
https://github.com/tensorflow/tensorflow/pull/25368
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reproducing Euler 2.0 distributed large-Embedding training recovery from a checkpoint with the integrated TensorFlow 1.12.0, then read TensorFlow pull request #25368 for the reported bug and proposed compatibility change. Done means checkpoint restoration works with the supported TensorFlow integration, or the compatibility limitation and upgrade requirements are documented.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- tensorflow
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100