alibaba / alibaba/euler

Euler2.0分布式训练大规模Embedding,从checkpoint恢复模型失败

Open
#315 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
2.9k
Forks
553
PR merge metrics
No merged PRs in 30d

Description

如下图
![image](https://user-images.githubusercontent.com/8065137/96850178-1a584680-1489-11eb-8cbf-a07248d21e23.png)

![image](https://user-images.githubusercontent.com/8065137/96850187-1d533700-1489-11eb-971d-5b26c27d4f28.png)

应该是tensorflow 1.12.0的Bug,Euler集成的tensorflow可以升级吗
https://github.com/tensorflow/tensorflow/pull/25368

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing Euler 2.0 distributed large-Embedding training recovery from a checkpoint with the integrated TensorFlow 1.12.0, then read TensorFlow pull request #25368 for the reported bug and proposed compatibility change. Done means checkpoint restoration works with the supported TensorFlow integration, or the compatibility limitation and upgrade requirements are documented.

Written by the indexing model from the issue text.

Assessment

Tech stack
tensorflow
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.