AngelException: operation is not support emerged after 0-th epoch beginning and then failed
未关闭
- 主要语言
- Java
- 星标
- 6.8k
- 派生
- 1.6k
- 平均合并
- 44 分钟
- 30 天内合并 PR
- 1
描述
I trained a deeepFM model with roughly 100GB samples with 12 PSs and 64workers,. The task could start successfully, but it failed after maintaining a state at the 0-th epoch for a long time. Then, I got an error:
AngelException: operation is not support!
The syslog showed that worker_54 failed with this error.
Below is my submit

I have no idea, please help me. Thanks!
贡献指南
评估
这个 Issue 还没有评估数据。