AngelException: operation is not support emerged after 0-th epoch beginning and then failed
Aberta
- Linguagem predominante
- Java
- Estrelas
- 6.8k
- Forks
- 1.6k
- Merge médio
- 44min
- PRs com merge (30d)
- 1
Descrição
I trained a deeepFM model with roughly 100GB samples with 12 PSs and 64workers,. The task could start successfully, but it failed after maintaining a state at the 0-th epoch for a long time. Then, I got an error:
AngelException: operation is not support!
The syslog showed that worker_54 failed with this error.
Below is my submit

I have no idea, please help me. Thanks!
Guia de contribuição
Avaliação
Esta issue ainda não foi avaliada.