allenai / allenai/CSR

Unable to train with DDP strategy

未关闭
#1 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
60
派生
10
PR 合并指标
30 天内没有已合并 PR

描述

I am trying to train edge representations on machine with 4 GPUs. However, training process hangs after validation sanity check. Training works well with single-gpu settings (and also with accelerator: 'dp' even though this is not the intended way of training).

Note: Debuging suggest that model get stuck when acessing registered buffer at: https://github.com/allenai/CSR/blob/main/src/lightning/modules/moco2_module.py#L153 . But I was unable to find a fix.

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。