Where is the global attention?
未关闭
- 主要语言
- Python
- 星标
- 2.2k
- 派生
- 286
- PR 合并指标
- 30 天内没有已合并 PR
描述
I searched some code in longformer and the related code in transformers,the invert_mask() function in BartEncoder destroies the integer 2 in the attention mask,but the longformer attention code regard the mask as it dosen't been inverted.
So I think the global attention is not enabled in the model, could you explain it to me ?
Hope that I'm wrong....
贡献指南
这个仓库没有索引到贡献指南
评估
这个 Issue 还没有评估数据。