Number of tokens per batch mismatch - longformer vs roberta
未关闭
- 主要语言
- Python
- 星标
- 2.2k
- 派生
- 286
- PR 合并指标
- 30 天内没有已合并 PR
描述
I see in your conversion notebook that you suggest that the number of tokens per batch should be the same as roberta: 2^18 = 260k
When I look at the roberta paper, it says it uses a sequence length of 512 and a batch size of 8k. This means that each batch has 512*8k = 4M tokens
Am I missing something?
贡献指南
这个仓库没有索引到贡献指南
评估
这个 Issue 还没有评估数据。