allenai / allenai/longformer

Is it possible to finetune the pretrained model on casual language modeling or text summarization?

未关闭
#29 11 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
2.2k
派生
285
PR 合并指标
30 天内没有已合并 PR

描述

Hi,
Thanks for providing and presenting this nice work.

As mentioned in your paper, your attention pattern for modeling long sequences can be plugged into any pretrained transformer model.
I wonder if this repo covers code to finetune a pretrained LM (e.g. gpt-2) or your own released pretrained model on a new dataset for language modeling task?

If so, is it possible through PyTorch implementation or the CUDA kernel?
I would appreciate if you can guide me in this respect.

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。