allenai / allenai/OLMoE

StarCoder data

未關閉
#38 2 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
主要語言
Jupyter Notebook
星號
1.1k
分支
120
PR 合併指標
30 天內沒有已合併 PR

描述

> I’m reaching out to ask about the preprocessing applied to the StarCoder data. The OLMoE paper says that “For the StarCoder subset, we also remove any document from a repository with fewer than 2 stars on GitHub, whose most frequent word constitutes over 30% of the document, or whose top-2 most frequent words constitute over 50% of the document.”
>
> Would it be possible for you to share the code used to compute these most frequent word statistics? I’ve tried several approaches, but I haven’t been able to reproduce the exact counts reported in the metadata. Having access to the original implementation would be extremely helpful for our work, and I will make sure to acknowledge your help in our paper once it is published.

Got this question from Tao; I think that code is accessible via https://github.com/allenai/dolma ; also cc @soldni

貢獻指南

這個儲存庫沒有索引到貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。