StarCoder data
- Linguagem predominante
- Jupyter Notebook
- Estrelas
- 1.1k
- Forks
- 120
- Métricas de merge de PRs
- Nenhum PR com merge em 30d
Descrição
> I’m reaching out to ask about the preprocessing applied to the StarCoder data. The OLMoE paper says that “For the StarCoder subset, we also remove any document from a repository with fewer than 2 stars on GitHub, whose most frequent word constitutes over 30% of the document, or whose top-2 most frequent words constitute over 50% of the document.”
>
> Would it be possible for you to share the code used to compute these most frequent word statistics? I’ve tried several approaches, but I haven’t been able to reproduce the exact counts reported in the metadata. Having access to the original implementation would be extremely helpful for our work, and I will make sure to acknowledge your help in our paper once it is published.
Got this question from Tao; I think that code is accessible via https://github.com/allenai/dolma ; also cc @soldni
Guia de contribuição
Nenhum guia de contribuição indexado para este repositório
Avaliação
Esta issue ainda não foi avaliada.