allenai / allenai/OLMoE

StarCoder data

Aberta
#38 2 comentários 0 reações 0 responsáveis Ver no GitHub
Linguagem predominante
Jupyter Notebook
Estrelas
1.1k
Forks
120
Métricas de merge de PRs
Nenhum PR com merge em 30d

Descrição

> I’m reaching out to ask about the preprocessing applied to the StarCoder data. The OLMoE paper says that “For the StarCoder subset, we also remove any document from a repository with fewer than 2 stars on GitHub, whose most frequent word constitutes over 30% of the document, or whose top-2 most frequent words constitute over 50% of the document.”
>
> Would it be possible for you to share the code used to compute these most frequent word statistics? I’ve tried several approaches, but I haven’t been able to reproduce the exact counts reported in the metadata. Having access to the original implementation would be extremely helpful for our work, and I will make sure to acknowledge your help in our paper once it is published.

Got this question from Tao; I think that code is accessible via https://github.com/allenai/dolma ; also cc @soldni

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Avaliação

Esta issue ainda não foi avaliada.

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.