allenai / allenai/OLMoE

StarCoder data

Aperta
#38 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Jupyter Notebook
Stelle
1.1k
Fork
120
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

> I’m reaching out to ask about the preprocessing applied to the StarCoder data. The OLMoE paper says that “For the StarCoder subset, we also remove any document from a repository with fewer than 2 stars on GitHub, whose most frequent word constitutes over 30% of the document, or whose top-2 most frequent words constitute over 50% of the document.”
>
> Would it be possible for you to share the code used to compute these most frequent word statistics? I’ve tried several approaches, but I haven’t been able to reproduce the exact counts reported in the metadata. Having access to the original implementation would be extremely helpful for our work, and I will make sure to acknowledge your help in our paper once it is published.

Got this question from Tao; I think that code is accessible via https://github.com/allenai/dolma ; also cc @soldni

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.