StarCoder data
- Lingua principale
- Jupyter Notebook
- Stelle
- 1.1k
- Fork
- 120
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
> I’m reaching out to ask about the preprocessing applied to the StarCoder data. The OLMoE paper says that “For the StarCoder subset, we also remove any document from a repository with fewer than 2 stars on GitHub, whose most frequent word constitutes over 30% of the document, or whose top-2 most frequent words constitute over 50% of the document.”
>
> Would it be possible for you to share the code used to compute these most frequent word statistics? I’ve tried several approaches, but I haven’t been able to reproduce the exact counts reported in the metadata. Having access to the original implementation would be extremely helpful for our work, and I will make sure to acknowledge your help in our paper once it is published.
Got this question from Tao; I think that code is accessible via https://github.com/allenai/dolma ; also cc @soldni
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Valutazione
Questa issue non è ancora stata valutata.