When were the most recent publications of pre-training data included?
Offen
- Vorherrschende Sprache
- Python
- Sterne
- 1.7k
- Forks
- 232
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beschreibung
I know that SciBERT is pre-trained by the Semantic Scholar corpus. I also know that the Semantic Scholar corpus is not publicly available.
I am wondering how many new papers are included in the pre-training data. For example, **are papers from ACL 2018 included?**
The Semantic Scholar Corpus paper was published in 2018 or so, so I'm guessing that's right around the borderline between having a paper...
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Bewertung
Dieses Issue wurde noch nicht bewertet.