allenai / allenai/scibert

JNLPBA dataset

Aperta
#26 3 commenti 0 reazioni 1 assegnatario Rivendicata da @kyleclo Vedi su GitHub
Lingua principale
Python
Stelle
1.7k
Fork
232
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

Hi,

thanks for releasing the `SciBERT` model and datasets :heart:

I'm currently integrating an importing method of the NER data into the [flair](https://github.com/zalandoresearch/flair) library.

I checked the number of imported sentences for all dataset splits and there's a mismatch of 2404 sentences compared to the total sentences number in table 2 of the paper (24,806). Then I checked the JNLPBA dataset and it seems that all `-DOCSTART- O` lines were also counted, which is I think a bit redundant.

The number of training and development sentences is also a bit different than the values reported in the BioBERT paper. The BioBERT uses a split of 14,690 / 3,856 / 3,856, whereas the provided data in this repository uses a split of 16,807 / 1,739 / 3,856. Could you confirm this?

Thanks + regards,

Stefan

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.