allenai / allenai/scibert

JNLPBA dataset

Ouverte
#26 3 commentaires 0 réactions 1 personne assignée Réclamée par @kyleclo Voir sur GitHub
Langage dominant
Python
Étoiles
1.7k
Forks
232
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

Hi,

thanks for releasing the `SciBERT` model and datasets :heart:

I'm currently integrating an importing method of the NER data into the [flair](https://github.com/zalandoresearch/flair) library.

I checked the number of imported sentences for all dataset splits and there's a mismatch of 2404 sentences compared to the total sentences number in table 2 of the paper (24,806). Then I checked the JNLPBA dataset and it seems that all `-DOCSTART- O` lines were also counted, which is I think a bit redundant.

The number of training and development sentences is also a bit different than the values reported in the BioBERT paper. The BioBERT uses a split of 14,690 / 3,856 / 3,856, whereas the provided data in this repository uses a split of 16,807 / 1,739 / 3,856. Could you confirm this?

Thanks + regards,

Stefan

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.