allenai / allenai/specter

Matching articles from SPECTER's dataset with S2ORC IDs

オープン
#37 コメント 1 件 リアクション 1 件 担当者 0 名 GitHub で見る
主要言語
Python
スター
592
フォーク
58
PR マージ指標
30日以内にマージされた PR はありません

説明

Hi,

I want to match articles used in SPECTER's training and validation sets with the articles from S2ORC.
The problem is that article IDs in SPECTER's [training and validation sets](https://github.com/allenai/specter/issues/2) are not used in the S2ORC dataset, i.e., S2ORC uses different paper IDs compared to SPECTER.

For example, [this article](https://www.semanticscholar.org/paper/793efec2096f6511c45430ff5f2f08a362dcf3eb) can be found in SPECTER's validation set and its ID there is: `793efec2096f6511c45430ff5f2f08a362dcf3eb`.
Corpus ID of this paper is `11967120` and this Corpus ID is used in S2ORC as `paper_id`. (I've found this Corpus ID on the Semantic Scholar's webpage linked above)

Is there any easy way to obtain these Corpus IDs for articles from SPECTER's dataset?
I'm aware I could use Semantic Scholar's API for this, but I think that would be very time-consuming (SPECTER's dataset contains over 165k unique article IDs if I calculated correctly).

Thanks!

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。