allenai / allenai/ir_datasets

Accept directories extracted from original compressed files

オープン
#60 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る
enhancement
主要言語
Python
スター
391
フォーク
58
PR マージ指標
30日以内にマージされた PR はありません

説明

Currently irds requires the original files from a dataset, such as a tar.gz file for the NYT corpus. It would be nice if the extracted directory could also be provided as input instead. This would be useful in cases where the original compressed file wasn't kept (as happened to me with NYT). Taking the directory as input is also closer to what most other tools do, so it potentially removes the need to have both the compressed file and its contents available.

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。