allenai / allenai/ir_datasets

Accept directories extracted from original compressed files

Aberta
#60 2 comentários 0 reações 0 responsáveis Ver no GitHub
enhancement
Linguagem predominante
Python
Estrelas
391
Forks
58
Métricas de merge de PRs
Nenhum PR com merge em 30d

Descrição

Currently irds requires the original files from a dataset, such as a tar.gz file for the NYT corpus. It would be nice if the extracted directory could also be provided as input instead. This would be useful in cases where the original compressed file wasn't kept (as happened to me with NYT). Taking the directory as input is also closer to what most other tools do, so it potentially removes the need to have both the compressed file and its contents available.

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Avaliação

Esta issue ainda não foi avaliada.

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.