allenai / allenai/ir_datasets

Direct access to all doc_ids

Aberta
#184 5 comentários 0 reações 0 responsáveis Ver no GitHub
Linguagem predominante
Python
Estrelas
391
Forks
58
Métricas de merge de PRs
Nenhum PR com merge em 30d

Descrição

This is something I was expecting to be quite straightforward (or at least better documented in the API) but it doesn't seem to be.
Say I want to gather all doc_ids from a given corpus (for instance, if I want to use a random negative sampler on run time).
Currently, this is what I do:
```
data = ir_datasets.load("msmarco-document/train")
all_doc_ids = list(data.docs._handler.docs_store().lookup.idx())
```
which is fine, but, from what I can get, this triggers an iteration over all docs in the collection (and is also not very intuitive).

Is there a better way to achieve this?

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Avaliação

Esta issue ainda não foi avaliada.

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.