bigscience-workshop / bigscience-workshop/promptsource

Find a way to not load all the tasks infos.

Open
#709 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
3k
Forks
375
PR merge metrics
No merged PRs in 30d

Description

When running `from promptsource.seqio_tasks import tasks` it takes a huge amount of time. One of the main reasons is this queries all dataset infos: https://github.com/bigscience-workshop/promptsource/blob/dba1d41e63a7af883fd7dc2727b4c7fd03e714c9/promptsource/seqio_tasks/tasks.py#L84 This is problematic for two reasons:
- One has to load ALL dataset infos as soon as one uses one task.
- Even when cached, it still queries urls to check that it didn't change. One can bypass this point by passing `HF_DATASETS_OFFLINE=1` as described in https://github.com/bigscience-workshop/promptsource/issues/703#issuecomment-1003061062

IMO both are unnecessary and should be fixed. Is there a reasons why one cannot load seqio tasks dynamically, in the sense of fetching only what is necessary? Something along the lines of:

```python
def add_seqio_task(task_name):
seqio.TaskRegistry.add(...)
```

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.