NVIDIA / NVIDIA/cudf

[FEA] Pass column indices as `index_col` in `read_csv`

Open
#15,127 7 comments 0 reactions 0 assignees View on GitHub
0 - Backlog feature request good first issue Python
Dominant language
C++
Stars
9.8k
Forks
1.1k
Avg merge
3d 6m
Merged PRs (30d)
278

Description

If I want to use the `index_col` parameter to set certain columns as indices when reading a csv file, I cannot pass a list of column indices (like in pandas). I can pass a list of column labels though:
```python
cudf.read_csv(filepath, index_col=[0])
KeyError: 'None of [0] are in the columns'

cudf.read_csv(filepath, index_col=['family'])
```
While this is not a huge issue, I imagine the following is a common scenario: You have know that the first 3 columns are index columns, but you don't exactly know how each are spelt (`'date'` vs `'Date'` etc.). In this case, if passing a list of column indices were possible, `index_col=[0,1,2]` would have worked fine; otherwise, you will have to read the file without specifying index columns and set index later (or require trial and error to guess the column labels).

Is it possible for `index_col` to accept list of indices like in pandas?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.