apache / apache/arrow

[R]: Unnamed columns cause issues when used in dplyr queries

Open
#40,303 0 comments 0 reactions 1 assignee Claimed by @thisisnic View on GitHub
Component: R Type: bug
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 13h
Merged PRs (30d)
88

Description

### Describe the bug, including details regarding any error messages, version, and platform.

When using the defaults for `write.csv()` the rownames are written and the column header is ``. If we were to read this in with R we would get a column name of X since an empty string is not a valid column name. But Arrow doesn't do this, and presents an empty string, which then causes issues with `rlang:::env_bind0`

```
> write.csv(mtcars, "mtcars.csv", row.names = FALSE)
> open_dataset("matcars.csv", format = "csv") |> mutate(gear_one = gear + 1) |> collect()
Error: IOError: Cannot list directory 'matcars.csv'. Detail: [errno 2] No such file or directory
> write.csv(mtcars, "mtcars.csv")
> open_dataset("mtcars.csv", format = "csv") |> mutate(gear_one = gear + 1) |> collect()
Error in env_bind0(env, data) : attempt to use zero-length variable name
> open_dataset("mtcars.csv", format = "csv") |> collect()
# A tibble: 32 × 12
`` mpg cyl disp hp drat wt qsec vs am gear carb

1 Mazda RX4 21 6 160 110 3.9 2.62 16.5 0 1 4 4
2 Mazda RX4 … 21 6 160 110 3.9 2.88 17.0 0 1 4 4
3 Datsun 710 22.8 4 108 93 3.85 2.32 18.6 1 1 4 1
4 Hornet 4 D… 21.4 6 258 110 3.08 3.22 19.4 1 0 3 1
5 Hornet Spo… 18.7 8 360 175 3.15 3.44 17.0 0 0 3 2
6 Valiant 18.1 6 225 105 2.76 3.46 20.2 1 0 3 1
7 Duster 360 14.3 8 360 245 3.21 3.57 15.8 0 0 3 4
8 Merc 240D 24.4 4 147. 62 3.69 3.19 20 1 0 4 2
9 Merc 230 22.8 4 141. 95 3.92 3.15 22.9 1 0 4 2
10 Merc 280 19.2 6 168. 123 3.92 3.44 18.3 1 0 4 4
# ℹ 22 more rows
# ℹ Use `print(n = ...)` to see more rows
```

Or, if we specifically select all but the column that is the empty string, it works:
```
> open_dataset("mtcars.csv", format = "csv") |> select(mpg, cyl, disp, hp, drat, wt, qsec, vs, am, gear, carb) |> mutate(gear_one = gear + 1) |> collect()
# A tibble: 32 × 12
mpg cyl disp hp drat wt qsec vs am gear carb gear_one

1 21 6 160 110 3.9 2.62 16.5 0 1 4 4 5
2 21 6 160 110 3.9 2.88 17.0 0 1 4 4 5
3 22.8 4 108 93 3.85 2.32 18.6 1 1 4 1 5
4 21.4 6 258 110 3.08 3.22 19.4 1 0 3 1 4
5 18.7 8 360 175 3.15 3.44 17.0 0 0 3 2 4
6 18.1 6 225 105 2.76 3.46 20.2 1 0 3 1 4
7 14.3 8 360 245 3.21 3.57 15.8 0 0 3 4 4
8 24.4 4 147. 62 3.69 3.19 20 1 0 4 2 5
9 22.8 4 141. 95 3.92 3.15 22.9 1 0 4 2 5
10 19.2 6 168. 123 3.92 3.44 18.3 1 0 4 4 5
# ℹ 22 more rows
# ℹ Use `print(n = ...)` to see more rows
```

But, if we don't write the un-named column, this works just fine:
```
> write.csv(mtcars, "mtcars.csv", row.names = FALSE)
> open_dataset("mtcars.csv", format = "csv") |> mutate(gear_one = gear + 1) |> collect()
# A tibble: 32 × 12
mpg cyl disp hp drat wt qsec vs am gear carb gear_one

1 21 6 160 110 3.9 2.62 16.5 0 1 4 4 5
2 21 6 160 110 3.9 2.88 17.0 0 1 4 4 5
3 22.8 4 108 93 3.85 2.32 18.6 1 1 4 1 5
4 21.4 6 258 110 3.08 3.22 19.4 1 0 3 1 4
5 18.7 8 360 175 3.15 3.44 17.0 0 0 3 2 4
6 18.1 6 225 105 2.76 3.46 20.2 1 0 3 1 4
7 14.3 8 360 245 3.21 3.57 15.8 0 0 3 4 4
8 24.4 4 147. 62 3.69 3.19 20 1 0 4 2 5
9 22.8 4 141. 95 3.92 3.15 22.9 1 0 4 2 5
10 19.2 6 168. 123 3.92 3.44 18.3 1 0 4 4 5
# ℹ 22 more rows
# ℹ Use `print(n = ...)` to see more rows
```

### Component(s)

R

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.