apache / apache/arrow

[R] Implement anonymous functions in calls to dplyr::across

Open
#43,207 5 comments 0 reactions 0 assignees View on GitHub
Component: R good-second-issue Status: needs champion Type: usage
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 13h
Merged PRs (30d)
88

Description

### Describe the usage question you have. Please include as many useful details as possible.

Hi Arrow devs.

I wanted to ask about something I noticed about using the column-wise operators with `dplyr` in `arrow` tables.

If I had an arrow table, and I wanted to run a basic function such as `mean`, `max`, or `min` using `summarize`, it appears that `arrow` does not currently accept the `na.rm = TRUE` argument, or that if it does, I can't seem to find it in the documentation.

Say I took the original dataset:

| Participant | Rating |
| ------------ | -------- |
| Donna | 17 |
| Donna | NA |
| Greg | 21 |
| Greg | NA |

If these were generic `R` dataframes, either of these two calls would work (though one is deprecated):

```
data.frame(
Participant = c('Greg', 'Greg', 'Donna', 'Donna'),
Rating = c(21, NA, 17, NA)
) |>
group_by(Participant) |>
summarize(across(matches("Rating"), \(x) max(x, na.rm = TRUE))) |>
as.data.frame()

data.frame(
Participant = c('Greg', 'Greg', 'Donna', 'Donna'),
Rating = c(21, NA, 17, NA)
) |>
group_by(Participant) |>
summarize(across(matches("Rating"), max, na.rm = TRUE)) |>
as.data.frame()

```
Producing:
| Participant | Rating |
| ------------ | -------- |
| Donna | 17 |
| Greg | 21 |

However, when I run the same commands as an arrow table, both throw errors:

```
data.frame(
Participant = c('Greg', 'Greg', 'Donna', 'Donna'),
Rating = c(21, NA, 17, NA)
) |>
as_arrow_table() |>
group_by(Participant) |>
summarize(across(matches("Rating"), \(x) max(x, na.rm = TRUE))) |>
as.data.frame()

Error in `across_setup()`:
! Anonymous functions are not yet supported in Arrow
Run `rlang::last_trace()` to see where the error occurred.

data.frame(
Participant = c('Greg', 'Greg', 'Donna', 'Donna'),
Rating = c(21, NA, 17, NA)
) |>
as_arrow_table() |>
group_by(Participant) |>
summarize(across(matches("Rating"), max, na.rm = TRUE)) |>
as.data.frame()

Error in `expand_across()`:
! `...` argument to `across()` is deprecated in dplyr and not supported in Arrow
Run `rlang::last_trace()` to see where the error occurred.
```
And the one that does work:
```
data.frame(
Participant = c('Greg', 'Greg', 'Donna', 'Donna'),
Rating = c(21, NA, 17, NA)
) |>
as_arrow_table() |>
group_by(Participant) |>
summarize(across(matches("Rating"), max)) |>
as.data.frame()
```

Returns `NA` values that are not what I want:
| Participant | Rating |
| ------------ | -------- |
| Donna | NA |
| Greg | NA |

Is there a way to pass the `na.rm = TRUE` argument to this call without having to manually drop the `NA` values for each column or row of interest I have in my data?

### Component(s)

R

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.