Extend dcast() to support different fill= values for different columns
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
Research direction
Start by reading the current dcast() fill handling and the setnafill() entry point, including the existing C logic mentioned in the issue. Define and test named-list fills for multiple value.var columns, including the proposed list(0) semantics and any default behavior. Done means dcast() supports distinct fill values without regressing existing fill= behavior.
Written by the indexing model from the issue text.
Description
As raised in #5980, this would bring us closer to parity with tidyr::pivot_wider()'s values_fill argument. This seems especially useful for length(value.var) > 1L use cases, when the different value.var can quite reasonably need different fill= values.
We are free to choose our own way for the extended argument to work, but the pivot_wider() approach seems clear enough: when fill is a named list, the names are columns and the elements are values to be filled. I'm not sure if {tidyr} supports a default unnamed entry so that you could do, e.g. fill = list(special = 1, 0) so that most columns are filled with 0, except for one special column. The only trouble is what to do with fill = list(0); I think we need to interpret that as "fill with the length-1 list with entry 0", not as "0 by default for all columns".
PS I also wonder if it wouldn't be easier to just re-use setnafill() to handle the fill= argument of dcast(), rather than the current logic in C which essentially duplicates that.
- Dominant language
- R
- Stars
- 3.9k
- Forks
- 1.1k
- Avg merge
- 14h 4m
- Merged PRs (30d)
- 4
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Rdatatable/data.table
-
as.data.table() recurses without end on a survival::Surv object (or any data.frame carrying one) Open
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Rdatatable/data.table#7887 ·
-
consistency tests
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Rdatatable/data.table#7853 · 3 comments ·
-
internals
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
Rdatatable/data.table#6938 · 1 comment ·
-
encoding fread
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
Rdatatable/data.table#5179 · 8 comments ·
-
documentation programming
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Rdatatable/data.table#3199 · 3 comments ·
All issues in Rdatatable/data.table
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
r-lib/pkgdepends#485 · 3 comments ·
-
Difficulty 1/5 Under an hour Newbie friendliness 92/100
-
beginners blocker
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
enviPathR OpenBuild Error Build OK Build Warning policies-accepted pre-review precheck-passed
Difficulty 1/5 Under an hour Newbie friendliness 84/100
Bioconductor/BiocContributions#207 · 6 comments ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
datacarpentry/semester-biology#1255 ·