dcast naming: 1 vs. 2+ value.var
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
Research direction
Start with the dcast examples in the issue and inspect how value.var affects generated column names for one versus multiple columns. Define whether the consistent foo_1/foo_2 naming should become the default or be controlled by an argument, then verify the requested output for both examples.
Written by the indexing model from the issue text.
Description
Suppose I have a datatable with columns id, part, etc. & want to widen it with id ~ part.
- If there is a single
value.varcolumn, the input & output look like this:
id part foo
<num> <num> <int>
1: 1 1 1
2: 1 2 2
3: 2 1 3
id 1 2
<num> <int> <int>
1: 1 1 2
2: 2 3 NA
- If there are multiple
value.varcolumns, the input & output look like this:
id part foo bar
<num> <num> <int> <int>
1: 1 1 1 4
2: 1 2 2 5
3: 2 1 3 6
id foo_1 foo_2 bar_1 bar_2
<num> <int> <int> <int> <int>
1: 1 1 2 4 5
2: 2 3 NA 6 NA
For interactive use this is fine, but when I'm using this code in an automated pipeline, I find myself wishing for a consistent treatment of column names, i.e. that the first result would be:
id foo_1 foo_2
<num> <int> <int>
1: 1 1 2
2: 2 3 NA
Could we switch to this treatment or provide an argument for doing so?
library(data.table) # 1.16.2
dt1 <- data.table(id = c(1, 1, 2), part = c(1, 2, 1), foo = 1:3)
dcast(dt1, id ~ part, value.var = 'foo')
dt1[, bar := 4:6]
dcast(dt1, id ~ part, value.var = c('foo', 'bar'))
- Dominant language
- R
- Stars
- 3.9k
- Forks
- 1.1k
- Avg merge
- 14h 4m
- Merged PRs (30d)
- 4
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Rdatatable/data.table
-
as.data.table() recurses without end on a survival::Surv object (or any data.frame carrying one) Open
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Rdatatable/data.table#7887 ·
-
consistency tests
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Rdatatable/data.table#7853 · 3 comments ·
-
internals
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
Rdatatable/data.table#6938 · 1 comment ·
-
encoding fread
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
Rdatatable/data.table#5179 · 8 comments ·
-
documentation programming
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Rdatatable/data.table#3199 · 3 comments ·
All issues in Rdatatable/data.table
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
r-lib/pkgdepends#485 · 3 comments ·
-
Difficulty 1/5 Under an hour Newbie friendliness 92/100
-
beginners blocker
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
enviPathR OpenBuild Error Build OK Build Warning policies-accepted pre-review precheck-passed
Difficulty 1/5 Under an hour Newbie friendliness 84/100
Bioconductor/BiocContributions#207 · 6 comments ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
datacarpentry/semester-biology#1255 ·