dcast naming: 1 vs. 2+ value.var

Open
#6,582 3 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Stale
Tech stack
r
Domain
data

Research direction

Start with the dcast examples in the issue and inspect how value.var affects generated column names for one versus multiple columns. Define whether the consistent foo_1/foo_2 naming should become the default or be controlled by an argument, then verify the requested output for both examples.

Written by the indexing model from the issue text.

Description

consistency reshape

Suppose I have a datatable with columns id, part, etc. & want to widen it with id ~ part.

  1. If there is a single value.var column, the input & output look like this:
      id  part   foo
   <num> <num> <int>
1:     1     1     1
2:     1     2     2
3:     2     1     3
      id     1     2
   <num> <int> <int>
1:     1     1     2
2:     2     3    NA
  1. If there are multiple value.var columns, the input & output look like this:
      id  part   foo   bar
   <num> <num> <int> <int>
1:     1     1     1     4
2:     1     2     2     5
3:     2     1     3     6
      id foo_1 foo_2 bar_1 bar_2
   <num> <int> <int> <int> <int>
1:     1     1     2     4     5
2:     2     3    NA     6    NA

For interactive use this is fine, but when I'm using this code in an automated pipeline, I find myself wishing for a consistent treatment of column names, i.e. that the first result would be:

      id foo_1 foo_2
   <num> <int> <int>
1:     1     1     2 
2:     2     3    NA

Could we switch to this treatment or provide an argument for doing so?


library(data.table) # 1.16.2
dt1 <- data.table(id = c(1, 1, 2), part = c(1, 2, 1), foo = 1:3)
dcast(dt1, id ~ part, value.var = 'foo')
dt1[, bar := 4:6]
dcast(dt1, id ~ part, value.var = c('foo', 'bar'))
Dominant language
R
Stars
3.9k
Forks
1.1k
Avg merge
14h 4m
Merged PRs (30d)
4

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Rdatatable/data.table

All issues in Rdatatable/data.table

Similar issues

More R issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.