Should .SD include a column used to create an ad-hoc 'by'?

Open
#6,388 11 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
25/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Stale
Tech stack
r
Domain
data

Research direction

Start by running the supplied R example and compare the .SD columns with the columns used to construct the ad-hoc by expression. Trace the .SD/by evaluation behavior, then determine and document a settled compatibility decision with regression coverage for both the current and proposed behavior.

Written by the indexing model from the issue text.

Description

breaking-change
DT = data.table(a = rep(letters, 2), b = 1)
DT[1:20, a := paste0(0, a)]
DT[15:26, a := paste0(a, 1)]
DT[, if (.N == 2L) .SD[1], by = .(trim0 = gsub("^0", "", a))]
#      trim0     b
#     <char> <num>
#  1:      a     1
#  2:      b     1
#  3:      c     1
#  4:      d     1
#  5:      e     1
#  6:      f     1
#  7:      g     1
#  8:      h     1
#  9:      i     1
# 10:      j     1
# 11:      k     1
# 12:      l     1
# 13:      m     1
# 14:      n     1

I might have expected .SD to include a here, since a is not in by -- it' just used in by to generate a new ad hoc column. Such mappings need not be invertible, as is the case here, so we can't be sure a could be recovered from trim0, i.e., the current behavior generates some loss of information.

In this simple table, the workaround is pretty easy (.SD[1] --> .(a = a[1], b = b[1]), but this will quickly get ugly.

If we changed behavior so that a is in .SD, getting the old behavior back (if so desired) would mean doing .SDcols = !"a", i.e., the workaround for that default would be a lot easier/more "canonical".

Then again I'm not sure if it's worth introducing a potentially breaking change for this, or if it's a bug, or if anyone else even agrees about this behavior being strange.

Dominant language
R
Stars
3.9k
Forks
1.1k
Avg merge
14h 4m
Merged PRs (30d)
4

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Rdatatable/data.table

All issues in Rdatatable/data.table

Similar issues

More R issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.