rolling join keeps only one `on` column, takes name from x table and values from i table

Open
#4,005 7 comments 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
35/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Stale
Tech stack
r
Domain
data

Research direction

Start with the reproducible R example and the rolling-join entry point x[y, roll = "nearest", on = .(x2 = y2)]. Investigate how the join selects and names the matching columns, then define and verify behavior that preserves both columns or uses the original-value column name.

Written by the indexing model from the issue text.

Description

joins non-equi joins

This got me very confused :

library(data.table)
x <- data.table(x1=letters[1:3], x2=c(10,20,30))
y <- data.table(y1=letters[4:6], y2=c(11,21,31))
y2 <- x[y, roll = "nearest", on = .(x2 = y2)]
y2
#>    x1 x2 y1
#> 1:  a 11  d
#> 2:  b 21  e
#> 3:  c 31  f

I would much prefer to keep both x2 and y2, and if we must keep only one column I would much rather have the column name fit the original values.

The workaround I found looks quite awful, can I do better ?

x_ <- copy(x)
x_[, x3 := x2]
y2 <- x_[y, roll = "nearest", on = .(x3 = y2)]
y2[, y2 := x3]
y2[, x3 := NULL]
rm(x_)
y2
#>    x1 x2 y1 y2
#> 1:  a 10  d 11
#> 2:  b 20  e 21
#> 3:  c 30  f 31
Dominant language
R
Stars
3.9k
Forks
1.1k
Avg merge
14h 4m
Merged PRs (30d)
4

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Rdatatable/data.table

All issues in Rdatatable/data.table

Similar issues

More R issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.