Rolling join on multiple columns
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Stale
- Tech stack
- r
- Domain
- data-engineering
Research direction
Start from the proposed buckets[dt, on=c("BinA"="A", "BinB"="B"), roll=c(-Inf, -Inf)] call and compare it with the shown result. Define the expected multi-column rolling semantics, including the supplied example, before locating the existing rolling-join implementation and relevant tests. Done means the example produces the requested BucketID, ID, BinA, and BinB values.
Written by the indexing model from the issue text.
Description
I have a feature request related to my SO post here. I would like to be able to be able to do a rolling join, where the roll applies to multiple columns.
Example
dt <- data.table(ID=1:5, A=c(1.3, 1.7, 2.4, 0.9, 0.6), B=c(3.3, 2.9, 3.0, 0.2, 0.2))
buckets <- CJ(BinA=as.numeric(1:4), BinB=as.numeric(1:4))
buckets[, BucketID := .I]
dt
ID A B
1: 1 1.3 3.3
2: 2 1.7 2.9
3: 3 2.4 3.0
4: 4 0.9 0.2
5: 5 0.6 0.2
buckets
BinA BinB BucketID
1: 1 1 1
2: 1 2 2
3: 1 3 3
4: 1 4 4
5: 2 1 5
6: 2 2 6
7: 2 3 7
8: 2 4 8
9: 3 1 9
10: 3 2 10
11: 3 3 11
12: 3 4 12
13: 4 1 13
14: 4 2 14
15: 4 3 15
16: 4 4 16
# I want to be able to do something like this
buckets[dt, on=c("BinA"="A", "BinB"="B"), roll=c(-Inf, -Inf)]
# And get back this
result <- data.table(BucketID=c(8, 7, 12, 1, 1), ID=1:5, BinA=c(1.3, 1.7, 2.4, 0.9, 0.6), BinB=c(3.3, 2.9, 3.0, 0.2, 0.2))
result
BucketID ID BinA BinB
1: 8 1 1.3 3.3
2: 7 2 1.7 2.9
3: 12 3 2.4 3.0
4: 1 4 0.9 0.2
5: 1 5 0.6 0.2
Can you get this done by tomorrow @arunsrinivasan? ... JK : )
- Dominant language
- R
- Stars
- 3.9k
- Forks
- 1.1k
- Avg merge
- 14h 4m
- Merged PRs (30d)
- 4
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Rdatatable/data.table
-
as.data.table() recurses without end on a survival::Surv object (or any data.frame carrying one) Open
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Rdatatable/data.table#7887 ·
-
consistency tests
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Rdatatable/data.table#7853 · 3 comments ·
-
internals
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
Rdatatable/data.table#6938 · 1 comment ·
-
encoding fread
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
Rdatatable/data.table#5179 · 8 comments ·
-
documentation programming
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Rdatatable/data.table#3199 · 3 comments ·
All issues in Rdatatable/data.table
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
r-lib/pkgdepends#485 · 3 comments ·
-
Difficulty 1/5 Under an hour Newbie friendliness 92/100
-
beginners blocker
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
enviPathR OpenBuild Error Build OK Build Warning policies-accepted pre-review precheck-passed
Difficulty 1/5 Under an hour Newbie friendliness 84/100
Bioconductor/BiocContributions#207 · 6 comments ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
datacarpentry/semester-biology#1255 ·