Add syntax for "subsetting join"
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 30/100
Research direction
Start with the X[Y, on=names(Y), nomatch=0] and X[!Y, on=names(Y)] examples in the issue, and inspect the existing join and subsetting entry points. Done should provide syntax that subsets X to unique matching rows, preserves X's order, and avoids duplicates caused by repeated rows in Y.
Written by the indexing model from the issue text.
Description
With X[!Y], we can subset to where X has rows not matching Y, but there is no analogue for subsetting to where X does match Y. A X[Y, nomatch=0] join will sort the result according to Y and recognize dupe rows in Y, so I need to do something like X[sort(unique(X[Y, nomatch=0, which=TRUE]))] instead (unless I'm forgetting some other way).
library(data.table)
X = data.table(id = c(3L, 1L, 2L, 1L, 1L), g = c("A", "A", "B", "B", "A"), v = (1:5)*10)
Y = data.table(id = c(1L, 1:3), g = "A")
X[Y, on=names(Y), nomatch=0]
# gives id = 1 before 3, contrary to X's ordering
# gives id = 1 twice, reflecting Y, but the goal is to subset X
# id g v
# 1: 1 A 20
# 2: 1 A 50
# 3: 1 A 20
# 4: 1 A 50
# 5: 3 A 10
X[ sort(unique(X[Y, on=names(Y), nomatch=0, which=TRUE])) ]
# desired result
# id g v
# 1: 3 A 10
# 2: 1 A 20
# 3: 1 A 50
X[!Y, on=names(Y)]
# analogous much simpler code for not join
# id g v
# 1: 2 B 30
# 2: 1 B 40
So I'm looking for new syntax to make this less awkward, maybe something like
X[Y, on=names(Y), subset.join = TRUE]
# or
X[subset.join(Y), on=names(Y)]
- Dominant language
- R
- Stars
- 3.9k
- Forks
- 1.1k
- Avg merge
- 14h 4m
- Merged PRs (30d)
- 4
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Rdatatable/data.table
-
as.data.table() recurses without end on a survival::Surv object (or any data.frame carrying one) Open
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Rdatatable/data.table#7887 ·
-
consistency tests
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Rdatatable/data.table#7853 · 3 comments ·
-
internals
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
Rdatatable/data.table#6938 · 1 comment ·
-
encoding fread
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
Rdatatable/data.table#5179 · 8 comments ·
-
documentation programming
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Rdatatable/data.table#3199 · 3 comments ·
All issues in Rdatatable/data.table
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
r-lib/pkgdepends#485 · 3 comments ·
-
Difficulty 1/5 Under an hour Newbie friendliness 92/100
-
beginners blocker
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
enviPathR OpenBuild Error Build OK Build Warning policies-accepted pre-review precheck-passed
Difficulty 1/5 Under an hour Newbie friendliness 84/100
Bioconductor/BiocContributions#207 · 6 comments ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
datacarpentry/semester-biology#1255 ·