Add syntax for "subsetting join"

Open
#2,158 3 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
30/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Stale
Tech stack
r
Domain
data

Research direction

Start with the X[Y, on=names(Y), nomatch=0] and X[!Y, on=names(Y)] examples in the issue, and inspect the existing join and subsetting entry points. Done should provide syntax that subsets X to unique matching rows, preserves X's order, and avoids duplicates caused by repeated rows in Y.

Written by the indexing model from the issue text.

Description

joins

With X[!Y], we can subset to where X has rows not matching Y, but there is no analogue for subsetting to where X does match Y. A X[Y, nomatch=0] join will sort the result according to Y and recognize dupe rows in Y, so I need to do something like X[sort(unique(X[Y, nomatch=0, which=TRUE]))] instead (unless I'm forgetting some other way).

library(data.table)
X = data.table(id = c(3L, 1L, 2L, 1L, 1L), g = c("A", "A", "B", "B", "A"), v = (1:5)*10)
Y = data.table(id = c(1L, 1:3), g = "A")

X[Y, on=names(Y), nomatch=0] 
# gives id = 1 before 3, contrary to X's ordering
# gives id = 1 twice, reflecting Y, but the goal is to subset X

#    id g  v
# 1:  1 A 20
# 2:  1 A 50
# 3:  1 A 20
# 4:  1 A 50
# 5:  3 A 10


X[ sort(unique(X[Y, on=names(Y), nomatch=0, which=TRUE])) ]
# desired result

#    id g  v
# 1:  3 A 10
# 2:  1 A 20
# 3:  1 A 50


X[!Y, on=names(Y)]
# analogous much simpler code for not join

#    id g  v
# 1:  2 B 30
# 2:  1 B 40

So I'm looking for new syntax to make this less awkward, maybe something like

X[Y, on=names(Y), subset.join = TRUE]
# or
X[subset.join(Y), on=names(Y)]
Dominant language
R
Stars
3.9k
Forks
1.1k
Avg merge
14h 4m
Merged PRs (30d)
4

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Rdatatable/data.table

All issues in Rdatatable/data.table

Similar issues

More R issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.