Why does column take precedence over an object with the same name in data table in cases of join with a subsetted data table

Open
#4,946 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
35/100
Issue type
Bug
Clarity
Needs clarification
Activity status
Stale
Tech stack
r
Domain
data

Research direction

Run the supplied reproducer in R and compare the two working expressions with the failing subsetted join. Start with ?data.table and the join/scoping behavior described in the issue; done means determining whether the precedence is intended and identifying the corresponding documentation or behavior change.

Written by the indexing model from the issue text.

Description

programming

Both operator_descriptions[condition] and operator_conditions[operator_descriptions, on="operator"] work without issue, yet (as you demonstrate) the merge fails.

operator_descriptions <- data.table(operator = "/", description = "/") operator_conditions <- data.table(operator = "/", condition = "Quotient non zero") condition <- c(TRUE) operator_conditions[operator_descriptions[condition], on = c("operator")]

Output Error: When i is a data.table (or character vector), the columns to join by must be specified either using 'on=' argument (see ?data.table) or by keying x (i.e. sorted, and, marked as sorted, see ?setkey). Keyed joins might have further speed benefits on very large data due to x being sorted in RAM.

Despite the fact that condition is in the scope of operator_descriptions, condition column of operator_conditions seems to take precedence over the condition object defined in previous statement.

This behavior requires one to be cautious about names etc. While in normal cases, this is fine as we intend to use column names mostly. But in cases like these, I am not sure if it is the intention.

# Output of sessionInfo()
R version 3.4.4 (2018-03-15) Platform: x86_64-pc-linux-gnu (64-bit) Running under: Ubuntu 16.04.7 LTS

Dominant language
R
Stars
3.9k
Forks
1.1k
Avg merge
14h 4m
Merged PRs (30d)
4

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Rdatatable/data.table

All issues in Rdatatable/data.table

Similar issues

More R issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.