lincc-frameworks / lincc-frameworks/nested-pandas
Wrap pd.merge with nested columns support
Open
Nobody has claimed this yet.
enhancement
interface
LSDB
- Dominant language
- Python
- Stars
- 26
- Forks
- 8
- Avg merge
- 2d 2h
- Merged PRs (30d)
- 9
Description
Feature request
It would be great if we allowed joining dataframes on sub-columns, something like that:
import nested_pandas as npd
merged1 = npd.merge(ndf1, df1, left_on="nested.id", right_on="id")
merged2 = npd.merge(ndf1, ndf2, left_on="nested.id", right_on="lc.oid")
Before submitting
Please check the following:
- I have described the purpose of the suggested change, specifying what I need the enhancement to accomplish, i.e. what problem it solves.
- I have included any relevant links, screenshots, environment information, and data relevant to implementing the requested feature, as well as pseudocode for how I want to access the new functionality.
- If I have ideas for how the new feature could be implemented, I have provided explanations and/or pseudocode and/or task lists for the steps.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue gives example calls to npd.merge using nested paths such as "nested.id" and "lc.oid". Start by locating the existing merge entry point and related tests, then trace how join keys are resolved. Done means the two demonstrated nested-column joins work and are covered by tests.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- pandas, python
- Domain
- data
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100