lincc-frameworks / lincc-frameworks/nested-pandas

Implement a parallel version of `map_rows`

Open
#286 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement LSDB performance
Dominant language
Python
Stars
26
Forks
8
Avg merge
2d 2h
Merged PRs (30d)
9

Description

**Feature request**

It would be cool if we can run UDF in parallel inside the `reduce` call. May also be partially achieved with #277

**Before submitting**
Please check the following:

- [x] I have described the purpose of the suggested change, specifying what I need the enhancement to accomplish, i.e. what problem it solves.
- [ ] I have included any relevant links, screenshots, environment information, and data relevant to implementing the requested feature, as well as pseudocode for how I want to access the new functionality.
- [ ] If I have ideas for how the new feature could be implemented, I have provided explanations and/or pseudocode and/or task lists for the steps.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating map_rows and the reduce call in the Python project, then review issue #277 for related work. Clarify the intended parallel UDF behavior and how it should be accessed; done should include an agreed implementation scope and verification that UDFs run in parallel within reduce.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.