[VL] Story: Improve PySpark support
Open
enhancement
Pyspark
- Dominant language
- Scala
- Stars
- 1.6k
- Forks
- 657
- Avg merge
- 2d 14h
- Merged PRs (30d)
- 80
Description
### Description
Gluten has basic support on Arrow UDF
https://github.com/apache/incubator-gluten/pull/5462
We need to verify and improve below methods use in PySpark:
Python UDF
Pandas UDF
Arrow UDF
Python UDTF
Contributor guide
Research direction
Start by reading the linked Arrow UDF pull request (apache/incubator-gluten#5462), then trace how PySpark handles Python UDF, Pandas UDF, Arrow UDF, and Python UDTF. Done means the four methods have been verified and any required support improvements are covered by appropriate tests.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, scala, spark
- Domain
- data-engineering, distributed-systems
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100