Memory usage of plotnine for big dataframe contain millions rows
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 4.8k
- Forks
- 254
- Avg merge
- 10h 35m
- Merged PRs (30d)
- 9
Description
There is a high memory usage for plotnine in _draw_layers step for big dataframe. And generate_data method in layer will copy the original data (self.data = plot_data.copy()). If this is true deep copy, it will cost many time and memory, especially for those contain many columns that weren't used for plot.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading the _draw_layers step and the generate_data method in layer, focusing on the self.data = plot_data.copy() operation. Profile memory and runtime with a dataframe containing millions of rows and many unused columns, then verify that the resulting behavior avoids unnecessary memory usage without changing plotting results.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data-visualization
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100