has2k1 / has2k1/plotnine

Memory usage of plotnine for big dataframe contain millions rows

Open
#483 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Enhancement Performance
Dominant language
Python
Stars
4.8k
Forks
254
Avg merge
10h 35m
Merged PRs (30d)
9

Description

There is a high memory usage for plotnine in _draw_layers step for big dataframe. And generate_data method in layer will copy the original data (self.data = plot_data.copy()). If this is true deep copy, it will cost many time and memory, especially for those contain many columns that weren't used for plot.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading the _draw_layers step and the generate_data method in layer, focusing on the self.data = plot_data.copy() operation. Profile memory and runtime with a dataframe containing millions of rows and many unused columns, then verify that the resulting behavior avoids unnecessary memory usage without changing plotting results.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-visualization
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.