[BUG]: px.sunburst / px.treemap / px.icicle with path give a different sector order on every run for Polars DataFrames
Ninguém assumiu esta issue ainda.
- Linguagem predominante
- Python
- Estrelas
- 18.8k
- Forks
- 2.8k
- Merge médio
- 16h 26min
- PRs com merge (30d)
- 21
Descrição
Description
When path= is used with a Polars DataFrame, px.sunburst, px.treemap and px.icicle build their ids / labels / parents / values arrays in a different order every time the script is run. The same data as a pandas DataFrame or a PyArrow table always gives the same order: the order in which the sectors first appear in the data.
The cause is in process_dataframe_hierarchy (plotly/express/_core.py). Each level of the hierarchy is built with df.group_by(path[i:]).agg(...). With pandas (narwhals uses sort=False) and PyArrow the groups come back in order of first appearance, but Polars' group_by does not guarantee any order, so the order of the output changes from run to run.
Consequences:
fig.to_json()/fig.write_html()output is not reproducible with Polars input (snapshot tests, caching, diffs of generated HTML).- With
sort=False, or when sectors have equal values, the chart itself is laid out differently on each run. - Polars results differ from pandas / PyArrow results for identical data.
Screenshots/Video
N/A: the difference is in the figure data; see the output below.
Steps to reproduce
import plotly
import plotly.express as px
import polars as pl
df = pl.DataFrame(
{
"region": ["South", "North", "South", "West", "North", "West"],
"sector": ["Tech", "Finance", "Finance", "Tech", "Tech", "Finance"],
"sales": [1, 2, 3, 4, 5, 6],
}
)
fig = px.sunburst(df, path=["region", "sector"], values="sales")
print(plotly.__version__, pl.__version__, list(fig.data[0].ids))
Running the script three times (plotly 7.1.0, polars 1.44.2):
7.1.0 1.44.2 ['West/Tech', 'West/Finance', 'North/Finance', 'South/Finance', 'North/Tech', 'South/Tech', 'South', 'North', 'West']
7.1.0 1.44.2 ['South/Tech', 'West/Tech', 'North/Tech', 'South/Finance', 'North/Finance', 'West/Finance', 'West', 'North', 'South']
7.1.0 1.44.2 ['South/Tech', 'West/Tech', 'North/Tech', 'North/Finance', 'South/Finance', 'West/Finance', 'West', 'South', 'North']
With pd.DataFrame(...) instead, every run prints:
['South/Tech', 'North/Finance', 'South/Finance', 'West/Tech', 'North/Tech', 'West/Finance', 'South', 'North', 'West']
Notes
I have a small fix with a regression test and will open a PR for it.
Guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Direção de pesquisa
Comece em plotly/express/_core.py, em process_dataframe_hierarchy, e rastreie as chamadas a group_by usadas para gráficos baseados em caminhos. Compare a ordenação hierárquica de Polars, pandas e PyArrow e, em seguida, examine o teste de regressão incluído ou proposto pela pessoa que reportou o problema. Considera-se concluído quando execuções repetidas de Polars produzirem uma ordenação estável pela primeira ocorrência, consistente com as outras entradas compatíveis.
Escrita pelo modelo de indexação a partir do texto da issue.
Avaliação
- Stack de tecnologia
- python
- Domínio
- data-visualization
- Tipo de issue
- Bug
- Dificuldade
- 3/5
- Tempo estimado
- 1-2 dias
- Status de atividade
- Ativa
- Clareza
- Claramente especificada
- Facilidade para iniciantes
- 35/100