plotly / plotly/plotly.py

px.imshow requires int for facet_col while px.scatter can be str

オープン
#3,296 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

bug P3
主要言語
Python
スター
18.8k
フォーク
2.8k
平均マージ
16時間 26分
マージ済み PR(30日)
21

説明

px.imshow produces an error if the dimension of a facet plot has strings as their labels instead of integers. I ran across this with xarray, however, I believe it is present regardless of the input data structure. The other px graphs, e.g. px.scatter, do allow strings in the facet dimension, so this seems like a bug with imshow. https://plotly.com/python/facet-plots/ has many plotting examples with strings as the facet_col, e.g. "sex=Female" and "sex=Male" from fig = px.scatter(df, x="total_bill", y="tip", color="smoker", facet_col="sex") and many others from px.histogram, px.choropleth, etc.

Versions in use:
python=3.9.1
plotly=4.14.3

Here is an example of the error.

>>> import numpy as np
>>> import xarray as xr
>>> import plotly.express as px
>>> row = xr.Variable('img_row', range(4))
>>> col = xr.Variable('img_col', range(4))
>>> fd = xr.Variable('facet_dim', 'QA QB'.split())
>>> a = [xr.DataArray(np.random.rand(4,4), coords=[row, col]) for i in range(2)]
>>> da_str = xr.concat(a, "facet_dim")
>>> da_var = xr.concat(a, fd)
>>> da_str
<xarray.DataArray (facet_dim: 2, img_row: 4, img_col: 4)>
array([[[0.17020614, 0.23346368, 0.55878844, 0.58773312],
        [0.80092718, 0.25899341, 0.12681188, 0.77129175],
        [0.85068779, 0.48145364, 0.72109667, 0.51325248],
        [0.2659319 , 0.62613397, 0.09588715, 0.59035821]],

       [[0.32991787, 0.59718269, 0.60165123, 0.63523444],
        [0.11699783, 0.19545503, 0.22478829, 0.56217593],
        [0.59725   , 0.34207063, 0.29841437, 0.30079022],
        [0.21422747, 0.74626584, 0.86025186, 0.61071694]]])
Coordinates:
  * img_row  (img_row) int32 0 1 2 3
  * img_col  (img_col) int32 0 1 2 3
Dimensions without coordinates: facet_dim
>>> da_var
<xarray.DataArray (facet_dim: 2, img_row: 4, img_col: 4)>
array([[[0.17020614, 0.23346368, 0.55878844, 0.58773312],
        [0.80092718, 0.25899341, 0.12681188, 0.77129175],
        [0.85068779, 0.48145364, 0.72109667, 0.51325248],
        [0.2659319 , 0.62613397, 0.09588715, 0.59035821]],

       [[0.32991787, 0.59718269, 0.60165123, 0.63523444],
        [0.11699783, 0.19545503, 0.22478829, 0.56217593],
        [0.59725   , 0.34207063, 0.29841437, 0.30079022],
        [0.21422747, 0.74626584, 0.86025186, 0.61071694]]])
Coordinates:
  * img_row    (img_row) int32 0 1 2 3
  * img_col    (img_col) int32 0 1 2 3
  * facet_dim  (facet_dim) <U2 'QA' 'QB'
>>> fig = px.imshow(da_str, facet_col='facet_dim', aspect='equal')
>>> [d.__class__.__name__ for d in fig.data]
['Heatmap', 'Heatmap']
>>> fig = px.imshow(da_var, facet_col='facet_dim', aspect='equal')
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "C:\Users\<NAME>\Miniconda3\envs\jupyter\lib\site-packages\plotly\express\_imshow.py", line 525, in imshow
    col_labels = ["%s=%d" % (slice_label, i) for i in facet_slices]
  File "C:\Users\<NAME>\Miniconda3\envs\jupyter\lib\site-packages\plotly\express\_imshow.py", line 525, in <listcomp>
    col_labels = ["%s=%d" % (slice_label, i) for i in facet_slices]
TypeError: %d format: a number is required, not numpy.str_

I think the culprit is the different string conversions in play. Can the px.imshow col_labels use a more forgiving conversion?
From File "C:\Users\<NAME>\Miniconda3\envs\jupyter\lib\site-packages\plotly\express\_imshow.py", line 525, in imshow
col_labels = ["%s=%d" % (slice_label, i) for i in facet_slices]
From File "C:\Users\<NAME>\Miniconda3\envs\jupyter\lib\site-packages\plotly\express\_core.py", line 1885, in make_figure
col_labels = [prefix + str(s) for s in sorted_group_values[m.grouper]]

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

facet ラベルがフォーマットされる plotly/express/_imshow.py の報告された 525 行目から始め、plotly/express/_core.py に示されているラベル処理と比較します。文字列値の facet_dim を使用して px.imshow の例を再現し、その後、文字列の facet ラベルで TypeError が発生しなくなり、期待される heatmap facet が引き続き生成されることを確認します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python
領域
data-visualization
issue の種類
バグ
難易度
2/5
見積もり時間
1〜3時間
活発さ
停滞
明瞭さ
明確に書かれている
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。