plotly / plotly/plotly.py

[FEATURE]: Make Plotly Express dataframe implementation-agnostic

オープン
#5,667 コメント 6 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

enhancement feature P2 size: 3
主要言語
Python
スター
18.8k
フォーク
2.8k
平均マージ
16時間 26分
マージ済み PR(30日)
21

説明

[FEATURE]: Make Plotly Express dataframe implementation-agnostic

Description

For a long time, there's been questions around how tightly Plotly Express depends on pandas. For example, here's a program from the Getting Started page that's supposed to show how easy it is to make a simple plot with Plotly Express:

import plotly.express as px

fig = px.bar(x=["a", "b", "c"], y=[1, 3, 2])
fig.show()

But if you run this with only plotly[express] installed, you get an error stating that pandas is required:

Traceback (most recent call last):
  File "/.../plotly/express/_core.py", line 1210, in to_named_series
    import pandas as pd
ModuleNotFoundError: No module named 'pandas'
...
NotImplementedError: Pandas installation is required if no dataframe is provided.

I think I understand this error; Plotly doesn't want to require pandas, when there are other dataframe-based libraries that people are choosing such as Polars. However, this message appears even if you have Polars installed. I think we can make Plotly Express more agnostic toward dataframe providers by looking for an available provider, rather than exiting if pandas is not found.

Why should this feature be added?

Plotly claims to not require pandas, but it still requires pandas in places where other dataframe implementations seem to work.

Mocks/ designs

I made some small changes to plotly/express/_core.py, which allows the above example to work with Polars installed. First, add a new function to look for an existing library that supports dataframes through narwhals:

def get_df_implementation():
    """Get an installed dataframe implementation."""
    for df_implementation in ("pandas", "polars"):
        try:
            return import_module(df_implementation)
        except ImportError:
            pass

    msg = (
        "A dataframe implementation such as Pandas or Polars is required if "
        "no dataframe is provided."
    )
    raise NotImplementedError(msg)

This POC version looks for pandas first, and then Polars.

The only change in the function to_named_series() is to call get_df_implemenation() in the else block, instead of doing its own check for pandas:

def to_named_series(x, name=None, native_namespace=None):
    ...
    if isinstance(x, nw.Series):
        return x.rename(name)
    elif native_namespace is not None:
        return nw.new_series(name=name, values=x, native_namespace=native_namespace)
    else:
        df_implementation = get_df_implementation()
        return nw.new_series(name=name, values=x, native_namespace=df_implementation)

The existing function process_args_into_dataframe() needs one call to get_df_implemenation() as well:

def process_args_into_dataframe(
    args, wide_mode, var_name, value_name, is_pd_like, native_namespace
):
    ...
    length = len(df_output[next(iter(df_output))]) if len(df_output) else 0

    if native_namespace is None:
        native_namespace = get_df_implementation()

    if ranges:
        import numpy as np
        ...

With these changes, the above example runs without errors when Polars is installed. When pandas and Polars are both unavailable, the error states that one of these libraries must be installed.

Notes

I'm not tied to this particular implementation in any way, this was just a proof of concept to see if we could make Plotly more agnostic about the dataframe provider.

Is this a direction that's worth pursuing? I'm happy to work on a PR if it is, or happy to see others more versed in Plotly's internals implement it.

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

plotly/express/_core.py の to_named_series() と process_args_into_dataframe() を読むところから始め、続いて Narwhals がネイティブの dataframe namespace をどのように受け取るかを確認します。pandas が利用できず Polars がインストールされている状態で Getting Started の棒グラフ例を実行し、例が動作することを確認するとともに、どちらのプロバイダーも利用できない場合に、プロバイダーがないことを示すエラーが有用な内容であることを確認します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
pandas, python
領域
data-visualization
issue の種類
機能追加
難易度
4/5
見積もり時間
3〜5日
活発さ
静か
明瞭さ
おおむね明確
初心者へのやさしさ
58/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。