plotly / plotly/plotly.py

significant speedup for "to_scalar_or_list"

未关闭
#2,938 4 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

P3 performance
主要语言
Python
星标
18.8k
派生
2.8k
平均合并
16 小时 26 分钟
30 天内合并 PR
21

描述

hello,

While investigating a slowness in plotly, I have stumbled upon the to_scalar_or_list function (https://github.com/plotly/plotly.py/blob/abd86092e048d5c8b02da65123824b75e4311838/packages/python/plotly/_plotly_utils/basevalidators.py#L30) that was taking much time.
After some tinkering, I came with the two following changes that vastly improves the performance:

So at the end, it looks like

np = get_module("numpy", should_load=False)
pd = get_module("pandas", should_load=False)
# Utility functions
# -----------------
def to_scalar_or_list(v):
    # Handle the case where 'v' is a non-native scalar-like type,
    # such as numpy.float32. Without this case, the object might be
    # considered numpy-convertable and therefore promoted to a
    # 0-dimensional array, but we instead want it converted to a
    # Python native scalar type ('float' in the example above).
    # We explicitly check if is has the 'item' method, which conventionally
    # converts these types to native scalars.

    # check first for the simple case
    if isinstance(v,(int,float,str)):
        return v
    if np and np.isscalar(v) and hasattr(v, "item"):
        return v.item()
    if isinstance(v, (list, tuple)):
        return [to_scalar_or_list(e) for e in v]
    elif np and isinstance(v, np.ndarray):
        if v.ndim == 0:
            return v.item()
        return [to_scalar_or_list(e) for e in v]
    elif pd and isinstance(v, (pd.Series, pd.Index)):
        return [to_scalar_or_list(e) for e in v]
    elif is_numpy_convertable(v):
        return to_scalar_or_list(np.array(v))
    else:
        return v

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

从 packages/python/plotly/_plotly_utils/basevalidators.py 中的 to_scalar_or_list 开始,检查附近对 get_module 和 is_numpy_convertable 的使用。比较 native scalar、列表或元组、NumPy 数组以及 pandas Series 或 Index 值的行为和性能。当提出的案例保留现有转换,同时避免不必要的重复工作时,即表示完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
numpy, pandas, python
领域
performance
Issue 类型
重构
难度
3/5
预计耗时
1-2 天
活跃度
停滞
描述清晰度
描述清楚
新手友好度
38/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。