plotly / plotly/plotly.py

make generated plot HTML reproducible

未关闭
#3,393 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

feature P3
主要语言
Python
星标
18.8k
派生
2.8k
平均合并
16 小时 26 分钟
30 天内合并 PR
21

描述

The HTML/js contains UUIDs generated by python's built-in uuid.uuid4. By nature, those are unique and totally random.

But if you generate documents as part of a pipeline (e.g. via https://dvc.org/ ) then it would be great to have them reproducible down to the byte.

import plotly.express as px

fig =px.scatter(x=range(10), y=range(10))
fig.write_html("file1.html")
fig.write_html("file2.html")

The files will be identical except from the UUIDs.

I "solved" this problem by monkey-patching the relevant function:

# monkey patch
import uuid
import random

uuid4_rnd = random.Random(0)
def uuid4_seeded():
    """Generate a random UUID using random.Random"""
    return uuid.UUID(bytes=uuid4_rnd.randbytes(16), version=4)

uuid.uuid4 = uuid4_seeded

fig =px.scatter(x=range(10), y=range(10))
fig.write_html("file.html")

re-executing this snippet will produce identical files.

It would be great to exchange one function (or set a seed for an internal version of uuid) in plotly instead of hijacking the python built-in.

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

fig.write_html 入口点开始,追踪 UUID 被插入生成的 HTML/JavaScript 的位置。比较对示例 figure 的重复写入,并确定 Plotly 应如何提供确定性的 UUID 生成或 seed 设置。完成标准是:无需对 Python 的 uuid.uuid4 进行 monkey-patching,即可生成逐字节完全相同的 HTML 文件。

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
data-visualization
Issue 类型
功能
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。