plotly / plotly/plotly.py

make generated plot HTML reproducible

未關閉
#3,393 2 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

feature P3
主要語言
Python
星號
18.8k
分支
2.8k
平均合併
16 小時 26 分鐘
30 天內合併 PR
21

描述

The HTML/js contains UUIDs generated by python's built-in uuid.uuid4. By nature, those are unique and totally random.

But if you generate documents as part of a pipeline (e.g. via https://dvc.org/ ) then it would be great to have them reproducible down to the byte.

import plotly.express as px

fig =px.scatter(x=range(10), y=range(10))
fig.write_html("file1.html")
fig.write_html("file2.html")

The files will be identical except from the UUIDs.

I "solved" this problem by monkey-patching the relevant function:

# monkey patch
import uuid
import random

uuid4_rnd = random.Random(0)
def uuid4_seeded():
    """Generate a random UUID using random.Random"""
    return uuid.UUID(bytes=uuid4_rnd.randbytes(16), version=4)

uuid.uuid4 = uuid4_seeded

fig =px.scatter(x=range(10), y=range(10))
fig.write_html("file.html")

re-executing this snippet will produce identical files.

It would be great to exchange one function (or set a seed for an internal version of uuid) in plotly instead of hijacking the python built-in.

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

fig.write_html 入口點開始,追蹤 UUID 被插入產生的 HTML/JavaScript 的位置。比較對範例 figure 的重複寫入,並確定 Plotly 應如何提供確定性的 UUID 產生或 seed 設定。完成標準是:無需對 Python 的 uuid.uuid4 進行 monkey-patching,即可產生逐位元組完全相同的 HTML 檔案。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
data-visualization
Issue 類型
功能
難度
4/5
預估耗時
3-5 天
活躍度
停滯
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。