python / python/cpython

json.dump(x,f) is much slower than f.write(json.dumps(x))

未關閉
#129,711 7 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

extension-modules performance stdlib type-bug
主要語言
Python
星號
77.2k
分支
36k
PR 合併指標
PR 指標待擷取

描述

Bug report

Bug description:

Experimentally I measured a huge performance improvement when I switched my code from

json.dump(x, f, **)

to

f.write(json.dumps(x, **))
Method

I essentially wrote the same contents to different files sequentially and measured the total amount of time taken. The json contents had 1, 300, and 400 entries per level, and 1, 5, and 6 levels of depth. There's quite a level of variance here but this wasn't what I was trying to measure in the first place. I discovered this by chance, so forgive the lack of precision. I also don't have the source code anymore because I wasn't originally planning to report this discovery.

Results
File Size Consecutive Files dump µs dumps µs
74 1 508 581
74 2 520 541
74 4 1153 1151
74 8 1930 1750
39184 1 6363 1086
39184 2 11261 1821
39184 4 38126 3521
39184 8 80411 6466
468218 1 82821 11921
468218 2 150234 38017
468218 4 302357 42137
468218 8 573450 78545
Conclusion

A cursory investigation into the cpython code suggests that the slow part is the sequential writing of the iterencode yield. The chunks are quite small.

CPython versions tested on:

3.10

Operating systems tested on:

macOS

Linked PRs
  • gh-130076

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

先檢視連結的 PR gh-130076,以及 CPython 中與 iterencode 的連續寫入相關的 JSON 程式碼。重現已回報的 dump-versus-dumps 計時結果,然後確認此變更能在更大的案例中改善 dump 效能,同時不改變輸出。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
backend
Issue 類型
缺陷
難度
4/5
預估耗時
3-5 天
活躍度
停滯
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。