python / python/cpython

json.dump(x,f) is much slower than f.write(json.dumps(x))

Aperta
#129,711 7 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

extension-modules performance stdlib type-bug
Lingua principale
Python
Stelle
77.2k
Fork
35.9k
Metriche di merge delle PR
Metriche PR in attesa

Descrizione

Bug report

Bug description:

Experimentally I measured a huge performance improvement when I switched my code from

json.dump(x, f, **)

to

f.write(json.dumps(x, **))
Method

I essentially wrote the same contents to different files sequentially and measured the total amount of time taken. The json contents had 1, 300, and 400 entries per level, and 1, 5, and 6 levels of depth. There's quite a level of variance here but this wasn't what I was trying to measure in the first place. I discovered this by chance, so forgive the lack of precision. I also don't have the source code anymore because I wasn't originally planning to report this discovery.

Results
File Size Consecutive Files dump µs dumps µs
74 1 508 581
74 2 520 541
74 4 1153 1151
74 8 1930 1750
39184 1 6363 1086
39184 2 11261 1821
39184 4 38126 3521
39184 8 80411 6466
468218 1 82821 11921
468218 2 150234 38017
468218 4 302357 42137
468218 8 573450 78545
Conclusion

A cursory investigation into the cpython code suggests that the slow part is the sequential writing of the iterencode yield. The chunks are quite small.

CPython versions tested on:

3.10

Operating systems tested on:

macOS

Linked PRs
  • gh-130076

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia esaminando la PR collegata gh-130076 e il codice JSON di CPython relativo alle scritture sequenziali da iterencode. Riproduci le misurazioni dei tempi riportate per dump-versus-dumps, quindi verifica che la modifica migliori le prestazioni di dump per i casi più grandi senza modificare l'output.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
backend
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Ferma
Chiarezza
Abbastanza chiara
Idoneità per principianti
35/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.