python / python/cpython

Speed up JSON string encoding with ensure_ascii=False for long string values

未关闭
#150,878 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

extension-modules performance type-feature
主要语言
Python
星标
77.2k
派生
35.9k
PR 合并指标
PR 指标待抓取

描述

Feature or enhancement

Proposal

When json.dumps runs with ensure_ascii=False, it sizes each escaped string one character at a time in escape_size (Modules/_json.c), after which write_escaped_unicode copies the string verbatim when nothing needs escaping. In this mode a character needs escaping only when c == '"', c == '\\', or c < 0x20; non-ASCII is kept verbatim. For a long string with no such character, which is the common case for text values including Western-European (Latin-1) text, that per-character sizing scan is pure overhead before the verbatim copy.

The proposal is to detect the no-escape case on the one-byte representation eight bytes at a time, returning the verbatim size after about one eighth of the work. A length guard keeps short strings, such as the typical dict key, on the existing per-character loop. Two-byte and four-byte strings (anything with a character above U+00FF) keep the current loop.

This is the ensure_ascii=False counterpart to the encoder change in #150875 (PR #150876); together with the decode-side scan in #150871 (PR #150872) the three cover JSON string scanning end to end. They touch different code paths and are separate changes.

How this differs from the SIMD backend in #142915

It is not the SIMD parsing architecture declined in #142915. It uses no SIMD intrinsics, no runtime CPU detection, and no build configuration, only portable 64-bit integer arithmetic with the same 0x0101… / 0x8080… masks that Objects/unicodeobject.c already applies for ASCII scanning. It changes one function and adds no infrastructure, so it does not depend on #125022 and needs no PEP.

When it helps, and when it does not

Measured json.dumps(..., ensure_ascii=False) speedups against the current encoder:

Document shape Effect
One long text field (~16 KB string) 5.8x faster
Long Western-European (Latin-1) text values 4.2x faster
Many 200-character ASCII string values 3.9x faster
Realistic mixed records (short and medium strings) 1.4x faster
Short keys, strings that need escaping no change
Strings with characters above U+00FF no change (scalar path)

The benefit applies only to ensure_ascii=False, which is the non-default mode, so it reaches fewer callers than the default-path change in #150876; within that mode the win matches.

Correctness

The encoded output is byte-identical to the current encoder. A patch is validated against test_json and a 199-case differential corpus that places each escape-relevant character at every offset across the eight-byte window, in both ensure_ascii modes. Every output matched.

A draft PR follows.

Benchmark

Built base and patched interpreters from this branch's main ancestor and the patch, ran the same script under each, and compared with pyperf compare_to (A/B by swapping Lib/json/encoder.py on the same build; macOS arm64, non-PGO).

import json, pyperf
d = lambda o: json.dumps(o, ensure_ascii=False)
objs = {
 "long_ascii":   [("x"*200) for _ in range(200)],
 "long_latin1":  [("café résumé naïve "*15) for _ in range(200)],   # 1-byte Latin-1, kept verbatim
 "text_blob":    {"body": "lorem ipsum dolor "*900},
 "short_keys":   {f"k{i}": i for i in range(2000)},
 "nonascii":     ["中文 текст 😀 "*30 for _ in range(200)],          # UCS-2/4 scalar
 "mixed_real":   [{"id":i,"name":f"user_{i}","bio":"hello world "*10} for i in range(300)],
}
r = pyperf.Runner()
for n,o in objs.items():
    r.bench_func(f"dumpsF/{n}", lambda o=o: d(o))
Linked PRs
  • gh-150879

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

拟议的更改围绕 Modules/_json.c 中的 escape_size 展开;首先阅读该函数及其周围的编码器代码,然后运行 test_json 测试套件。完成的标准是 ensure_ascii=False 的输出保持逐字节一致,包括所述的差分用例,同时基准测试的长字符串路径得到改进,并且不改变需要转义的字符串。

由索引模型根据 Issue 内容生成。

评估

技术栈
c, python
领域
backend
Issue 类型
功能
难度
3/5
预计耗时
1-2 天
活跃度
停滞
描述清晰度
描述清楚
新手友好度
25/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。