Keep PyLong loop carries as twodigits in shifts and division
还没有人认领这个 Issue。
- 主要语言
- Python
- 星标
- 77.2k
- 派生
- 35.9k
- PR 合并指标
- PR 指标待抓取
描述
Feature or enhancement
Proposal:
Several functions in longobject.c (v_lshift, v_rshift, and x_divrem) narrowed a loop carry value to digit or sdigit, then widened it again on the next iteration.
On AArch64, that 64-to-32-to-64 conversion inserts an extra mov on the loop-carried critical path. Keeping the carry at two-digit width until the function returns drops that mov and shortens the carry chain. Results are unchanged for valid limbs.
Use v_lshift on AArch64 as an example. The code inside the loop is:
ldr w0, [x4, x2, lsl #2] ; a[i]
mov w3, w3 ; carry chain
lsl x0, x0, x24 ; a[i] << d
orr x0, x0, x3 ; carry chain
and w1, w0, #0x3fffffff
ubfx x3, x0, #30, #32 ; carry chain
str w1, [x26, x2, lsl #2]
add x2, x2, #1
cmp x25, x2
b.ne
removing the narrowing, the code will be optimized to:
ldr w0, [x4, x2, lsl #2] ; a[i]
lsl x0, x0, x24 ; a[i] << d
orr x0, x0, x3 ; carry chain
and w1, w0, #0x3fffffff
str w1, [x26, x2, lsl #2]
add x2, x2, #1
lsr x3, x0, #30 ; carry chain
cmp x25, x2
b.ne
The instruction count on the loop carried chain is reduced from 3 to 2.
Has this already been discussed elsewhere?
No response given
Links to previous discussion of this feature:
No response
Linked PRs
- gh-157060
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
从 longobject.c 开始,检查 v_lshift、v_rshift 和 x_divrem 中的进位处理。比较保留两位数进位值前后生成的 AArch64 循环,然后验证对于有效 limb 结果保持不变,并确认进位链中的指令数量如描述的那样得到改善。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- c
- 领域
- performance
- Issue 类型
- 功能
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 停滞
- 描述清晰度
- 描述清楚
- 新手友好度
- 45/100