python / python/pyperformance

Benchmarks for common python I/O patterns

未关闭
#399 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

主要语言
Python
星标
1k
派生
203
平均合并
1 小时 20 分钟
30 天内合并 PR
2

描述

As work happens on I/O pieces I've been building specialized micro-benchmarks (ex. gh-120754 Speed up open().read() pattern by reducing the number of system calls and others have gh-117151: IO performance improvement, increase io.DEFAULT_BUFFER_SIZE to 128k), it would be nice to have more general benchmarks to validate I/O performance for common cases.

Talking a little with people at PyConUS there was some interest in the tests, and a general desire for I/O tests not to be enabled by default, but to be a group which can be manually run.

General I/O shapes I'm hoping to add benchmarks for:

  • read/write all of the byes of a file in a single call (including pathlib.Path.read_text, pathlib.Path.write_text)
  • read/write many small files (ex. .pyc files, maybe just compile_all?)
  • streaming bytes read/write (ex. to a pipe / console such as stdin/stdout/stderr, non-seekable devices)
  • read/write a zipfile, tarfile (read + seek, write + seek, in particular buffering behavior)
  • use zipimport
  • Create a zipapp
  • Multi-threaded write to stdout, stderr (ex. logging in a large application/codebase)

Note: With these aiming to stay at the Binary / Bytes IO layer as much as possible (not touch Text I/O for now)

贡献指南

这个仓库没有索引到贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

审查现有的 pyperformance 基准测试组织方式以及本 issue 中提出的 I/O 形式,将范围限定在 Binary/Bytes IO 层。定义一个可手动运行的组,覆盖选定的常见用例,然后验证它默认未启用,并且这些基准测试会执行预期的文件、流、归档、导入或并发模式。

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
performance
Issue 类型
功能
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。