python / python/cpython

http.client accepts Content-Length and chunk-size values that RFC 9112 forbids

未關閉
#150,751 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

stdlib type-bug
主要語言
Python
星號
77.2k
分支
35.9k
PR 合併指標
PR 指標待擷取

描述

http.client derives the response body framing from int() of the Content-Length header and the chunked chunk-size line:

  • HTTPResponse.begin: self.length = int(length)
  • HTTPResponse._read_next_chunk_size: return int(line, 16)

RFC 9112 defines Content-Length = 1*DIGIT and chunk-size = 1*HEXDIG, but int() is more permissive: it accepts a leading +/-, underscores, surrounding whitespace and, in base 16, an 0x prefix and non-ASCII digits. So values like Content-Length: +5 / 5_0 and chunk sizes -5, +5, 0x5, 1_f are accepted and used to frame the body, while an RFC-compliant front end would reject them or frame the message differently (CWE-444).

Reproducer:

import http.client, io
class S:
    def __init__(s, d): s.f = io.BytesIO(d)
    def makefile(s, *a, **k): return s.f
def parse(raw):
    r = http.client.HTTPResponse(S(raw)); r.begin(); return r
raw = b'HTTP/1.1 200 OK\r\nTransfer-Encoding: chunked\r\n\r\n+5\r\nHELLO\r\n0\r\n\r\n'
print(parse(raw).read())            # b'HELLO' -- '+5' is not a HEXDIG
raw = b'HTTP/1.1 200 OK\r\nContent-Length: 5_0\r\n\r\n' + b'A'*50
print(parse(raw).length)            # 50 -- '5_0' is not 1*DIGIT

The body-framing tokens should be validated against the grammar before being passed to int().

Linked PRs
  • gh-150752

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

從 HTTPResponse.begin 和 HTTPResponse._read_next_chunk_size 開始,這裡會使用 int() 轉換 Content-Length 和 chunk-size 值。在轉換前,根據 RFC 9112 的語法驗證每個 body framing token,並確認 reproducer 中不符合規範的值會被拒絕,或不再用於 framing。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
networking, security
Issue 類型
缺陷
難度
3/5
預估耗時
1-2 天
活躍度
停滯
描述清晰度
描述清楚
新手友好度
35/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。