email.parser.BytesParser.parse() cannot handle binary data that include \x0d \x0a correctly.
Chưa có ai nhận issue này.
- Ngôn ngữ chính
- Python
- Star
- 77.2k
- Fork
- 35.9k
- Chỉ số merge pull request
- Chỉ số pull request đang chờ
Mô tả
Bug report
Bug description:
I would like to extract a binary file in a multipart MIME file using email.parser.BytesParser, but a byte sequence "0x0d 0x0a" (CR + LF) in the binary file is replaced by "0x0a" (LF). Below is a minimal reproducible example.
from email.parser import BytesParser
from email.policy import default
from io import BytesIO
mime_file_byte_array = b'MIME-Version: 1.0\r\nContent-Type: multipart/mixed; boundary="MIME\
_boundary-1";\r\n\r\n--MIME_boundary-1\r\nContent-Type: application/octet-stream\r\nContent\
-Location: test.bin\r\n\r\na\r\nb\r\n--MIME_boundary-1--\r\n\r\n'
fp = BytesIO(mime_file_byte_array)
parser = BytesParser(policy=default)
msg = parser.parse(fp)
parts = [part for part in msg.walk()]
binary_data = parts[1].get_payload(decode=True)
print('===== Beginning of Original MIME File =====')
print(mime_file_byte_array.decode())
print('===== End of Original MIME File =====')
print('')
print('===== test.bin after parse =====')
print(binary_data)
print('===== test.bin after parse =====')
As can be seen in the fifth line, the multipart MIME file includes a binary file "test.bin". The contents of the binary file is b"a\r\nb".
Therefore, the variable binary_data is supposed to contain b"a\r\nb", but it was actually b"a\nb".
It is probably because TextIOWrapper in BytesParser.parse() translates CR+LF to LF on Linux.
https://github.com/python/cpython/blob/767c89ba7c5a70626df6e75eb56b546bf911b997/Lib/email/parser.py#L103
When I replaced the above line with the line below, this problem was fixed. However, this fix may have a side effect which I cannot foresee.
fp = TextIOWrapper(fp, encoding='ascii', errors='surrogateescape', newline='')
CPython versions tested on:
3.10
Operating systems tested on:
Linux
Linked PRs
- gh-157726
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Hướng nghiên cứu
Bắt đầu trong Lib/email/parser.py tại dòng TextIOWrapper được đề cập trong báo cáo, sau đó tái hiện vấn đề bằng ví dụ BytesIO được cung cấp. So sánh payload đã được phân tích với dữ liệu b"a\r\nb" ban đầu và xem xét PR liên quan gh-157726; công việc hoàn tất khi BytesParser giữ nguyên chuỗi CRLF nhị phân.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- python
- Lĩnh vực
- backend
- Loại issue
- Lỗi
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức độ hoạt động
- Đình trệ
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 35/100