Error messages for `{}` style formatters for int, float, str, and complex
Chưa có ai nhận issue này.
- Ngôn ngữ chính
- Python
- Star
- 77.2k
- Fork
- 35.9k
- Chỉ số merge pull request
- Chỉ số pull request đang chờ
Mô tả
Bug report
Bug description:
The {}-style formatters for int, float, str, and complex have several Unicode handling problems and confusion when generating ValueError messages.
There is an issue #142037 (and PR #142801) that focuses on %-style formatting, but the {}-style formatter has similar problems.
I am drafting a PR to fix the problems in the {} style formatter, and I would like to gather feedback on the proposed changes.
Invalid format specifier string
>>> f"{1:\n\\}"
Traceback (most recent call last):
File "<python-input-0>", line 1, in <module>
f"{1:\n\\}"
^^^^^^^^
ValueError: Invalid format specifier '
\' for object of type 'int'
In the above example, the error message is concatenated (%U) from the specifier string.
Suggestion: Use repr() (%R) to escape special characters.
Invalid type code in format specifier
>>> f"{1:\x7f}"
Traceback (most recent call last):
File "<python-input-2>", line 1, in <module>
f"{1:\x7f}"
^^^^^^^^
ValueError: Unknown format code '' for object of type 'int'
>>> f"{1:\t}"
Traceback (most recent call last):
File "<python-input-3>", line 1, in <module>
f"{1:\t}"
^^^^^^
ValueError: Unknown format code '\x9' for object of type 'int'
>>> f"{1:🐍}"
Traceback (most recent call last):
File "<python-input-4>", line 1, in <module>
f"{1:🐍}"
^^^^^^
ValueError: Unknown format code '\x1f40d' for object of type 'int'
The unprintable \x7f is concatenated, and \t and 🐍 are escaped improperly.
Suggestion: Use similar logic as in the %-style formatter in recent PR #142801:
Conversion specifier
Same behavior as above:
>>> "{0!🐍}".format(1)
Traceback (most recent call last):
File "<python-input-5>", line 1, in <module>
"{0!🐍}".format(1)
~~~~~~~~~~~~~~~^^^
ValueError: Unknown conversion specifier \x1f40d
Suggestion: Apply the same logic as above, and so it is also consistent with the compile-time SyntaxError:
>>> f"{1!🐍}"
File "<python-input-6>", line 1
f"{1!🐍}"
^^
SyntaxError: invalid character '🐍' (U+1F40D)
Check for fractional part grouping separator
Currently, the check for thousands separator occurs when parsing the format specifier string, but may be omitted for the fractional part, so it passes to format a str (The parser processes the string without knowledge of the object type).
The 1st and 3rd are expected, the 2nd may be not:
>>> f'{123456.123456:.,}'
'123456.123,456'
>>> f'{"x":.,s}'
'x'
>>> f'{"x":,s}'
Traceback (most recent call last):
File "<python-input-8>", line 1, in <module>
f'{"x":,s}'
^^^^^^^^
ValueError: Cannot specify ',' with 's'.
Suggestion: I didn't find any PEP or document about where we can use the fractional part grouping separator, so I suggest that it can (and only can) be used for e, f, g, E, G, %, and F. For comparison, integer part supports above plus d, b, o, x, and X. None of these five has a fractional part.
Related reference:
- Thousands separator was added in PEP 378.
- Fractional part grouping separator was added in issue #87790.
Check type field before thousands separator
>>> f'{100000:+#020,🐍}'
Traceback (most recent call last):
File "<python-input-10>", line 1, in <module>
f'{100000:+#020,🐍}'
^^^^^^^^^^^^^^^^^
ValueError: Cannot specify ',' with '\x1f40d'.
This should be a type field issue that reports unknown format code, instead of misusing the thousands separator. Also, the character is not escaped properly.
Suggestion: Check the type field is valid first, then check the thousands separator. So we can also avoid repeating the Unicode codepoint logic when forming the error message.
I would appreciate any feedback or suggestions on these proposed changes. Thank you!
CPython versions tested on:
CPython main branch
Operating systems tested on:
macOS
Linked PRs
- gh-144326
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Hướng nghiên cứu
Bắt đầu với Objects/unicode_format.c, đặc biệt là logic định dạng được tham chiếu từ PR #142801, và so sánh với PR #144326 được liên kết. Tái hiện các trường hợp bộ chỉ định định dạng và chuyển đổi không hợp lệ được liệt kê, sau đó xác minh rằng Unicode được escape nhất quán, đồng thời việc nhóm phần phân số và kiểm tra hợp lệ trường kiểu tuân theo các quy tắc được đề xuất.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- python
- Lĩnh vực
- compilers
- Loại issue
- Lỗi
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức độ hoạt động
- Đình trệ
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 25/100