Error messages for `{}` style formatters for int, float, str, and complex
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 77.2k
- Forks
- 35.9k
- PR merge metrics
- PR metrics pending
Description
Bug report
Bug description:
The {}-style formatters for int, float, str, and complex have several Unicode handling problems and confusion when generating ValueError messages.
There is an issue #142037 (and PR #142801) that focuses on %-style formatting, but the {}-style formatter has similar problems.
I am drafting a PR to fix the problems in the {} style formatter, and I would like to gather feedback on the proposed changes.
Invalid format specifier string
>>> f"{1:\n\\}"
Traceback (most recent call last):
File "<python-input-0>", line 1, in <module>
f"{1:\n\\}"
^^^^^^^^
ValueError: Invalid format specifier '
\' for object of type 'int'
In the above example, the error message is concatenated (%U) from the specifier string.
Suggestion: Use repr() (%R) to escape special characters.
Invalid type code in format specifier
>>> f"{1:\x7f}"
Traceback (most recent call last):
File "<python-input-2>", line 1, in <module>
f"{1:\x7f}"
^^^^^^^^
ValueError: Unknown format code '' for object of type 'int'
>>> f"{1:\t}"
Traceback (most recent call last):
File "<python-input-3>", line 1, in <module>
f"{1:\t}"
^^^^^^
ValueError: Unknown format code '\x9' for object of type 'int'
>>> f"{1:🐍}"
Traceback (most recent call last):
File "<python-input-4>", line 1, in <module>
f"{1:🐍}"
^^^^^^
ValueError: Unknown format code '\x1f40d' for object of type 'int'
The unprintable \x7f is concatenated, and \t and 🐍 are escaped improperly.
Suggestion: Use similar logic as in the %-style formatter in recent PR #142801:
Conversion specifier
Same behavior as above:
>>> "{0!🐍}".format(1)
Traceback (most recent call last):
File "<python-input-5>", line 1, in <module>
"{0!🐍}".format(1)
~~~~~~~~~~~~~~~^^^
ValueError: Unknown conversion specifier \x1f40d
Suggestion: Apply the same logic as above, and so it is also consistent with the compile-time SyntaxError:
>>> f"{1!🐍}"
File "<python-input-6>", line 1
f"{1!🐍}"
^^
SyntaxError: invalid character '🐍' (U+1F40D)
Check for fractional part grouping separator
Currently, the check for thousands separator occurs when parsing the format specifier string, but may be omitted for the fractional part, so it passes to format a str (The parser processes the string without knowledge of the object type).
The 1st and 3rd are expected, the 2nd may be not:
>>> f'{123456.123456:.,}'
'123456.123,456'
>>> f'{"x":.,s}'
'x'
>>> f'{"x":,s}'
Traceback (most recent call last):
File "<python-input-8>", line 1, in <module>
f'{"x":,s}'
^^^^^^^^
ValueError: Cannot specify ',' with 's'.
Suggestion: I didn't find any PEP or document about where we can use the fractional part grouping separator, so I suggest that it can (and only can) be used for e, f, g, E, G, %, and F. For comparison, integer part supports above plus d, b, o, x, and X. None of these five has a fractional part.
Related reference:
- Thousands separator was added in PEP 378.
- Fractional part grouping separator was added in issue #87790.
Check type field before thousands separator
>>> f'{100000:+#020,🐍}'
Traceback (most recent call last):
File "<python-input-10>", line 1, in <module>
f'{100000:+#020,🐍}'
^^^^^^^^^^^^^^^^^
ValueError: Cannot specify ',' with '\x1f40d'.
This should be a type field issue that reports unknown format code, instead of misusing the thousands separator. Also, the character is not escaped properly.
Suggestion: Check the type field is valid first, then check the thousands separator. So we can also avoid repeating the Unicode codepoint logic when forming the error message.
I would appreciate any feedback or suggestions on these proposed changes. Thank you!
CPython versions tested on:
CPython main branch
Operating systems tested on:
macOS
Linked PRs
- gh-144326
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with Objects/unicode_format.c, particularly the formatting logic referenced from PR #142801, and compare the linked PR #144326. Reproduce the listed invalid format-specifier and conversion cases, then verify that Unicode is escaped consistently and fractional grouping and type-field validation follow the proposed rules.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- compilers
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100