Error messages for `{}` style formatters for int, float, str, and complex
Dieses Issue hat noch niemand übernommen.
- Vorherrschende Sprache
- Python
- Sterne
- 77.2k
- Forks
- 35.9k
- PR-Merge-Kennzahlen
- PR-Kennzahlen ausstehend
Beschreibung
Bug report
Bug description:
The {}-style formatters for int, float, str, and complex have several Unicode handling problems and confusion when generating ValueError messages.
There is an issue #142037 (and PR #142801) that focuses on %-style formatting, but the {}-style formatter has similar problems.
I am drafting a PR to fix the problems in the {} style formatter, and I would like to gather feedback on the proposed changes.
Invalid format specifier string
>>> f"{1:\n\\}"
Traceback (most recent call last):
File "<python-input-0>", line 1, in <module>
f"{1:\n\\}"
^^^^^^^^
ValueError: Invalid format specifier '
\' for object of type 'int'
In the above example, the error message is concatenated (%U) from the specifier string.
Suggestion: Use repr() (%R) to escape special characters.
Invalid type code in format specifier
>>> f"{1:\x7f}"
Traceback (most recent call last):
File "<python-input-2>", line 1, in <module>
f"{1:\x7f}"
^^^^^^^^
ValueError: Unknown format code '' for object of type 'int'
>>> f"{1:\t}"
Traceback (most recent call last):
File "<python-input-3>", line 1, in <module>
f"{1:\t}"
^^^^^^
ValueError: Unknown format code '\x9' for object of type 'int'
>>> f"{1:🐍}"
Traceback (most recent call last):
File "<python-input-4>", line 1, in <module>
f"{1:🐍}"
^^^^^^
ValueError: Unknown format code '\x1f40d' for object of type 'int'
The unprintable \x7f is concatenated, and \t and 🐍 are escaped improperly.
Suggestion: Use similar logic as in the %-style formatter in recent PR #142801:
Conversion specifier
Same behavior as above:
>>> "{0!🐍}".format(1)
Traceback (most recent call last):
File "<python-input-5>", line 1, in <module>
"{0!🐍}".format(1)
~~~~~~~~~~~~~~~^^^
ValueError: Unknown conversion specifier \x1f40d
Suggestion: Apply the same logic as above, and so it is also consistent with the compile-time SyntaxError:
>>> f"{1!🐍}"
File "<python-input-6>", line 1
f"{1!🐍}"
^^
SyntaxError: invalid character '🐍' (U+1F40D)
Check for fractional part grouping separator
Currently, the check for thousands separator occurs when parsing the format specifier string, but may be omitted for the fractional part, so it passes to format a str (The parser processes the string without knowledge of the object type).
The 1st and 3rd are expected, the 2nd may be not:
>>> f'{123456.123456:.,}'
'123456.123,456'
>>> f'{"x":.,s}'
'x'
>>> f'{"x":,s}'
Traceback (most recent call last):
File "<python-input-8>", line 1, in <module>
f'{"x":,s}'
^^^^^^^^
ValueError: Cannot specify ',' with 's'.
Suggestion: I didn't find any PEP or document about where we can use the fractional part grouping separator, so I suggest that it can (and only can) be used for e, f, g, E, G, %, and F. For comparison, integer part supports above plus d, b, o, x, and X. None of these five has a fractional part.
Related reference:
- Thousands separator was added in PEP 378.
- Fractional part grouping separator was added in issue #87790.
Check type field before thousands separator
>>> f'{100000:+#020,🐍}'
Traceback (most recent call last):
File "<python-input-10>", line 1, in <module>
f'{100000:+#020,🐍}'
^^^^^^^^^^^^^^^^^
ValueError: Cannot specify ',' with '\x1f40d'.
This should be a type field issue that reports unknown format code, instead of misusing the thousands separator. Also, the character is not escaped properly.
Suggestion: Check the type field is valid first, then check the thousands separator. So we can also avoid repeating the Unicode codepoint logic when forming the error message.
I would appreciate any feedback or suggestions on these proposed changes. Thank you!
CPython versions tested on:
CPython main branch
Operating systems tested on:
macOS
Linked PRs
- gh-144326
Beitragsleitfaden
Erste Schritte
- Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
- Forke das Repository und arbeite in einem Branch.
- Öffne einen Pull Request, der die Issue-Nummer nennt.
Rechercherichtung
Beginnen Sie mit Objects/unicode_format.c, insbesondere mit der Formatierungslogik, auf die in PR #142801 verwiesen wird, und vergleichen Sie sie mit dem verlinkten PR #144326. Reproduzieren Sie die aufgeführten ungültigen Formatbezeichner- und Konvertierungsfälle und überprüfen Sie anschließend, dass Unicode konsistent maskiert wird und die Gruppierung von Nachkommastellen sowie die Validierung des Typfelds den vorgeschlagenen Regeln folgen.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python
- Bereich
- compilers
- Issue-Typ
- Bug
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Veraltet
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 25/100