Error messages for `{}` style formatters for int, float, str, and complex
Personne n'a encore pris cette issue.
- Langage dominant
- Python
- Étoiles
- 77.2k
- Forks
- 35.9k
- Métriques de merge des PR
- Métriques de PR en attente
Description
Bug report
Bug description:
The {}-style formatters for int, float, str, and complex have several Unicode handling problems and confusion when generating ValueError messages.
There is an issue #142037 (and PR #142801) that focuses on %-style formatting, but the {}-style formatter has similar problems.
I am drafting a PR to fix the problems in the {} style formatter, and I would like to gather feedback on the proposed changes.
Invalid format specifier string
>>> f"{1:\n\\}"
Traceback (most recent call last):
File "<python-input-0>", line 1, in <module>
f"{1:\n\\}"
^^^^^^^^
ValueError: Invalid format specifier '
\' for object of type 'int'
In the above example, the error message is concatenated (%U) from the specifier string.
Suggestion: Use repr() (%R) to escape special characters.
Invalid type code in format specifier
>>> f"{1:\x7f}"
Traceback (most recent call last):
File "<python-input-2>", line 1, in <module>
f"{1:\x7f}"
^^^^^^^^
ValueError: Unknown format code '' for object of type 'int'
>>> f"{1:\t}"
Traceback (most recent call last):
File "<python-input-3>", line 1, in <module>
f"{1:\t}"
^^^^^^
ValueError: Unknown format code '\x9' for object of type 'int'
>>> f"{1:🐍}"
Traceback (most recent call last):
File "<python-input-4>", line 1, in <module>
f"{1:🐍}"
^^^^^^
ValueError: Unknown format code '\x1f40d' for object of type 'int'
The unprintable \x7f is concatenated, and \t and 🐍 are escaped improperly.
Suggestion: Use similar logic as in the %-style formatter in recent PR #142801:
Conversion specifier
Same behavior as above:
>>> "{0!🐍}".format(1)
Traceback (most recent call last):
File "<python-input-5>", line 1, in <module>
"{0!🐍}".format(1)
~~~~~~~~~~~~~~~^^^
ValueError: Unknown conversion specifier \x1f40d
Suggestion: Apply the same logic as above, and so it is also consistent with the compile-time SyntaxError:
>>> f"{1!🐍}"
File "<python-input-6>", line 1
f"{1!🐍}"
^^
SyntaxError: invalid character '🐍' (U+1F40D)
Check for fractional part grouping separator
Currently, the check for thousands separator occurs when parsing the format specifier string, but may be omitted for the fractional part, so it passes to format a str (The parser processes the string without knowledge of the object type).
The 1st and 3rd are expected, the 2nd may be not:
>>> f'{123456.123456:.,}'
'123456.123,456'
>>> f'{"x":.,s}'
'x'
>>> f'{"x":,s}'
Traceback (most recent call last):
File "<python-input-8>", line 1, in <module>
f'{"x":,s}'
^^^^^^^^
ValueError: Cannot specify ',' with 's'.
Suggestion: I didn't find any PEP or document about where we can use the fractional part grouping separator, so I suggest that it can (and only can) be used for e, f, g, E, G, %, and F. For comparison, integer part supports above plus d, b, o, x, and X. None of these five has a fractional part.
Related reference:
- Thousands separator was added in PEP 378.
- Fractional part grouping separator was added in issue #87790.
Check type field before thousands separator
>>> f'{100000:+#020,🐍}'
Traceback (most recent call last):
File "<python-input-10>", line 1, in <module>
f'{100000:+#020,🐍}'
^^^^^^^^^^^^^^^^^
ValueError: Cannot specify ',' with '\x1f40d'.
This should be a type field issue that reports unknown format code, instead of misusing the thousands separator. Also, the character is not escaped properly.
Suggestion: Check the type field is valid first, then check the thousands separator. So we can also avoid repeating the Unicode codepoint logic when forming the error message.
I would appreciate any feedback or suggestions on these proposed changes. Thank you!
CPython versions tested on:
CPython main branch
Operating systems tested on:
macOS
Linked PRs
- gh-144326
Guide de contribution
Ouvrir le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Piste de recherche
Commencez par Objects/unicode_format.c, en particulier par la logique de formatage référencée depuis PR #142801, et comparez-la avec le PR #144326 indiqué. Reproduisez les cas listés de spécificateurs de format et de conversions invalides, puis vérifiez que Unicode est échappé de manière cohérente et que le groupement fractionnaire ainsi que la validation du champ de type suivent les règles proposées.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- python
- Domaine
- compilers
- Type d'issue
- Bug
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Activité
- À l'abandon
- Clarté
- Plutôt claire
- Accessibilité débutants
- 25/100