Error messages for `{}` style formatters for int, float, str, and complex
Ninguém assumiu esta issue ainda.
- Linguagem predominante
- Python
- Estrelas
- 77.2k
- Forks
- 36k
- Métricas de merge de PRs
- Métricas de PR pendentes
Descrição
Bug report
Bug description:
The {}-style formatters for int, float, str, and complex have several Unicode handling problems and confusion when generating ValueError messages.
There is an issue #142037 (and PR #142801) that focuses on %-style formatting, but the {}-style formatter has similar problems.
I am drafting a PR to fix the problems in the {} style formatter, and I would like to gather feedback on the proposed changes.
Invalid format specifier string
>>> f"{1:\n\\}"
Traceback (most recent call last):
File "<python-input-0>", line 1, in <module>
f"{1:\n\\}"
^^^^^^^^
ValueError: Invalid format specifier '
\' for object of type 'int'
In the above example, the error message is concatenated (%U) from the specifier string.
Suggestion: Use repr() (%R) to escape special characters.
Invalid type code in format specifier
>>> f"{1:\x7f}"
Traceback (most recent call last):
File "<python-input-2>", line 1, in <module>
f"{1:\x7f}"
^^^^^^^^
ValueError: Unknown format code '' for object of type 'int'
>>> f"{1:\t}"
Traceback (most recent call last):
File "<python-input-3>", line 1, in <module>
f"{1:\t}"
^^^^^^
ValueError: Unknown format code '\x9' for object of type 'int'
>>> f"{1:🐍}"
Traceback (most recent call last):
File "<python-input-4>", line 1, in <module>
f"{1:🐍}"
^^^^^^
ValueError: Unknown format code '\x1f40d' for object of type 'int'
The unprintable \x7f is concatenated, and \t and 🐍 are escaped improperly.
Suggestion: Use similar logic as in the %-style formatter in recent PR #142801:
Conversion specifier
Same behavior as above:
>>> "{0!🐍}".format(1)
Traceback (most recent call last):
File "<python-input-5>", line 1, in <module>
"{0!🐍}".format(1)
~~~~~~~~~~~~~~~^^^
ValueError: Unknown conversion specifier \x1f40d
Suggestion: Apply the same logic as above, and so it is also consistent with the compile-time SyntaxError:
>>> f"{1!🐍}"
File "<python-input-6>", line 1
f"{1!🐍}"
^^
SyntaxError: invalid character '🐍' (U+1F40D)
Check for fractional part grouping separator
Currently, the check for thousands separator occurs when parsing the format specifier string, but may be omitted for the fractional part, so it passes to format a str (The parser processes the string without knowledge of the object type).
The 1st and 3rd are expected, the 2nd may be not:
>>> f'{123456.123456:.,}'
'123456.123,456'
>>> f'{"x":.,s}'
'x'
>>> f'{"x":,s}'
Traceback (most recent call last):
File "<python-input-8>", line 1, in <module>
f'{"x":,s}'
^^^^^^^^
ValueError: Cannot specify ',' with 's'.
Suggestion: I didn't find any PEP or document about where we can use the fractional part grouping separator, so I suggest that it can (and only can) be used for e, f, g, E, G, %, and F. For comparison, integer part supports above plus d, b, o, x, and X. None of these five has a fractional part.
Related reference:
- Thousands separator was added in PEP 378.
- Fractional part grouping separator was added in issue #87790.
Check type field before thousands separator
>>> f'{100000:+#020,🐍}'
Traceback (most recent call last):
File "<python-input-10>", line 1, in <module>
f'{100000:+#020,🐍}'
^^^^^^^^^^^^^^^^^
ValueError: Cannot specify ',' with '\x1f40d'.
This should be a type field issue that reports unknown format code, instead of misusing the thousands separator. Also, the character is not escaped properly.
Suggestion: Check the type field is valid first, then check the thousands separator. So we can also avoid repeating the Unicode codepoint logic when forming the error message.
I would appreciate any feedback or suggestions on these proposed changes. Thank you!
CPython versions tested on:
CPython main branch
Operating systems tested on:
macOS
Linked PRs
- gh-144326
Guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Direção de pesquisa
Comece por Objects/unicode_format.c, particularmente pela lógica de formatação referenciada no PR #142801, e compare-a com o PR #144326 vinculado. Reproduza os casos listados de especificadores de formato e conversões inválidos e, em seguida, verifique se Unicode é escapado de forma consistente e se o agrupamento fracionário e a validação do campo de tipo seguem as regras propostas.
Escrita pelo modelo de indexação a partir do texto da issue.
Avaliação
- Stack de tecnologia
- python
- Domínio
- compilers
- Tipo de issue
- Bug
- Dificuldade
- 4/5
- Tempo estimado
- 3-5 dias
- Status de atividade
- Estagnada
- Clareza
- Razoavelmente clara
- Facilidade para iniciantes
- 25/100