python / python/cpython

Py_UNICODE_TOUPPER() and friends do not return the simple case mappings

Open
#156,518 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

extension-modules topic-C-API topic-unicode type-bug
Dominant language
Python
Stars
77.2k
Forks
35.9k
PR merge metrics
PR metrics pending

Description

Bug report

Py_UNICODE_TOUPPER(), Py_UNICODE_TOLOWER() and Py_UNICODE_TOTITLE() return the first character of the full case mapping when it is longer than one character.

UnicodeData.txt leaves the simple uppercase of 'ß' (LATIN SMALL LETTER SHARP S) empty, which means the character is its own simple uppercase; the full uppercase "SS" is in SpecialCasing.txt. Py_UNICODE_TOUPPER() returns 'S'. The same for 'ǰ' -> 'J', 'fi' -> 'F' and 'ΐ' -> 'Ι'.

For 27 Greek letters with ypogegrammeni the UCD does define a simple uppercase, and it is not the first character of the full one:

1F80;GREEK SMALL LETTER ALPHA WITH PSILI AND YPOGEGRAMMENI;Ll;0;L;1F00 0345;;;;N;;;1F88;;1F88
1F80; 1F80; 1F88; 1F08 0399; # GREEK SMALL LETTER ALPHA WITH PSILI AND YPOGEGRAMMENI

The simple uppercase of 'ᾀ' is 'ᾈ' (U+1F88), the full one is + Ι, and the macro returns (U+1F08). Tools/unicode/makeunicodedata.py stores only the full mappings for a character which has a SpecialCasing.txt entry, so the simple ones are not in the table at all.

102 characters are affected for the uppercase, 48 for the titlecase, and one (LATIN CAPITAL LETTER I WITH DOT ABOVE, whose simple lowercase is 'i') for the lowercase. All of them are Unicode 1.1, and no character added since has joined them.

This is a side effect of b2bf01d824e (bpo-12736, 3.3), which added the full mappings: before it the fields held the simple mappings.

_sre is the only user in CPython, and it compensates: _sre.unicode_iscased() reports whether the simple lowercase or uppercase differs from the character, which is right for 'ß' only because the returned 'S' differs from it. Fixing the mappings requires comparing the full mappings there instead.

Linked PRs
  • gh-156519

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with Tools/unicode/makeunicodedata.py, which currently stores full mappings for SpecialCasing.txt entries, then inspect the _sre unicode_iscased() compensation. Done means simple case mappings are preserved for the affected characters and _sre still determines casing correctly using the full mappings.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
internationalization
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.