python / python/cpython

Make `re.Match` a well-rounded `Sequence` type

Ouverte
#133,546 5 commentaires 1 réaction 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

extension-modules topic-regex type-feature
Langage dominant
Python
Étoiles
77.2k
Forks
35.9k
Métriques de merge des PR
Métriques de PR en attente

Description

Feature or enhancement

Proposal:

It would be nice if the following worked as expected:

m = re.match(r"(a)(b)(c)", "abc")

assert isinstance(m, Sequence)
assert len(m) == 4
assert list(m) == ["abc", "a", "b", "c"]

abc, a, b, c = m
assert abc == "abc" and a == "a" and b == "b" and c == "c"

match re.match(r"(\d+)-(\d+)-(\d+)", "2025-05-07"):
    case [_, year, month, day]:
        assert year == "2025" and month == "05" and day == "07"

If you also work with Javascript this will feel very familiar:

let m = "abc".match(/(a)(b)(c)/)

console.log(m instanceof Array) // true
console.log(m.length) // 4
console.log(Array.from(m)) // [ 'abc', 'a', 'b', 'c' ]

let [abc, a, b, c] = m
console.log(abc) // abc
console.log(a) // a
console.log(b) // b
console.log(c) // c

Back in 2016, the re.Match object API was expanded to include __getitem__ as a shortcut for .group(...).

The goal was to improve usability and approachability by making re.Match objects fit a bit more seamlessly into python's core data model. Accessing groups via subscripting is now intuitive, but because re.Match objects only have a __getitem__ and no __len__, they can't be used as a proper Sequence type.

To me, this always felt a bit awkward. After digging up the original discussion, it seems like the reason why __len__ didn't make it was that it was still undecided whether the returned value should take into account group 0 or not.

Almost a decade later, as a user, the way I see it is that the __getitem__ implementation we're now used to suggests a regular Sequence type that also happens to transparently translate group names provided as subscript to their corresponding group index. In fact, this is actually how it works in the underlying C code.

With this in mind, we can simply define __len__ taking into account group 0, and we'll finally be able to enjoy coherent re.Match objects that behave as proper Sequence types.

Has this already been discussed elsewhere?

https://discuss.python.org/t/make-re-match-a-well-rounded-sequence-type/91039

Links to previous discussion of this feature:

Improve the usability of the match object named group API (https://github.com/python/cpython/commit/605bdae0780a1ff8a88b80e6e67f3f355fb2ddb4)

Linked PRs
  • gh-133549

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Piste de recherche

Commencez par le point d’entrée re.Match et son comportement getitem existant, puis lisez les discussions liées sur group 0 et la sémantique de Sequence. Utilisez les assertions de la proposition et l’exemple de pattern matching comme critères d’acceptation ; terminé signifie que le comportement attendu pour la longueur, l’itération, le déballage et les patterns de séquence est établi, sans questions d’API non résolues.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
backend
Type d'issue
Fonctionnalité
Difficulté
5/5
Temps estimé
Plus d'une semaine
Activité
À l'abandon
Clarté
Plutôt claire
Accessibilité débutants
25/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.