python / python/cpython

Align the grammar documentation with Python's actual grammar

Offen
#127,833 2 Kommentare 3 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

docs
Vorherrschende Sprache
Python
Sterne
77.2k
Forks
35.9k
PR-Merge-Kennzahlen
PR-Kennzahlen ausstehend

Beschreibung

Documentation

The current documentation of Python syntax (the later chapters of the language reference) uses hand-maintained production lists, like this:

A)

compound_stmt ::=  if_stmt
                   | while_stmt
                   | for_stmt
                   | try_stmt
                   | with_stmt
                   | match_stmt
                   | funcdef
                   | classdef
                   | async_with_stmt
                   | async_for_stmt
                   | async_funcdef
suite         ::=  stmt_list NEWLINE | NEWLINE INDENT statement+ DEDENT
statement     ::=  stmt_list NEWLINE | compound_stmt
stmt_list     ::=  simple_stmt (";" simple_stmt)* [";"]

There is no mechanism to ensure that these are in sync with the actual grammar, and they inevitably do get out of sync.
See some of the “docs” issues mentioning “grammar”.

It's not easy to write an automatic tool to keep them in sync, because we do want to elide some details -- the parser rules, unnecessary lookaheads, cuts, etc. But, it's possible to write it, and we wrote a proof of concept, which will need to be rewritten, tuned, and reviewed. Before introducing it, I'd like to go through all the docs, correct the existing documentation, bring it closer to what a tool could generate, and discuss what the ideal presentation would look like. That needs to be a manual process, and it will also need to touch the prose that's next to the grammar snippets.

As a first step, I propose an update to the tooling, which brings the presentation a bit closer to the python.gram syntax.

From the existing ReST source, we can get this:

B)

compound_stmt: if_stmt
               | while_stmt
               | for_stmt
               | try_stmt
               | with_stmt
               | match_stmt
               | funcdef
               | classdef
               | async_with_stmt
               | async_for_stmt
               | async_funcdef
suite:         stmt_list NEWLINE | NEWLINE INDENT statement+ DEDENT
statement:     stmt_list NEWLINE | compound_stmt
stmt_list:     simple_stmt (";" simple_stmt)* [";"]

Since Sphinx hard-codes the productionlist formatting (the ::= symbol and the aligning), we'll need to override the productionlist directive to achieve this.

Then, by changing the ReST and using a different directive, we can get to something like:

C)

compound_stmt:
    | if_stmt
    | while_stmt
    | for_stmt
    | try_stmt
    | with_stmt
    | match_stmt
    | funcdef
    | classdef
    | async_with_stmt
    | async_for_stmt
    | async_funcdef
suite:
    | stmt_list NEWLINE | NEWLINE INDENT statement+ DEDENT
statement:
    | stmt_list NEWLINE | compound_stmt
stmt_list:
    | simple_stmt (";" simple_stmt)* [";"]

I propose to go from A) to B) at once (by overriding productionlist), and from B) to C) gradually, while also updating the content (including changing rule names to match the grammar, and adjusting/reorganizing nearby prose).
I think that the B) and C) styles are similar enough that mixing them in a single version of the docs should not be jarring.

By the way, one additional benefit of a custom directive is that we can add syntax highlighting. (Ideally, with support from the theme.) I think that making strings stand out makes the listings more readable:

image


As a second step, I'd like to rewrite token documentation, and then the lexical analysis chapter, on which the grammar chapters build: #135676

Then, continue with the grammar chapters in Language reference:

  • #141984
  • Simple statements
  • Compound statements
  • Top-level components
  • Full Grammar specification

And after that, use generated grammar snippets wherever possible.

Linked PRs
  • gh-127835
  • gh-129689
  • gh-129690
  • gh-129692
  • gh-130376
  • gh-131468
  • gh-131474
  • gh-132407
  • gh-133632
  • gh-133633
  • gh-133652
  • gh-134423
  • gh-134443
  • gh-134850
  • gh-135301
Other related PRs
  • gh-130588
  • gh-134034

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Lies Grammar/python.gram und Tools/peg_generator/docs_generator.py und prüfe anschließend die ReST-Produktionslisten der Sprachreferenz sowie die verknüpften PRs. Die Arbeit ist abgeschlossen, wenn die Grammatikdokumentation und der umgebende Text mit der tatsächlichen Grammatik übereinstimmen und die vorgeschlagene Darstellung der Produktionslisten sowie die Änderungen an den Tools geprüft und angewendet wurden.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
documentation
Issue-Typ
Dokumentation
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
25/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.