Align the grammar documentation with Python's actual grammar
まだ誰も着手していません。
- 主要言語
- Python
- スター
- 77.2k
- フォーク
- 36k
- PR マージ指標
- PR 指標を取得中
説明
Documentation
The current documentation of Python syntax (the later chapters of the language reference) uses hand-maintained production lists, like this:
A)
compound_stmt ::= if_stmt
| while_stmt
| for_stmt
| try_stmt
| with_stmt
| match_stmt
| funcdef
| classdef
| async_with_stmt
| async_for_stmt
| async_funcdef
suite ::= stmt_list NEWLINE | NEWLINE INDENT statement+ DEDENT
statement ::= stmt_list NEWLINE | compound_stmt
stmt_list ::= simple_stmt (";" simple_stmt)* [";"]
There is no mechanism to ensure that these are in sync with the actual grammar, and they inevitably do get out of sync.
See some of the “docs” issues mentioning “grammar”.
It's not easy to write an automatic tool to keep them in sync, because we do want to elide some details -- the parser rules, unnecessary lookaheads, cuts, etc. But, it's possible to write it, and we wrote a proof of concept, which will need to be rewritten, tuned, and reviewed. Before introducing it, I'd like to go through all the docs, correct the existing documentation, bring it closer to what a tool could generate, and discuss what the ideal presentation would look like. That needs to be a manual process, and it will also need to touch the prose that's next to the grammar snippets.
As a first step, I propose an update to the tooling, which brings the presentation a bit closer to the python.gram syntax.
From the existing ReST source, we can get this:
B)
compound_stmt: if_stmt
| while_stmt
| for_stmt
| try_stmt
| with_stmt
| match_stmt
| funcdef
| classdef
| async_with_stmt
| async_for_stmt
| async_funcdef
suite: stmt_list NEWLINE | NEWLINE INDENT statement+ DEDENT
statement: stmt_list NEWLINE | compound_stmt
stmt_list: simple_stmt (";" simple_stmt)* [";"]
Since Sphinx hard-codes the productionlist formatting (the ::= symbol and the aligning), we'll need to override the productionlist directive to achieve this.
Then, by changing the ReST and using a different directive, we can get to something like:
C)
compound_stmt:
| if_stmt
| while_stmt
| for_stmt
| try_stmt
| with_stmt
| match_stmt
| funcdef
| classdef
| async_with_stmt
| async_for_stmt
| async_funcdef
suite:
| stmt_list NEWLINE | NEWLINE INDENT statement+ DEDENT
statement:
| stmt_list NEWLINE | compound_stmt
stmt_list:
| simple_stmt (";" simple_stmt)* [";"]
I propose to go from A) to B) at once (by overriding productionlist), and from B) to C) gradually, while also updating the content (including changing rule names to match the grammar, and adjusting/reorganizing nearby prose).
I think that the B) and C) styles are similar enough that mixing them in a single version of the docs should not be jarring.
By the way, one additional benefit of a custom directive is that we can add syntax highlighting. (Ideally, with support from the theme.) I think that making strings stand out makes the listings more readable:
As a second step, I'd like to rewrite token documentation, and then the lexical analysis chapter, on which the grammar chapters build: #135676
Then, continue with the grammar chapters in Language reference:
- #141984
- Simple statements
- Compound statements
- Top-level components
- Full Grammar specification
And after that, use generated grammar snippets wherever possible.
Linked PRs
- gh-127835
- gh-129689
- gh-129690
- gh-129692
- gh-130376
- gh-131468
- gh-131474
- gh-132407
- gh-133632
- gh-133633
- gh-133652
- gh-134423
- gh-134443
- gh-134850
- gh-135301
Other related PRs
- gh-130588
- gh-134034
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
調査の方向性
Grammar/python.gram と Tools/peg_generator/docs_generator.py を読み、その後 language-reference の ReST プロダクションリストとリンクされた PR を調べます。文法ドキュメントと周辺の文章が実際の文法と一致し、提案されたプロダクションリストの表示方法とツールの変更がレビューされ適用された時点で、作業は完了です。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python
- 領域
- documentation
- issue の種類
- ドキュメント
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 活発さ
- 停滞
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 25/100