Separate syntactic and semantic nodes?

Aperta
#5,159 2 commenti 1 reazione 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
25/100
Tipo di issue
Refactoring
Chiarezza
Da chiarire
Stato di attività
Ferma
Stack tecnologico
python
Ambito
compilers

Direzione di ricerca

Start with fixup.py and its use of NodeVisitor, then inspect the existing ClassDef/TypeInfo and Var/AssignmentStmt split described in the issue. Compare the current AST, symbol-table, and type visitors with the proposed semantic-node hierarchy; done would require an agreed, incremental refactoring plan rather than a single localized change.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

needs discussion priority-2-low refactoring size-large

There is something that I was thinking about for a year or so that became more apparent while working on type alias refactoring (and planning the module refactoring). We have two different things mixed together in mypy code: syntactic nodes and "semantic" nodes. This is most obvious in fixup.py where we use NodeVisitor interface (intended to traverse syntactic tree) to patch symbol tables (something that carries purely semantic info). As a result, there are various weird things like visit_var, visit_type_info, etc. The idea is to have clear separation between syntactic and semantic nodes and have separate visitors for them. We currently already have this partially, we have TypeInfo vs ClassDef, and Var vs AssignmentStmt.

Here is a short (and approximate) summary of the proposal:

  • There are following semantic nodes: Var, Function, Class (non-leaf, has symbol table), Module (non-leaf, ditto), TypeVar, TypeAlias, ConditionalNode (see below).
  • The above nodes inherit from SymbolNode, while syntactic nodes inherit from Node, these two both inherit from Context.
  • Only the above nodes (and types) are serialized in cache and deserialized in incremental runs.
  • Semantic nodes can have attributes that point to the relevant (defining) syntactic nodes (for example variable can have defining assignment, or function statement if property), but we should limit this to minimize cache size.
  • The above nodes will have their separate visitor, so that in total we have three: for AST, for symbol tables, and for types.
  • ConditionalNode exists to avoid having .nodes instead of .node in SymbolTableNodes (the latter are thin wrappers around semantic nodes), while still supporting certain conditional definitions like conditional imports. This node will have an attribute that is a list of other semantic nodes.

This is a very large refactoring, but IMO this will add robustness and clarity, and will simplify addition of new features (e.g. conditional imports). This is probably a low priority (long term) thing. We just discussed this with @JukkaL, he likes the idea but is concerned about the size of this refactoring. A possible way to go forward with this is to split this in several separate PRs.

Lingua principale
Python
Stelle
20.6k
Fork
3.3k
Merge medio
1g 18h
PR unite (30g)
54

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di python/mypy

Tutte le issue di python/mypy

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.