apache / apache/datafusion-sqlparser-rs

[EPIC] Improve sqlparser performance

Abierto
#1,557 5 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Rust
Estrellas
3.5k
Forks
772
Merge medio
4 d 9 h
PR fusionados (30 d)
17

Descripción

## What problem are you trying to solve?
Normally, in a SQL processing system, parsing SQL is not a major bottleneck compared to actually processing data. That being said, given how many SQL strings are parsed by this crate, I think there is significant benefit to improving the performance of the SQL parser in this crate.

That being said, I also think it is important to minimize the impact on downstream crates as much as possible.

Recently, we started [introducing locations into the parser](https://github.com/apache/datafusion-sqlparser-rs/pull/1435) (thanks again @Nyrox!), which we found slows things down a bit (see https://github.com/apache/datafusion-sqlparser-rs/pull/1435#issuecomment-2500664144).

Thankfully, I think there is significant room for improvement. As as part of the adding location information, I spent some time profiling and I think there are some obvious ways to improve the performance without impacting downstream crates.

Here is the flamegraph for anyone who is interested (you can download it locally to get zoom / etc):

[fixed-flamegraph](https://github.com/user-attachments/assets/7ceae3a0-dec2-462d-9a00-0b49afd9d2bc)

![fixed-flamegraph](https://github.com/user-attachments/assets/7ceae3a0-dec2-462d-9a00-0b49afd9d2bc)

## What would you like to see?
The idea would be
1. Run the benchmarks (instructions in https://github.com/apache/datafusion-sqlparser-rs/pull/1555)
2. Maybe add additional benchmarks so they are more representative
3. Improve the benchmarks

## Ideas to improve performance:
- The most obvious one is to next_token / peek to not clone each `Token` (which involves copying strings)L: https://github.com/apache/datafusion-sqlparser-rs/issues/1558
- https://github.com/apache/datafusion-sqlparser-rs/issues/1381

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Línea de trabajo

Comienza con las instrucciones de benchmarks de PR #1555 y ejecuta los benchmarks existentes; después, revisa el flamegraph enlazado y las ideas de rendimiento de los issues #1558 y #1381. Se considera completado cuando se hayan añadido benchmarks representativos donde sea necesario y se haya mejorado el rendimiento del parser sin afectar de forma significativa a los crates posteriores.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
rust
Área
databases, performance
Tipo de issue
Refactorización
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Estancado
Claridad
Necesita aclaración
Aptitud para principiantes
25/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.