apache / apache/datafusion-sqlparser-rs

Optimize `Token::make_word`

オープン
#1,588 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Rust
スター
3.5k
フォーク
772
平均マージ
4日 9時間
マージ済み PR(30日)
17

説明

While working on #1587 I noticed that Instruments is showing `Token::make_word` as the second hottest single function, right after `alloc::raw_vec::finish_grow`.

Looking into the implementation I saw that its just doing a binary search across all keywords to find if its a known keyword or not. This is a fairly classical case where we have a known set of strings and want to check if a given string is in that list. There are a bunch of ways that we could speed this up. This issue is to figure out a good compromise between those possible speedups and other project constraints like maintaining a `no_std` ability.

My [first approach](https://github.com/apache/datafusion-sqlparser-rs/commit/4551933dc0a9e892e412be5ca0022a124859dad0) at speeding this up was to create a table for the first byte in every keyword to reduce the number of entries that need to be searched. This small optimization managed to shave off about 400ms of time (of the 1.4ish seconds total).

However, there are other approaches that could speed this up even more. Either by generating parsing/lookup tables or using something like [phf](https://crates.io/crates/phf) to do the heavy lifting for us.

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

Start with Token::make_word and the linked first approach commit to understand the current keyword lookup and its performance impact. Compare possible lookup-table or phf-based approaches while preserving the project's no_std constraint, then use the reported Instruments timing to verify that the chosen design improves performance without changing keyword recognition.

索引モデルが issue の本文から書いたものです。

評価

技術スタック
rust
領域
compilers
issue の種類
リファクタリング
難易度
5/5
見積もり時間
1週間以上
活発さ
停滞
明瞭さ
おおむね明確
初心者へのやさしさ
30/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。