Using `sql_tokenize()` table function for tokenization.
- Dominant language
- Python
- Stars
- 186
- Forks
- 113
- Avg merge
- 13h 29m
- Merged PRs (30d)
- 17
Description
### What happens?
Original issue here: https://github.com/duckdblabs/duckdb-internal/issues/6733
https://github.com/duckdb/duckdb/pull/20171 introduced a new table function `sql_tokenize`, which takes as input the query and returns the tokens and their category.
Currently in the python client we use a custom implementation somewhere near `PyTokenize` (see [here](https://github.com/duckdb/duckdb-python/blob/89ed9a1d66dcab2455cfd667462f1b90619b2aac/src/duckdb_py/duckdb_python.cpp#L35)). I think this can now be moved over to use the table function instead, making it a bit simpler, and possible to provide more information about tokenization in the future.
### To Reproduce
See above
### OS:
MacOS
### DuckDB Package Version:
latest
### Python Version:
3.14?
### Full Name:
Daniel ten Wolde
### Affiliation:
DuckDB Labs
### What is the latest build you tested with? If possible, we recommend testing with the latest nightly build.
I have tested with a stable release
### Did you include all relevant data sets for reproducing the issue?
Yes
### Did you include all code required to reproduce the issue?
- [x] Yes, I have
### Did you include all relevant configuration to reproduce the issue?
- [x] Yes, I have
Contributor guide
Research direction
Start in src/duckdb_py/duckdb_python.cpp near PyTokenize and trace how the Python client currently tokenizes queries. Compare that path with the sql_tokenize table function introduced by duckdb/duckdb#20171; done means the client uses the table function while preserving the returned tokens and categories.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, sql
- Domain
- api, database
- Issue type
- Refactor
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100