lance-format / lance-format/lance

Allow the part files to be skipped when training FTS

Open
#5,970 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

A-index enhancement
Dominant language
Rust
Stars
7.1k
Forks
852
Avg merge
3d 18h
Merged PRs (30d)
272

Description

The FTS index runs in two phases.

First, workers scan the column and tokenize the input. If this tokenized input gets too large then the data is spilled to disk (part files). This is controlled by LANCE_FTS_PARTITION_SIZE.

Second, the part files are scanned one at a time and written into one large index. However, this large index is sharded. As the large index gets too large it writes out a shard. This is controlled by LANCE_FTS_TARGET_SIZE. At search time these shards are searched in parallel.

We could skip the intermediate write if we interleave tokenizing with the construction of the final index. When a part is big enough to spill, instead of spilling, it could be sent on a shared channel to a writer thread. The writer thread would immediately flush the part into the index builder. The index builder would then spill as normal (we would still have LANCE_FTS_TARGET_SIZE but no longer use LANCE_FTS_PARTITION_SIZE). From my experiments this would cut the index building time roughly in half (perhaps even a better perf. boost on systems with many cores)

Perhaps more significantly, it would reduce the temporary disk space required to build an FTS index.

The downside is that it would possibly result in higher RAM consumption. Although we could presumably limit the size of the channel between the tokenizer threads and the writer thread which should still bound the total RAM usage.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating the FTS training implementation and the code paths controlled by LANCE_FTS_PARTITION_SIZE and LANCE_FTS_TARGET_SIZE. Compare the current part-file workflow with the proposed tokenizer-to-writer channel and bounded buffering. Done should mean intermediate part-file writes can be skipped while index sharding remains correct, temporary disk use is reduced, and RAM consumption stays bounded.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
search
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.