ozontech / ozontech/seq-db

Optimizations for `re` filter

Open
#355 0 comments 0 reactions 1 assignee View on GitHub

@dkharms is already working on this.

Since Feb 16, 2026.

performance
Dominant language
Go
Stars
131
Forks
16
Avg merge
2d 4h
Merged PRs (30d)
11

Description

In #348, we decided to use anchored regular expressions by default.
This change was necessary to enable several optimizations for regular expression matching.

Prefix and Suffix Literals

Before matching all tokens for a field from a search query, we can extract two literals from the regular expression: prefix and suffix literals. Using these two literals, we can apply two key optimizations:

  • Reduce the number of token blocks to examine by performing a binary search over token blocks using the (strings|bytes).HasPrefix() function (as we already do for literals);
  • For each token we match against, use (strings|bytes).HasSuffix() before running the regular expression NFA.
Set Lookups

Sometimes we can transform a regular expression into a basic search query. For example, consider this query using the re filter: k8s_pod:re("pod-1|pod-2"). It is easy to see that we can rewrite it as k8s_pod:'pod-1' OR k8s_pod:'pod-2'.

I am pretty sure there are many more optimizations we can implement, so this is a topic worth researching further.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.