apache / apache/lucene

SynonymGraphFilter: wrong output token position when input positions overlap

Open
#12,080 1 comment 0 reactions 0 assignees View on GitHub
type:bug
Dominant language
Java
Stars
3.6k
Forks
1.4k
Avg merge
2d 11h
Merged PRs (30d)
88

Description

### Description

In my example, the query is 'test polskie'.
I use MorfologikFilter for Polish stemming, it turns 'polskie' into 'polski' + 'polskie'.
I also use SynonymGraphFilter which turns 'polski' into 'pol'. It's applied **only for query**.
Here's what I see in quey analysis (token position in parenthesis):
```
Tokenizer: test(1) polskie(2)
MF: test(1) polskie(2) polski(2)
SGF: test(1) polskie(2) pol(3) polski(3).
```
When I search for "test polskie" with quotation marks, a document with the same text doesn't match, because SGF changes positions of tokens in query compared to index.

In documentation, the description for the old `SynonymFilter` says "_The position value of the new tokens are set such they all occur at the same position as the original token._" In `SynonymGraphFilter` instead they are set to a position after the previous token. Is that an intentional change? Doesn't seem so, because it doesn't work as expected in my example.

Looking at the code, it seems the problem is in https://github.com/apache/lucene/blob/main/lucene/analysis/common/src/java/org/apache/lucene/analysis/synonym/SynonymGraphFilter.java#L246:
`nextNodeOut = lastNodeOut + posLenAtt.getPositionLength();`

`nextNodeOut` is always set as the position after the current token, and that is later used as position of output token.
I tried to remove this line and instead set this field right after the call to `input.incrementToken()`, in line 340:
`nextNodeOut = lastNodeOut + posIncrAtt.getPositionIncrement();`
This sets it to the original token's position. This way the final positions are `SGF: test(1) polskie(2) pol(2) polski(2).` and my document does match. I didn't experience any unexpected side effects.

Hope this helps. I'm not familiar with the project enough to easilly submit a proper pull request, with tests and all.

### Version and environment details

lucene 9.4

Contributor guide

Open the contributing guide

Research direction

Start in lucene/analysis/common/src/java/org/apache/lucene/analysis/synonym/SynonymGraphFilter.java around the reported nextNodeOut assignment and reproduce the Polish analyzer example from the issue. Check that overlapping input positions preserve the expected output positions and that the quoted query matches the equivalent document without introducing regressions.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
search
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.