apache / apache/lucene

Document Length Normalization in BM25Similarity correct? [LUCENE-8000]

Open
#9,048 11 comments 0 reactions 0 assignees View on GitHub
legacy-jira-priority:Minor type:bug
Dominant language
Java
Stars
3.6k
Forks
1.4k
Avg merge
2d 11h
Merged PRs (30d)
88

Description

Length of individual documents only counts the number of positions of a document since discountOverlaps defaults to true.

```Java
`@Override`
public final long computeNorm(FieldInvertState state) {
final int numTerms = discountOverlaps ? state.getLength() - state.getNumOverlap() : state.getLength();
int indexCreatedVersionMajor = state.getIndexCreatedVersionMajor();
if (indexCreatedVersionMajor >= 7) {
return SmallFloat.intToByte4(numTerms);
} else {
return SmallFloat.floatToByte315((float) (1 / Math.sqrt(numTerms)));
}
}}
```

Measureing document length this way seems perfectly ok for me. What bothers me is that
average document length is based on sumTotalTermFreq for a field. As far as I understand that sums up totalTermFreqs for all terms of a field, therefore counting positions of terms including those that overlap.

```Java
protected float avgFieldLength(CollectionStatistics collectionStats) {
final long sumTotalTermFreq = collectionStats.sumTotalTermFreq();
if (sumTotalTermFreq <= 0) {
return 1f; // field does not exist, or stat is unsupported
} else {
final long docCount = collectionStats.docCount() == -1 ? collectionStats.maxDoc() : collectionStats.docCount();
return (float) (sumTotalTermFreq / (double) docCount);
}
}
}
```

Are we comparing apples and oranges in the final scoring?

I haven't run any benchmarks and I am not sure whether this has a serious effect. It just means that documents that have synonyms or in my use case different normal forms of tokens on the same position are shorter and therefore get higher scores than they should and that we do not use the whole spectrum of relative document lenght of BM25.

I think for BM25 discountOverlaps should default to false.

---
Migrated from [LUCENE-8000](https://issues.apache.org/jira/browse/LUCENE-8000) by Christoph Goller, updated Oct 23 2017

Contributor guide

Open the contributing guide

Research direction

Start with BM25Similarity's computeNorm and avgFieldLength methods shown in the issue, then trace how discountOverlaps, field lengths, and sumTotalTermFreq are used during scoring. Compare overlap handling in both calculations and determine whether the default should change; document the conclusion and validate it with a focused scoring or benchmark test.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
search
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.