Stored fields and term vectors do not detect a corrupt chunk on the read path
- Dominant language
- Java
- Stars
- 3.6k
- Forks
- 1.4k
- Avg merge
- 2d 11h
- Merged PRs (30d)
- 88
Description
### Description
A corrupt byte inside a compressed chunk of stored fields or term vectors is not detected on the read path. It surfaces as whatever the decompressor happens to do with it — an `ArrayIndexOutOfBoundsException` from LZ4, or no error at all and a different document than the one that was stored.
The per-file CRC32 footer does not help here. It exists for every format, but it is verified either wholesale from `checkIntegrity` — which runs at merge time and in `CheckIndex` — or incidentally while reading a file that is read whole at open, like `.fnm` or `.si`. Neither fires when a document is fetched from `.fdt`.
This was proposed on #6331 in 2013, while diagnosing exactly such an `AIOOBE` from a corrupt `.fdt`. @jpountz:
> maybe we should add 4 bytes of checksum per chunk in order to be able to distinguish index corruptions from bugs in the compression layer
It was never built. #10396 then added a test that flipped bytes and expected corruption to be detected, and that test was disabled in `cf8f83c` — *"Disable test, some corruptions are still not detected as corruptions"* — because it is not, without a checksum.
### Measurement
Corrupting one byte at each of 102 sampled positions across the `.fdt` of a 500-document index, reading every document back and comparing:
| outcome | current |
|---|---|
| `ArrayIndexOutOfBoundsException` from LZ4 | 0 |
| some other error | 82 |
| **wrong document returned, no error at all** | **16** |
The 16 are the case that matters: a query returns a document whose stored fields are not what was indexed, and nothing reports a problem. A validity check inside the decompressor cannot catch that — most corrupt bytes still decode to *something*, just not to what was compressed.
### What the format already specifies
The LZ4 frame format defines an optional 4-byte xxHash-32 per compressed block, and states its purpose directly:
> calculated by using the xxHash-32 algorithm on the raw (compressed) data block […] The intention is to detect data corruption (storage or transmission errors) immediately, **before decoding**.
Lucene implements the LZ4 *block* format, which carries no checksum of its own, so the mechanism has never been available — the 2013 proposal was asking for something the surrounding format already standardises.
### Proposal
Record a CRC32C of each chunk's compressed bytes and verify it before the chunk is decompressed, in both `Lucene90CompressingStoredFieldsFormat` and `Lucene90CompressingTermVectorsFormat`.
CRC32C rather than xxHash-32: it is in the JDK, hardware-accelerated (measured at 8,900 MB/s here, against 1–3 GB/s for the LZ4 decompression it guards), and already the primitive behind `CodecUtil`'s footers.
Cost is 4 bytes per chunk — 0.024% at the 16 KB chunks of `BEST_SPEED`, 0.0065% at the 60 KB chunks of `BEST_COMPRESSION`.
Existing indexes are unaffected: the format version gates the check, older segments are read exactly as before, and a segment acquires checksums when it is next merged.
I have this working with tests and back-compat verified in both directions; PR to follow.
Contributor guide
Research direction
Start with Lucene90CompressingStoredFieldsFormat and Lucene90CompressingTermVectorsFormat, then inspect the existing corruption and back-compat tests mentioned in the issue. Trace chunk reads and format-version handling; done means CRC32C checks are verified before decompression in both formats, with older segments still readable and the tests passing.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- search
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Clearly specified
- Newbie friendliness
- 35/100