infiniflow / infiniflow/infinity
[Bug]: WAL replay of an entry referencing a deleted data file crashes InitPhase2 (exit 139), still on main with #3429 — is a supported way to discard unreplayable entries possible?
- Dominant language
- C++
- Stars
- 4.7k
- Forks
- 445
- Avg merge
- 2d 2h
- Merged PRs (30d)
- 7
Description
### Is there an existing issue for the same bug?
- [x] I have checked the existing issues.
Related: **#3435** (same reporter, same instance — WAL files are never recycled and manual cleanup by timestamp/size crashes InitPhase2). This issue reports a related but different failure mode, and specifically tests whether **#3429** covers it. As far as we can measure, it does not.
### Version or Commit ID
Two builds, tested side by side on identical fresh copies of the same data volume:
- **Control:** `v0.7.3-0-d755c5a` (image `infiniflow/infinity:v0.7.3-x64-v3`, build 2026-08-06) — the build our production instance runs.
- **Treatment:** `main@-0-eca7266` (image `infiniflow/infinity:nightly-x64-v3`, digest `sha256:9203ca30ac080e62f1386bb18fa9912d72906abced1ba59c7d4ec4963473bcb7`, build 2026-09-16 13:14), which includes `807af1b05` (#3429, merged 2026-09-04).
### Other environment information
```Markdown
Infinity deployment: Docker, standalone mode, via RAGFlow v0.27.1
Data layout: default /var/infinity layout on a named Docker volume (wal/, data/, persistence/, catalog/)
Same instance as #3435: at the time of this test, wal/ held 58,742 files (1.6 GB), 51,147 of them exactly 82 bytes
Config: stock infinity_conf.toml, default paths, log_level=info
```
### Actual behavior and How to reproduce it
**Context — how the volume got into this state**
The engine crashed during a drop/cleanup operation. The WAL entries written after the last checkpoint reference data files that the crashed operation had already deleted. Since then, every boot of this volume crashes; the only recovery we have found is manually quarantining `wal.log*` files (moving, not deleting), which forfeits the entire post-checkpoint tail.
**What happens on boot (6/6 runs, deterministic)**
1. Startup reads the catalog fine and finds the latest checkpoint: `Find and set checkpoint max commit ts: 499869`.
2. WAL replay phases 1–2 complete: **15 entries are found after the checkpoint and parse without any checksum/size error** — the WAL stream itself is intact.
3. While *executing* the replay of one of those entries (a `WalCmdAppendV2`), the engine fails to load the data file the entry references:
```
[critical] GetObjCache Failed to find object for local path db_1/tbl_11/seg_0/blk_0/col_3_out: Key not found: pm|object|db_1/tbl_11/seg_0/blk_0/col_3_out
[critical] Error: Failed to find object for local path /var/infinity/data/db_1/tbl_11/seg_0/blk_0/col_3_out@src/storage/buffer/file_worker/file_worker_impl.cpp:284
...
[critical] Error: TerminateHandler: Unhandled Exception: Failed to find object for local path /var/infinity/data/db_1/tbl_11/seg_0/blk_0/col_3_out@src/storage/buffer/file_worker/file_worker_impl.cpp:284
```
4. The terminate handler logs the unhandled exception and the process dies with exit code 139 (SIGSEGV) ~60 seconds into every start. Stack trace from the nightly build:
```
0# infinity::PrintStacktrace@infinity_core(...) at /infinity/src/common/utility/exception_impl.cpp:48
1# infinity::UnrecoverableError@infinity_core(...) at /infinity/src/common/utility/exception_impl.cpp:85
2# infinity::FileWorker@infinity_core::GetFilePathInner[abi:cxx11](bool) at /infinity/src/storage/buffer/file_worker/file_worker_impl.cpp:284
3# infinity::FileWorker@infinity_core::ReadFromFile(bool) at /infinity/src/storage/buffer/file_worker/file_worker_impl.cpp:207
4# infinity::BufferObj@infinity_core::LoadNoLock(bool) at /infinity/src/storage/buffer/buffer_obj_impl.cpp:117
5# infinity::BufferObj@infinity_core::Load(bool) at /infinity/src/storage/buffer/buffer_obj_impl.cpp:178
6# infinity::VarBufferManager@infinity_core::GetInnerNoLock() at /infinity/src/storage/column_vector/var_buffer_impl.cpp:187
7# infinity::VarBufferManager@infinity_core::TotalSize() at /infinity/src/storage/column_vector/var_buffer_impl.cpp:211
8# infinity::NewTxn@infinity_core::AppendInColumn(...) at /infinity/src/storage/new_txn/new_txn_data_impl.cpp:1016
9# infinity::NewTxn@infinity_core::AppendInBlock(...) at /infinity/src/storage/new_txn/new_txn_data_impl.cpp:948
10# infinity::NewTxn@infinity_core::CommitBottomAppend(infinity::WalCmdAppendV2@infinity_core*) at /infinity/src/storage/new_txn/new_txn_data_impl.cpp:1882
11# infinity::NewTxn@infinity_core::CommitBottom() at /infinity/src/storage/new_txn/new_txn_impl.cpp:4578
12# infinity::NewTxn@infinity_core::CommitReplay() at /infinity/src/storage/new_txn/new_txn_impl.cpp:2151
13# infinity::NewTxnManager@infinity_core::CommitReplayTxn(...) at /infinity/src/storage/new_txn/new_txn_manager_impl.cpp:327
14# infinity::WalManager@infinity_core::ReplayWalEntries(...) at /infinity/src/storage/wal/wal_manager_impl.cpp:719
15# infinity::Storage@infinity_core::AdminToWriter() at /infinity/src/storage/storage_impl.cpp:326
16# infinity::Storage@infinity_core::SetStorageMode(...) at /infinity/src/storage/storage_impl.cpp:835
17# infinity::InfinityContext@infinity_core::ChangeServerRole(...) at /infinity/src/main/infinity_context_impl.cpp:132
18# infinity::InfinityContext@infinity_core::InitPhase2(bool) at /infinity/src/main/infinity_context_impl.cpp:90
19# main at /infinity/src/bin/infinity_main.cpp:285
```
**Why this looks like #3429 territory but isn't covered by it**
#3429 fixes the WAL *reading* layer (out-of-bounds reads on corrupt entries, over-aggressive purge) — and in our repro the reading layer works: all 15 post-checkpoint entries are found and parsed cleanly on **both** builds. The crash is in the *execution* of replay — `CommitReplay` → `CommitBottomAppend` → append into a column whose backing file no longer exists. We verified this empirically with a controlled A/B:
- Two **fresh identical clones** of the poisoned volume (`cp -a` into disposable volumes; source mounted read-only; one clone per arm, so the control run could not contaminate the treatment run).
- **Control arm** — `v0.7.3-0-d755c5a`: 3/3 boots → exit 139.
- **Treatment arm** — `main@-0-eca7266` (with #3429): 3/3 boots → exit 139, identical critical line (`file_worker_impl.cpp:284`) and the same stack (the only difference between builds is one frame's source line: `new_txn_data_impl.cpp:1881` vs `:1882`).
| Arm | Build | Boot 1 | Boot 2 | Boot 3 |
|---|---|---|---|---|
| Control | v0.7.3-0-d755c5a | exit 139 (~67 s) | exit 139 (~61 s) | exit 139 (~60 s) |
| Treatment | main@-0-eca7266 (incl. #3429) | exit 139 (~64 s) | exit 139 (~60 s) | exit 139 (~61 s) |
So this is not a one-off crash report: the same poisoned WAL that crash-loops `v0.7.3` also crash-loops the current tip of `main`.
**How we ran each arm (for reference)**
```
docker volume create inf_rehearsal_ctrl
docker run --rm -v :/src:ro -v inf_rehearsal_ctrl:/dst \
alpine sh -c 'cp -a /src/. /dst/'
docker create --name inf-test --network inf-test-net \
--memory 8073741824 --restart=no \
-v inf_rehearsal_ctrl:/var/infinity \
infiniflow/infinity:v0.7.3-x64-v3 -f /infinity_conf.toml
docker cp infinity_conf.toml inf-test:/infinity_conf.toml
docker start inf-test # ~60 s later: exit 139
docker start inf-test # repeatable
```
(For the treatment arm, same commands with the nightly image referenced by digest. Standalone mode; `--memory` mirrors what our production container runs with; `--restart=no`, boots restarted manually to confirm determinism.)
### Expected behavior
Two things we'd love your thoughts on:
1. **Is there, or could there be, a supported way to selectively discard a WAL entry whose replay references a missing data file — without discarding the rest of the post-checkpoint WAL?** The engine already detects the exact situation and logs it (`GetObjCache Failed to find object ... Key not found`) right before throwing `UnrecoverableError`. Skipping that one transaction — with a loud warning and some durable marker, e.g. in the catalog — would let such a volume boot with a bounded, explicit loss instead of requiring manual WAL surgery. #3429's purge is intentionally conservative at file granularity (and its tests argue well for that); this would be one level more surgical: entry-level, at replay execution.
2. **Is the failure mode itself intended?** `UnrecoverableError` during replay propagates through the terminate handler and the process dies with SIGSEGV (exit 139) rather than exiting cleanly with the message. A clean non-zero exit with the same message would already make automated boot-loop detection and restart policies much easier to handle.
For our part: the operational workaround we use today (also referenced in #3435) is moving — never deleting — post-checkpoint `wal.log*` files into a quarantine directory. That boots the instance but forfeits the whole tail. We are not set up to contribute a fix ourselves — this report is mainly to (a) document, with a deterministic reproduction, that #3429 does not cover this path, and (b) ask whether a supported "discard unreplayable entry" path exists or is planned.
Happy to share more logs/details from the test environment if useful.
Thanks!
Gisselle
Contributor guide
Research direction
Start with src/storage/wal/wal_manager_impl.cpp:719 and trace replay through src/storage/new_txn/new_txn_manager_impl.cpp:327 and new_txn_impl.cpp:2151. Inspect the missing-file path at src/storage/buffer/file_worker/file_worker_impl.cpp:284 and the existing replay tests; done should define and verify safe behavior for an unreplayable entry, including a clear boot outcome.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp, docker
- Domain
- databases
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100