quickwit-oss / quickwit-oss/quickwit
Restart Pod of the control plane on a panic?
Open
Nobody has claimed this yet.
bug
- Dominant language
- Rust
- Stars
- 11.7k
- Forks
- 597
- Avg merge
- 2d 22h
- Merged PRs (30d)
- 37
Description
Describe the bug
Once panicked, would it be better to restart Pod?
thread 'tokio-runtime-worker' panicked at /usr/local/cargo/git/checkouts/chitchat-xxx/dxxx/chitchat/src/delta.rs:413:9:
assertion failed: mtu >= 100
we use quickwit/quickwit:v0.8.2
current liveness probe is:
livenessProbe:
failureThreshold: 3
httpGet:
path: /health/livez
port: rest
scheme: HTTP
periodSeconds: 10
successThreshold: 1
timeoutSeconds: 1
which leaves pods in a partially-failed state.
Possible solution is to use RUST_PANIC=abort.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by tracing the control plane's /health/livez behavior after the reported Rust panic and reviewing the provided Kubernetes livenessProbe configuration. Evaluate the proposed RUST_PANIC=abort approach and compare it with other restart behavior. Done means a panic leaves the pod in a recoverable state with a clear, tested restart outcome.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- kubernetes, rust
- Domain
- devops, infrastructure
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 52/100