NethermindEth / NethermindEth/pluto

Revisit DKG timeout strategy and add test coverage

Open
#465 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Rust
Stars
8
Forks
5
Avg merge
4d 16h
Merged PRs (30d)
37

Description

Description

The DKG timeout works differently from Charon. Charon splits the timeout per phase (conf.Timeout / 6); Pluto uses one overall timeout for the whole run, because our sync service runs across all phases and can't be cut per-phase (we tried — it killed healthy runs).

Two gaps come from this. A stuck peer is only caught after the full timeout (~60s) instead of one phase (~10s). And with many validators (untested) a healthy run takes longer and could hit the timeout and be killed by mistake — made worse by parsigex already using the full conf.timeout per exchange.

Works today: stuck peer aborts cleanly, and a small healthy cluster finishes fine. Not covered: many-validator runs, and there's no automated test for the timeout.

Benefit (The "Why")

A DKG should never hang forever on a stuck peer, but it also must never kill a healthy ceremony by mistake. Closing these gaps makes the timeout reliable for real clusters (many validators), gives faster failure detection, and protects it with an automated test.

Acceptance Criteria
  • A healthy DKG with many validators completes without a false timeout under the default --timeout.
  • The relationship between the overall timeout and the per-exchange parsigex timeout is consistent (one budget is not silently exceeded by the other).
  • A stuck/absent peer still aborts cleanly with DkgError::Timeout.
  • An automated test covers the overall timeout (stuck peer aborts) and the healthy path (no false timeout).

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by tracing the DKG run's overall --timeout handling and the per-exchange parsigex timeout, including the DkgError::Timeout path. Add coverage for a stuck or absent peer and for a healthy run with many validators. Done means both paths are reliable and the timeout budgets are consistent without false aborts.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
blockchain, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.