Clearly explain that draining a swarm node does not wait for replcas to be started on an active node before stopping tasks on a node being drained
Nobody has claimed this yet.
- Dominant language
- Markdown
- Stars
- 4.7k
- Forks
- 8.5k
- Avg merge
- 2d 18h
- Merged PRs (30d)
- 108
Description
File: engine/swarm/swarm-tutorial/drain-node.md
States:
"Sometimes, such as planned maintenance times, you need to set a node to DRAIN availability. DRAIN availability prevents a node from receiving new tasks from the swarm manager. It also means the manager stops tasks running on the node and launches replica tasks on a node with ACTIVE availability."
This is misleading in that a drain operation has no logic to maintain the configured number of replicas during a drain operation.
This should be clearly explained.
If you have a two worker node swarm and have performed maintenance on worker node 1, this has all replicas running on worker node 2.
If you then drain worker node 2 for patching, it causes downtime because swarm doesn't for example, stop replica 1 on node 2, start replica 1 on node 1 before moving on to do the same for replica 2.
The current design causes downtime for applications.
Support advised this is expected behavior and a workaround is to reconfigure all running services to have more replicas to force them to start on another worker node before issuing a drain command for a node.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Open engine/swarm/swarm-tutorial/drain-node.md and read the section describing DRAIN availability and replica tasks. Clarify that draining does not wait for replacement replicas to start before stopping tasks, and include the documented workaround described in the issue; the page should accurately set expectations about possible downtime.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- docker
- Domain
- devops, documentation
- Issue type
- Documentation
- Difficulty
- 1/5
- Estimated time
- Under an hour
- Activity status
- Stale
- Clarity
- Clearly specified
- Newbie friendliness
- 55/100