ClusterLabs / ClusterLabs/resource-agents
Cause for FileSystem monitor timeout
- Dominant language
- Shell
- Stars
- 519
- Forks
- 608
- Avg merge
- 6d 1h
- Merged PRs (30d)
- 7
Description
I am using Pacemaker to manage a Postgres cluster, with 2 servers and a shared storage disk.
The disks are mounted on the master/active node using the resource agent Filesystem.
The Filesystem monitor with default settings (interval=20s timeout=40s), timed out and caused the Postgres to failover.
I checked the disks usage, memory usage, Network IO, Disk IO etc., and everything looks normal before the aforementioned monitoring timeout.
I have been scratching my head to find out what caused that timeout. Information in the cluster logs only mentions about the timeout, but no reason. Is there a way to further investigate this issue? Could it have been the DELL storage issue?
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue identifies the Filesystem monitor, its interval and timeout settings, and the cluster logs as starting points; review the timeout and failover evidence around the reported event. The payload names no files, tests, reproducible case, or specific code change, so a clear definition of done is missing.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- postgresql, shell
- Domain
- databases, distributed-systems, infrastructure
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100