ClusterLabs / ClusterLabs/resource-agents

Cause for FileSystem monitor timeout

Open
#1,695 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Shell
Stars
519
Forks
608
Avg merge
6d 1h
Merged PRs (30d)
7

Description

I am using Pacemaker to manage a Postgres cluster, with 2 servers and a shared storage disk.
The disks are mounted on the master/active node using the resource agent Filesystem.

The Filesystem monitor with default settings (interval=20s timeout=40s), timed out and caused the Postgres to failover.

I checked the disks usage, memory usage, Network IO, Disk IO etc., and everything looks normal before the aforementioned monitoring timeout.

I have been scratching my head to find out what caused that timeout. Information in the cluster logs only mentions about the timeout, but no reason. Is there a way to further investigate this issue? Could it have been the DELL storage issue?

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue identifies the Filesystem monitor, its interval and timeout settings, and the cluster logs as starting points; review the timeout and failover evidence around the reported event. The payload names no files, tests, reproducible case, or specific code change, so a clear definition of done is missing.

Written by the indexing model from the issue text.

Assessment

Tech stack
postgresql, shell
Domain
databases, distributed-systems, infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.