saltstack / saltstack/salt

[BUG] 3006.7 Minion makes the pillar-schedules persistent in failover process

Open
#67,980 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug needs-triage
Dominant language
Python
Stars
15.7k
Forks
5.6k
Avg merge
2d 44m
Merged PRs (30d)
80

Description

Description
Salt Minion dump the pillar-schedules to persistent store(_schedule.conf) in the failover process

Setup
Run 2 masters and minion in the multi-master setup.

On masters set the pillar-schedule:

$ cat /srv/pillar/schedule.sls
schedule:
  get-grains:
    enabled: true
    function: grains.items
    jid_include: true
    maxrunning: 1
    metadata: {get-common-grains: true}
    seconds: 60
    splay: 1
    name: get-grains
    run: true

$ cat /srv/pillar/top.sls
base:
  '*':
    - schedule

Minion multi-master conf:

$ cat /etc/salt/minion.d/failover.conf
master:
- 10.10.10.1
- 10.10.10.2
master_type: failover
verify_master_pubkey_sign: True
random_master: True
master_alive_interval: 30
master_failback: False
master_failback_interval: 0
retry_dns: 0

The Minion's _schedule.conf includes only default scheduled jobs:

$ cat /etc/salt/minion.d/_schedule.conf
schedule:
  __master_alive_10.10.10.1:
    enabled: true
    function: status.master
    jid_include: true
    kwargs: {connected: true, master: 10.10.10.1}
    maxrunning: 1
    return_job: false
    seconds: 30
  __mine_interval: {enabled: true, function: mine.update, jid_include: true, maxrunning: 2,
    minutes: 60, return_job: false, run_on_start: true}
  • on-prem machine
  • VM (Virtualbox, KVM, etc. please specify)
  • VM running on a cloud service, please be explicit and add details
  • container (Kubernetes, Docker, containerd, etc. please specify)
  • or a combination, please be explicit
  • jails if it is FreeBSD
  • classic packaging
  • onedir packaging
  • used bootstrap to install

Steps to Reproduce the behavior
Shutdown the master that the minion is connected.
Wait some time required for the failover process.
Ensure wat minion is connected to the new master.
Power up the first master and repeat failover case but now the second master must be stopped.
Look at /etc/salt/minion.d/_schedule.conf and u will see the schedule from pillar.

$ cat /etc/salt/minion.d/_schedule.conf | grep "get-grains:" -A 9
  get-grains:
    enabled: true
    function: grains.items
    jid_include: true
    maxrunning: 1
    metadata: {get-common-grains: true}
    name: get-grains
    run: true
    seconds: 60
    splay: 1

Expected behavior
The /etc/salt/minion.d/_schedule.conf not contains pillar-schedule

$ cat /etc/salt/minion.d/_schedule.conf | grep "get-grains:" -A 9
$

Versions Report

salt --versions-report
Salt Version:
          Salt: 3006.7

Python Version:
        Python: 3.10.15 (main, Jan 28 2025, 19:31:30) [GCC 8.5.0]

Dependency Versions:
          cffi: 1.14.6
      cherrypy: unknown
      dateutil: 2.8.1
     docker-py: Not Installed
         gitdb: Not Installed
     gitpython: Not Installed
        Jinja2: 3.1.3
       libgit2: 1.9.0
  looseversion: 1.0.2
      M2Crypto: Not Installed
          Mako: Not Installed
       msgpack: 1.0.2
  msgpack-pure: Not Installed
  mysql-python: Not Installed
     packaging: 22.0
     pycparser: 2.21
      pycrypto: Not Installed
  pycryptodome: 3.19.1
        pygit2: 1.17.0
  python-gnupg: 0.4.8
        PyYAML: 6.0.1
         PyZMQ: 23.2.0
        relenv: 0.18.0
         smmap: Not Installed
       timelib: 0.2.4
       Tornado: 4.5.3
           ZMQ: 4.3.4

System Versions:
          dist: astra 1.7_x86-64 1.7_x86-64
        locale: utf-8
       machine: x86_64
       release: 6.1.50-1-generic
        system: Linux
       version: Astra Linux 1.7_x86-64 1.7_x86-64

Additional context
The patch what helps me to fix the issue:

clear-pillar-schedules-in-failover-schedules-dump.patch
Subject: [PATCH] fix(minion/failover): clear pillar-schedules in schedules
 dump

---
 salt/minion.py | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/salt/minion.py b/salt/minion.py
index c2e1551137..e2e3c4e6a6 100644
--- a/salt/minion.py
+++ b/salt/minion.py
@@ -2890,7 +2890,7 @@ class Minion(MinionBase):
                         )
 
                         # put the current schedule into the new loaders
-                        self.opts["schedule"] = self.schedule.option("schedule")
+                        self.opts["schedule"] = self.schedule._get_schedule(include_pillar=False, remove_hidden=True)
                         (
                             self.functions,
                             self.returners,
-- 

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in salt/minion.py around the failover logic that puts the current schedule into the new loaders, then compare the existing schedule handling with the provided patch. Reproduce the two-master failover sequence and inspect /etc/salt/minion.d/_schedule.conf; done means pillar-defined schedules such as get-grains are absent after failover while default jobs remain.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
devops, infrastructure
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.