hashicorp / hashicorp/nomad

Running tasks with Vault templates should survive Vault outages using last-rendered secrets

Open
#28,366 3 comments 4 reactions 0 assignees View on GitHub
hcc/jira stage/waiting-reply theme/template theme/vault type/enhancement
Dominant language
Go
Stars
17k
Forks
2.1k
Avg merge
1d 9h
Merged PRs (30d)
105

Description

### Proposal

When Vault is unreachable, Nomad should keep already-running tasks alive and keep using the last successfully rendered template/env values. Secret *sync* can stop; the workload shouldn't.

Right now a Vault outage can still cascade into restarts / failed templates / pending allocs for anything that depends on `vault` + `template`. `vault_retry { attempts = 0 }` helps by waiting longer, but it's not the same as "stay up on last known good secrets."

### Background

This is related to #11209. That issue described the same class of outage (Vault 503 → jobs with templates restart / get stuck pending). It was closed by #11606 exposing consul-template knobs (`vault_retry`, etc.).

That improvement is useful and we use it (or plan to). It doesn't fully address the operator need from #11209: **don't convert a Vault outage into a full application outage for tasks that already have secrets on disk / in process env.**

### Use-cases

We run a large number of long-lived Nomad jobs that inject app config/secrets from Vault via templates (`env = true`, continuous render). Vault is a dependency for *changes* and for *new starts*, which is fine. What hurts is:

1. Vault blips / network partitions / maintenance.
2. Template watches or token renewal fail.
3. Tasks that were healthy get restarted or fail.
4. They can't start again until Vault is back → wide blast radius even if the apps themselves were fine with the last env.

We've hit this more than once. Prefer degraded mode: no secret updates, apps keep running, only deploys/restarts that need a fresh render are blocked.

### Attempted Solutions

- `template { change_mode = "noop" }` — helps for secret *content* changes, not for "Vault is down / render fails."
- Client `template.vault_retry` with high/unlimited attempts — delays failure, doesn't define last-good-secret semantics; also doesn't help "I need a clear emergency behavior" under prolonged outage.
- Making Vault HA — necessary, but doesn't remove the hard dependency on live Vault for template health.

### Possible directions (not prescriptive)

Something like:

- Client or job option: on Vault error after successful initial render, **log + keep last file/env**, don't fail the running task.
- Or explicit "persist last render" / emergency mode (similar ideas were floated on #11209).
- Clear docs on what is *guaranteed* during a Vault outage for running vs new allocs.

Happy to provide more detail on our jobspecs / client config if useful (redacted).

### Nomad version

Nomad v2.0.3

### Reproduction (high level)

1. Job with `vault {}` + `template` reading Vault (`env = true`, `once = false`).
2. Start job successfully.
3. Take Vault offline / return 503.
4. Observe template/renew paths eventually disrupting running work or blocking restarts; compare to desired "keep last secrets, stay up."

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the high-level reproduction: a job using vault plus a template with env=true and continuous rendering, then inspect the existing template, Vault retry, renewal, and task lifecycle behavior. Done should define and implement the behavior for running tasks after a successful render during a Vault outage, including logging, last-known values, and the effect on new allocations or restarts.

Written by the indexing model from the issue text.

Assessment

Tech stack
go
Domain
infrastructure
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.