bottlerocket-os / bottlerocket-os/bottlerocket

Make zram swap size proportional to node RAM (currently hardcoded to 1 GiB)

Open
#4,903 0 comments 10 reactions 0 assignees View on GitHub
status/needs-triage type/enhancement
Dominant language
Rust
Stars
9.7k
Forks
586
Avg merge
1d 11h
Merged PRs (30d)
11

Description

**What I'd like:**

That the zram swap device introduced in
[core-kit#590](https://github.com/bottlerocket-os/bottlerocket-core-kit/pull/590) should be sized such that it cannot mask kubelet's `memory.available` eviction signal — i.e. sized relative to the node's RAM *and* to the configured eviction thresholds, rather than as a fixed constant.

**What actually happened:**

`prepare-swap.service` creates a **fixed 1 GiB** device on every node, regardless of effective RAM and regardless of the eviction thresholds in effect:

```
ExecStart=/usr/sbin/zramctl /dev/zram0 --size 1G
```

Verified unchanged on 1.63.0 (16 GiB `r7g.large`):

```
# cat /sys/block/zram0/disksize
1073741824
# cat /proc/sys/vm/swappiness /proc/sys/vm/page-cluster
200
0
```

Because swapped-out anonymous pages leave `workingSet`, and kubelet computes `memory.available = capacity − workingSet`, the device is a **blind spot in the eviction signal of up to its own size**. When it is large relative to the eviction threshold, neither eviction nor the OOM killer fires and the node thrashes indefinitely instead of shedding load.

### The default configuration cannot trigger hard eviction

This is the part I'd most like maintainers to look at, because it is not a small-instance edge case — it affects **every** node running stock settings. Unless of course, my analysis is wrong ;-).

The kubelet config template emits eviction thresholds *only* if the user set
them ([`kubelet-config`](https://github.com/bottlerocket-os/bottlerocket-core-kit/blob/develop/packages/kubernetes-1.36/kubelet-config)):

```handlebars
{{#if settings.kubernetes.eviction-hard}}
evictionHard:
{{#each settings.kubernetes.eviction-hard}}
{{@key}}: "{{this}}"
{{/each}}
{{/if}}
```

Bottlerocket ships **no default**, so a stock node falls back to kubelet's own built-in default — `memory.available<100Mi`, absolute — **and no soft thresholds at all**.

Now consider a full zram device. Its pages are compressed into `zsmalloc` memory, which is unaccounted kernel memory: it occupies physical RAM but appears in nobody's `workingSet`. So with a device of size `Z` and compression ratio `r`, kubelet's reported `memory.available` has a hard floor of roughly `Z / r`, and cannot go below it no matter how much memory the workload demands:

| compression ratio | floor on `memory.available` with `Z` = 1 GiB |
|---|---|
| `r = 1` (incompressible; zram stores such pages raw) | 1024 MiB |
| `r = 3.0` (measured, low end) | 341 MiB |
| `r = 4.1` (measured, high end) | 250 MiB |

Every one of those is **above the 100 MiB default threshold**. So on a stock Bottlerocket node with a saturated zram device, `memory.available` never reaches `evictionHard`, at *any* instance size — 2 GiB or 768 GiB alike.

**Two defaults that ship together are mutually incompatible:** swap-on-by-default and kubelet's 100 MiB absolute hard-eviction default. Whichever way this issue is resolved, I think those two need to be reconciled.

### Why a fixed size is the wrong shape of knob

For users who *do* configure percentage thresholds, the failure becomes size-dependent instead of universal. A threshold is only reachable if its band exceeds the device. With our thresholds (`hard 5%` / `soft 15%`):

| Nominal RAM | hard band (5%) | soft band (15%) | hard reachable? | soft reachable? |
|---|---|---|---|---|
| 4 GiB | 205 MiB | 614 MiB | ❌ | ❌ |
| 8 GiB | 410 MiB | 1229 MiB | ❌ | ⚠️ 205 MiB margin |
| 16 GiB | 819 MiB | 2458 MiB | ❌ | ✅ |
| 32 GiB | 1638 MiB | 4915 MiB | ✅ | ✅ |

(Nominal RAM; kubelet's `capacity` is ~4–6% lower.)

**The core point is scale invariance.** A fixed 1 GiB device is safe only above some node size — ~20 GiB for a 5% hard threshold, ~6.7 GiB for a 15% soft one. A device sized as a fraction of the *threshold* is safe at every size, which is the property a node OS needs, since it cannot know what instance types or thresholds its users will run.

Note also that #4735's own validation was performed on a 7.6 GiB node — right at the boundary where the soft band stops exceeding the device.

### What this looked like in production

A `c5a.large` (3.7 GiB) node went into permanent reclaim thrash. kubelet reported **46% memory available** while only **91 MiB was actually free**, at ~29,700 page ops/s. No eviction and no OOM kill ever occurred; the node simply degraded until we replaced it.

Measurements, including an honest accounting of what masked what

- `memory` PSI `full` avg300: **21.71%**; `io` PSI `full`: **6.99%**
- zram: **318 MiB** of anon parked; compression ratio 3.0–4.1× (`mm_stat`)
- page cache: **~1.8 GiB**, hot and constantly refaulting
- A healthy node of identical uptime and instance family: `pswpin`/`pswpout` = 0
- Collateral damage: our monitoring agent's probes timed out → `exitCode 137`,
**1524 restarts**. That is what led us to investigate.

**To be precise about causation:** on this node the *larger* share of the
`memory.available` inflation was inactive page cache, not zram — that is the long-standing [kubernetes#43916](https://github.com/kubernetes/kubernetes/issues/43916) blind spot, which is not Bottlerocket's doing. zram contributed the smaller share directly.

Its real contribution was to make the state *stable*. With `vm.swappiness = 200` and a 1 GiB absorber, the kernel had somewhere to keep putting anon pages, so it ping-ponged indefinitely rather than being forced into reclaim that would have resolved the situation one way or the other. The two blind spots also **compose** — which is the main reason the proposal below sizes the device at a *fraction* of the threshold rather than up to it.

After moving the same workloads to 16 GiB nodes (no other change): `memory` PSI `full` **0.00** (`total = 0`), zram used **0**, 0 restarts.

We worked around this by enforcing a **≥16 GiB floor** on all node pools — a blunt and expensive workaround for what is really a sizing constant, and one that is unavailable to anyone whose workloads legitimately belong on small nodes.

### Existing workarounds are all unpleasant

- There is **no supported setting** to resize or disable the device.
`settings.kubernetes.memory-swap-behavior` (v1.42.0) governs *pod* swap only
and has no effect on the host device.
- The workarounds circulating in #4075 are a bootstrap container running
`swapoff` / `zramctl -r` / `modprobe -r zram` (with a retry loop for
`Device or resource busy`), or `vm.swappiness=60` via `EC2NodeClass` userData.
Both fight the OS rather than configure it, and the latter does not address the
accounting blind spot at all.
- Upstream Kubernetes will not solve this: making the eviction manager
swap-aware is an explicit **Non-Goal** of KEP-2400,
[kubernetes#129578](https://github.com/kubernetes/kubernetes/pull/129578) was
closed unmerged, and no PSI-based eviction signal is proposed.

## Proposed change

> [!NOTE]
> Maybe there are better solutions to this problem than what I propose here (I am really not a subject matter exper neither on ZRAM nor on kubelet's eviction feature). So, please do not hesitate to bring up more suitable alternatives.

### The invariant

Size the device so that a *full* device still leaves `memory.available` able to cross the tightest threshold that is meant to fire:

```
Z ≤ α · T where T = min over the configured memory.available
eviction thresholds (hard and soft),
resolved to bytes against MemTotal
```

`α = 0.5` is the value I'd suggest. The floor on `memory.available` is `Z / r`, so even at the worst case `r = 1` (incompressible data, which zram stores raw) `α = 0.5` leaves a 2× margin — and that margin is what absorbs the page-cache blind spot the two effects compose into.

Taking `min(hard, soft)` rather than one or the other is what makes *both* thresholds keep working, which is the property we actually want.

### Option A — derive the size from the thresholds

```
Z = min( α · T , Z_max ) Z_max = 4 GiB
```

With `hard 5% / soft 15%`, so `T = 5% · MemTotal`:

| Nominal RAM | `T` | proposed `Z` | today's `Z` |
|---|---|---|---|
| 4 GiB | 205 MiB | 102 MiB | 1024 MiB (10.0× too large) |
| 8 GiB | 410 MiB | 205 MiB | 1024 MiB (5.0× too large) |
| 16 GiB | 819 MiB | 410 MiB | 1024 MiB (2.5× too large) |
| 32 GiB | 1638 MiB | 819 MiB | 1024 MiB (1.25× too large) |
| 64 GiB | 3277 MiB | 1638 MiB | 1024 MiB (too *small*) |
| 128 GiB | 6554 MiB | 3277 MiB | 1024 MiB (too *small*) |

The constant is wrong in both directions — it defeats eviction below ~32 GiB and under-provisions the absorber above ~64 GiB. That asymmetry is, I think, the clearest argument that this should not be a constant.

If the computed size falls below a small floor (~64 MiB), I'd suggest simply not enabling swap: at that point the device cannot help, and silently clamping it upward would reintroduce the bug.

**Implementation.** The thresholds are already settings, so this needs no new
API surface:

1. Render the existing values into an env file via the normal template
mechanism, e.g. `/etc/zram.env`:
```
EVICTION_MEMORY_AVAILABLE_HARD={{#if settings.kubernetes.eviction-hard}}...{{/if}}
EVICTION_MEMORY_AVAILABLE_SOFT={{#if settings.kubernetes.eviction-soft}}...{{/if}}
```
2. Replace the literal `--size 1G` with a ~25-line shell helper that reads that
file plus `MemTotal` from `/proc/meminfo`, handles both `N%` and `NMi`/`NGi`
forms, and computes `Z`. No new Handlebars helper needed — the arithmetic
stays in shell, the template only passes the raw strings through.
3. Reorder the unit. This is the one real cost, and it is small:
```diff
-DefaultDependencies=no
-Wants=modprobe@zram.service
-After=modprobe@zram.service dev-zram0.device
+Wants=modprobe@zram.service
+After=modprobe@zram.service dev-zram0.device settings-applier.service

[Install]
-WantedBy=sysinit.target
+RequiredBy=preconfigured.target
```
`settings-applier.service` (`thar-be-settings --all`) is where every other
templated config is written, and it is `RequiredBy=preconfigured.target`,
while `kubelet.service` waits on `configured.target`. So the device would
still be online well before kubelet starts, in the same phase as the rest of
the settings-derived configuration. Nothing in early boot needs swap — the
non-swap variants have never had any at that point.

### The default case still needs a decision

Deriving from thresholds does not by itself fix a stock node, because a stock node has `T = 100 MiB` absolute, which yields `Z = 50 MiB` — effectively disabling the feature. That is a real tension and I don't think it can be papered over: **a meaningfully sized zram device is incompatible with kubelet's default 100 MiB absolute hard-eviction threshold.**

So I'd suggest pairing the change with a **default percentage threshold shipped alongside the swap feature**, chosen so the invariant holds by construction — e.g. defaulting `settings.kubernetes.eviction-hard` to `memory.available: 10%` when `release-swap` is in the variant. Then `Z = 5% · MemTotal` satisfies `Z ≤ 0.5·T` on every node, at every size, with no user configuration.

If shipping a default threshold is unpalatable, the fallback is a plain RAM-proportional size — `Z = min(5% · MemTotal, 4 GiB)` — which is still strictly better than a constant, but leaves stock hard eviction unreachable.

For prior art, systemd's `zram-generator` defaults to `min(RAM/2, 4096 MiB)` and Fedora has sized zram proportionally since F33/F34. As far as I can tell, no other Kubernetes node OS (AL2023, Flatcar, Talos, COS) enables zram by default at all.

### Option B — expose it as a setting

`settings.swap.zram-size`, accepting an absolute size or a percentage, with `"auto"` meaning the Option A derivation. More flexible, and now a smaller delta since Option A already introduces the template and the ordering change — but it needs a model type in `sources/models` and boot-only restart semantics. I'd suggest Option A as the default behaviour regardless, with B layered on later if there is demand.

### One caveat to be explicit about

`zramctl --size` is *uncompressed* capacity. At the 3–4× ratios we measured, a device of size `Z` costs only `Z/r` physically — but it hides up to the full `Z` from `workingSet`. All the math above is driven by the uncompressed figure, which is the conservative and correct one for eviction reasoning; the physical cost of the proposed sizes is 3–4× smaller again.

### Related

- #4075 — the de-facto thread for these symptoms; multiple reporters plus a
maintainer confirm `kswapd` saturating CPU
- #4735 / core-kit#590, core-kit#792 — where zram was introduced and defaulted
- #4808 — requests swap *files*; does not cover zram sizing
- kubernetes#43916 — the page-cache half of the accounting blind spot

Contributor guide

Open the contributing guide

Research direction

Start with prepare-swap.service and the kubelet-config template, then inspect settings-applier.service and the /proc/meminfo inputs described in the proposal. Trace how eviction-hard and eviction-soft values reach boot-time configuration and how zramctl is invoked. Done means an agreed sizing policy is implemented, boot ordering remains valid, and behavior is tested across configured thresholds and node memory sizes.

Written by the indexing model from the issue text.

Assessment

Tech stack
kubernetes, linux, shell
Domain
devops, operating-systems
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.