intel / intel/ethernet-linux-ice
ice: PF MSI-X / queue oversubscription causes "Affinity broken due to vector space exhaustion" on systems with a narrow IRQ-affinity pool
- Dominant language
- C
- Stars
- 60
- Forks
- 21
- PR merge metrics
- No merged PRs in 30d
Description
## Description
On systems with a large CPU count and a narrow IRQ-affinity pool
(e.g. PREEMPT_RT controllers using `isolcpus=managed_irq` to keep most
cores RT-isolated), the `ice` driver requests `MAX_DEFAULT_VECTORS=64`
MSI-X vectors and 64 Tx/Rx queues per PF at probe time, regardless of
how many CPUs are actually allowed to service IRQs.
When several PFs probe in a row, the resulting demand exhausts the
per-CPU MSI-X vector budget inside the IRQ-affinity mask. The kernel
emits
irq : Affinity broken due to vector space exhaustion.
and `__assign_irq_vector()` falls back to `cpu_online_mask`, silently
placing ice IRQs on **isolated / RT-reserved cores** or on the
housekeeping core. That fallback is not a cosmetic warning - it
breaks the isolation contract for any RT/NFV workload running on the
isolated set, because production network IRQs end up on those cores.
The `ethtool -L combined N` udev workaround mitigates the symptom but
not the cause, because it runs **after** MSI-X allocation has already
happened: the first PF probed still has its vectors trapped on the
fallback CPU even after the cap is applied.
## Environment
- Driver version: ice 2.6.6 (OOT) — also reproduces on 2.3.10.x and
the in-tree driver as of v6.12.
- Kernel version: 6.12.0-1-rt-amd64 (PREEMPT_RT)
- Hardware: Intel Xeon 6776P-B (Granite Rapids-D), 144 LCPU,
Dell PowerEdge XR8720t
- NICs: 24 ice PFs across 4 PCI roots
(E825-C and E830-CC, mix of 4- and 8-port adapters)
- SR-IOV: 2 VFs / PF (~49 VFs total), 16 queue pairs each
- Boot cmdline (relevant bits):
```
irqaffinity=69-71,141-143 # 6 IRQ CPUs
kthread_cpus=0-1,72-73 # 4 housekeeping CPUs
isolcpus=nohz,domain,managed_irq,2-68,74-140
nohz_full=2-68,74-140
rcu_nocbs=2-71,74-143
intel_iommu=on iommu=pt
```
## Steps to reproduce
1. Boot a host with many ice PFs (e.g. 24) on a CPU with high logical
core count (e.g. 144 LCPU).
2. Boot with `irqaffinity=` (e.g. 6 CPUs) and
`isolcpus=managed_irq,`.
3. Let ice probe all PFs and enable SR-IOV (2 VFs / PF is enough to
trigger it on a 24-PF system; more VFs reproduce sooner).
4. Watch `dmesg`:
```
irq 2587: Affinity broken due to vector space exhaustion.
irq 2586: Affinity broken due to vector space exhaustion.
...
```
5. Inspect `/proc/irq//effective_affinity_list` for the ice TxRx
IRQs - many land on CPUs **outside** the `irqaffinity=` set,
including on `isolcpus`-isolated CPUs.
## Expected behavior
- The driver should size its per-PF MSI-X reservation to a value that
the housekeeping IRQ pool can actually accommodate, or at least
expose a knob so the operator can do so before any vector is
allocated.
- No ice IRQ should land on an `isolcpus=managed_irq` CPU.
## Actual behavior
On the reference 144-LCPU box with 24 PFs and `irqaffinity=69-71,141-143`
(6 CPUs ≈ 1140 free vectors), one boot snapshot shows:
| Metric | Value |
|----------------------------------------------|-------:|
| Total ice/iavf MSI-X IRQs | 1 928 |
| → on IRQ-pool CPUs `{69-71,141-143}` | 1 159 |
| → on housekeeping CPUs `{0,1}` | 199 |
| → **on isolated CPUs (managed_irq, 2-68/74-140)** | **570** |
| Per-PF default queue count | 64 |
| Per-CPU vector demand on the 6-CPU pool | ~321 |
| Per-CPU MSI-X budget (IR enabled, x86) | ~190 |
So the housekeeping pool is oversubscribed ~1.7× and the driver loses
570 IRQs onto isolated cores. Affected PFs include 6 of the 24 (e.g.
`enp35s0f{2..7}`) which end up with their 64 Tx/Rx IRQs fanned out
linearly across 60+ isolated cores each; the first PF probed
(`enp19s0f0`) has all 64 queues trapped on the housekeeping CPU 1.
After the `ethtool -L combined 6` udev workaround:
| Metric | Value |
|----------------------------------------------|------:|
| Total ice/iavf MSI-X IRQs | 536 |
| → on IRQ-pool CPUs | 529 |
| → on housekeeping CPUs | 7 |
| → **on isolated CPUs** | **0** |
i.e. the cap works *but only because* it reduces the post-probe
queue count; the MSI-X vectors of the first PF stay stuck on the
housekeeping core because `ethtool -L combined` does not migrate the
surviving vectors.
## Root-cause analysis
Both call sites that decide the per-PF MSI-X / queue count consult the
same helper, `ice_normalize_cpu_count()` in `src/ice_lib.c`:
```c
int ice_normalize_cpu_count(int num_cpus)
{
if (num_cpus > MAX_DEFAULT_VECTORS)
num_cpus = MAX_DEFAULT_VECTORS; /* clamp at 64 */
else if (num_cpus < MIN_DEFAULT_VECTORS)
num_cpus = MIN_DEFAULT_VECTORS; /* floor at 8 */
return num_cpus;
}
```
It is used by:
- `ice_ena_msix_range()` in `src/ice_irq.c` — drives `pf->msix.eth`
(the actual MSI-X budget requested via `pci_enable_msix_*`);
- `ice_vsi_set_num_qs()` in `src/ice_lib.c` — drives
`vsi->alloc_txq`, `vsi->alloc_rxq` and `vsi->num_q_vectors` for
ICE_VSI_PF.
The clamp considers the number of local CPUs only - it does not
consult `housekeeping_cpumask()` / the `irqaffinity=` mask, so on a
big-core RT box it will happily request 64 vectors per PF even when
the housekeeping pool has room for only a few hundred vectors total.
## Suggested fix
Add a single module parameter, `max_default_qps`, that further clamps
the value returned by `ice_normalize_cpu_count()`. Because that helper
is the *only* code path used by both the MSI-X budget calculation and
the queue-count calculation, one knob fixes both demand and queue
count atomically, without touching call sites.
```c
/* operator-tunable cap, defined in ice_main.c, declared in ice_lib.h */
int ice_max_default_qps = -1;
module_param_named(max_default_qps, ice_max_default_qps, int, 0644);
MODULE_PARM_DESC(max_default_qps,
"Cap per-PF default Tx/Rx queues and MSI-X vectors (1..64; <=0 disables, default -1)");
int ice_normalize_cpu_count(int num_cpus)
{
int upper = MAX_DEFAULT_VECTORS;
if (ice_max_default_qps > 0 && ice_max_default_qps < upper)
upper = ice_max_default_qps;
if (num_cpus > upper)
num_cpus = upper;
else if (num_cpus < MIN_DEFAULT_VECTORS &&
upper >= MIN_DEFAULT_VECTORS)
num_cpus = MIN_DEFAULT_VECTORS;
return num_cpus;
}
```
Default value `-1` preserves the existing behaviour exactly. With
`max_default_qps=6` on the reference 144-LCPU controller, the
"Affinity broken" warnings disappear from the boot log entirely and
every ice IRQ lands inside the `irqaffinity=` mask from the very first
probe - **no `ethtool -L` workaround needed**.
## Workaround
Until a kernel-side or upstream-driver fix is in, two options:
1. **Userspace cap (already deployed)**:
```
# /etc/udev/rules.d/91-ice-queues.rules
ACTION=="add", SUBSYSTEM=="net", DRIVERS=="ice", \
ATTR{device/sriov_totalvfs}!="", \
RUN+="/usr/sbin/ethtool -L $name combined "
```
Caveat: only fully effective for PFs probed *after* the rule has
loaded; the first PF probed typically keeps its surviving vectors
on the housekeeping CPU.
2. **Module param (this issue / attached patch)**:
```
options ice max_default_qps=
```
Applied at probe time, no residual.
Attached: `0001-ice-add-max_default_qps-module-param-to-cap-pf-queues.patch`
[0001-ice-add-max_default_qps-module-param-to-cap-pf-queues.patch](https://github.com/user-attachments/files/29957642/0001-ice-add-max_default_qps-module-param-to-cap-pf-queues.patch)
Contributor guide
Research direction
Start with ice_normalize_cpu_count() in src/ice_lib.c, then trace its use from ice_ena_msix_range() in src/ice_irq.c and the PF queue setup in ice_vsi_set_num_qs(). Review the attached patch alongside ice_main.c and ice_lib.h, then reproduce with the documented multi-PF configuration and inspect dmesg plus /proc/irq/*/effective_affinity_list. Done means the cap applies before probe, avoids affinity warnings, and keeps ice IRQs within the configured IRQ pool.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- c, linux
- Domain
- networking, operating-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Clearly specified
- Newbie friendliness
- 55/100