oxidecomputer / oxidecomputer/propolis
"rcu_preempt self-detected stall on CPU"
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 270
- Forks
- 42
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 6
Description
@JustinAzoff saw this on colo:
Jun 11 14:12:22 oxcolo kernel: rcu: INFO: rcu_preempt self-detected stall on CPU
Jun 11 14:12:22 oxcolo kernel: rcu: 0-....: (1 GPs behind) idle=3514/0/0x3 softirq=671080/671080 fqs=3762
Jun 11 14:12:22 oxcolo kernel: rcu: (t=21002 jiffies g=1292137 q=85 ncpus=2)
Jun 11 14:12:22 oxcolo kernel: CPU: 0 UID: 0 PID: 0 Comm: swapper/0 Not tainted 6.18.33 #1-NixOS PREEMPT(lazy)
Jun 11 14:12:22 oxcolo kernel: Hardware name: Oxide OxVM, BIOS v0.8 The Aftermath 30, 3185 YOLD
Jun 11 14:12:22 oxcolo kernel: RIP: 0010:handle_softirqs+0x8b/0x2a0
Jun 11 14:12:22 oxcolo kernel: Code: 00 89 5c 24 14 48 89 6c 24 08 c7 44 24 10 0a 00 00 00 44 89 6c 24 04 45 89 f5 31 c0 65 66 89 05 73 a4 7b 02 fb 0f 1f 44 00 00 <bb> ff ff ff ff 49 c7 c2 c0 80 40 b3 44 89 ed 41 0f bc dd 83 c3 01
Jun 11 14:12:22 oxcolo kernel: RSP: 0018:ffffd492c0003f90 EFLAGS: 00000246
Jun 11 14:12:22 oxcolo kernel: RAX: 0000000000000000 RBX: 0000000004200002 RCX: 0000000000000051
Jun 11 14:12:22 oxcolo kernel: RDX: ffff8e2783d41000 RSI: fffffffe6e487784 RDI: 0000000000000002
Jun 11 14:12:22 oxcolo kernel: RBP: 000000013e473d5e R08: 0000000000000002 R09: ffff8e2737c228c0
Jun 11 14:12:22 oxcolo kernel: R10: 0003b6906640e23d R11: 0000000000bf57d7 R12: 0000000000000000
Jun 11 14:12:22 oxcolo kernel: R13: 0000000000000080 R14: 0000000000000080 R15: 0000000000000000
Jun 11 14:12:22 oxcolo kernel: FS: 0000000000000000(0000) GS:ffff8e2783d41000(0000) knlGS:0000000000000000
Jun 11 14:12:22 oxcolo kernel: CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
Jun 11 14:12:22 oxcolo kernel: CR2: 000057f722a24ff8 CR3: 0000000102130000 CR4: 00000000003506f0
Jun 11 14:12:22 oxcolo kernel: Call Trace:
Jun 11 14:12:22 oxcolo kernel: <IRQ>
Jun 11 14:12:22 oxcolo kernel: __irq_exit_rcu+0xc8/0xf0
Jun 11 14:12:22 oxcolo kernel: sysvec_apic_timer_interrupt+0x73/0x80
Jun 11 14:12:22 oxcolo kernel: </IRQ>
Jun 11 14:12:22 oxcolo kernel: <TASK>
Jun 11 14:12:22 oxcolo kernel: asm_sysvec_apic_timer_interrupt+0x1a/0x20
Jun 11 14:12:22 oxcolo kernel: RIP: 0010:pv_native_safe_halt+0xf/0x20
Jun 11 14:12:22 oxcolo kernel: Code: 20 d0 e9 6f b4 06 ff 0f 1f 40 00 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 f3 0f 1e fa eb 07 0f 00 2d a3 9d 18 00 fb f4 <e9> 47 b4 06 ff 90 66 66 2e 0f 1f 84 00 00 00 00 00 90 90 90 90 90
Jun 11 14:12:22 oxcolo kernel: RSP: 0018:ffffffffb3403e70 EFLAGS: 00000292
Jun 11 14:12:22 oxcolo kernel: RAX: 0000000000000000 RBX: 0000000000000000 RCX: ffffffffb3412940
Jun 11 14:12:22 oxcolo kernel: RDX: 4000000000000000 RSI: 0000000000000000 RDI: 000000007a6034f4
Jun 11 14:12:22 oxcolo kernel: RBP: ffffffffb3412940 R08: 000000007a6034f4 R09: ffff8e2737c2f7c0
Jun 11 14:12:22 oxcolo kernel: R10: 00000000fffffffb R11: 0000000000000000 R12: 0000000000000000
Jun 11 14:12:22 oxcolo kernel: R13: 0000000000000000 R14: 000000000000000a R15: 00000000bdf89000
Jun 11 14:12:22 oxcolo kernel: default_idle+0x9/0x20
Jun 11 14:12:22 oxcolo kernel: default_idle_call+0x2a/0x100
Jun 11 14:12:22 oxcolo kernel: do_idle+0x23a/0x270
Jun 11 14:12:22 oxcolo kernel: cpu_startup_entry+0x29/0x30
Jun 11 14:12:22 oxcolo kernel: rest_init+0xcc/0xd0
Jun 11 14:12:22 oxcolo kernel: start_kernel+0x759/0x760
Jun 11 14:12:22 oxcolo kernel: x86_64_start_reservations+0x24/0x30
Jun 11 14:12:22 oxcolo kernel: x86_64_start_kernel+0xd2/0xe0
Jun 11 14:12:22 oxcolo kernel: common_startup_64+0x13e/0x141
Jun 11 14:12:22 oxcolo kernel: </TASK>
Jun 11 14:12:49 oxcolo kernel: watchdog: BUG: soft lockup - CPU#0 stuck for 49s! [swapper/0:0]
I'm not optimistic we can tell from this what or why Propolis was stalled: Linux set up watchdog_timer_fn to tick every ~1/5th (NSEC_PER_SEC / NUM_SAMPLE_PERIODS, out of kernel/watchdog.c) but the vCPU didn't get ticked for 49 seconds. what were we doing? it seems possible that the vCPU was off in Propolis (as it was in #1008) doing some other I/O, but certainly that should not get the CPU stuck for any number of milliseconds.
I'm opening this mostly as I'm not sure I fully understand what the circumstances are for a guest to produce this, and as a reference for a follow-up issue about how I'd like to at least narrow down who's off doing what for these kinds of problems.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the guest log in this issue and the watchdog_timer_fn timing described from kernel/watchdog.c. Compare the circumstances with issue #1008 and trace what Propolis was doing while the vCPU was not being ticked. Done means identifying or narrowing down the operation responsible for the stall and documenting the relevant circumstances.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- linux, rust
- Domain
- operating-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100