NVIDIA / NVIDIA/open-gpu-kernel-modules
NV_ERR_GPU_IS_LOST while alt tabbing between the game and KDE Plasma environment
Nobody has claimed this yet.
- Dominant language
- C
- Stars
- 17.4k
- Forks
- 1.9k
- PR merge metrics
- No merged PRs in 30d
Description
NVIDIA Open GPU Kernel Modules Version
565.77
Please confirm this issue does not happen with the proprietary driver (of the same version). This issue tracker is only for bugs specific to the open kernel driver.
- I confirm that this does not happen with the proprietary driver package.
Operating System and Version
ArchLinux
Kernel Release
6.12.10-arch1-1
Please confirm you are running a stable release kernel (e.g. not a -rc). We do not accept bug reports for unreleased kernels.
- I am running on a stable kernel release.
Hardware: GPU
NVIDIA RTX A2000 8GB Laptop GPU
Describe the bug
While running RimWorld through prime-run and alt-tabbing between the game and Firefox/LibreOffice Calc running on KDE Plasma environment running on Intel GPU after tens of such switches my game freezes and dGPU dies.
On Firefox I from time to time I also get some temporary browser freezes while having the game running in the background.
Feb 02 16:44:36 staballoy kernel: NVRM: GPU at PCI:0000:01:00: GPU-12c5ff8f-0e23-23a4-6e1e-8931aee48997
Feb 02 16:44:36 staballoy kernel: NVRM: Xid (PCI:0000:01:00): 79, pid='<unknown>', name=<unknown>, GPU has fallen off the bus.
Feb 02 16:44:36 staballoy kernel: NVRM: GPU 0000:01:00.0: GPU has fallen off the bus.
Feb 02 16:44:36 staballoy kernel: NVRM: kgspRcAndNotifyAllChannels_IMPL: RC all channels for critical error 79.
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _issueRpcAndWait: rpcSendMessage failed with status 0x0000000f for fn 78!
Feb 02 16:44:36 staballoy kernel: NVRM: nvCheckOkFailedNoLog: Check failed: GPU lost from the bus [NV_ERR_GPU_IS_LOST] (0x0000000F) returned from nvdEngineDumpCallbackHelper(pGpu, pPrbEnc, pNvDumpState, pEngineCallback) @ nv_debug_dump.c:274
Feb 02 16:44:36 staballoy kernel: NVRM: _issueRpcAndWait: rpcSendMessage failed with status 0x0000000f for fn 78!
Feb 02 16:44:36 staballoy kernel: NVRM: nvCheckOkFailedNoLog: Check failed: GPU lost from the bus [NV_ERR_GPU_IS_LOST] (0x0000000F) returned from nvdEngineDumpCallbackHelper(pGpu, pPrbEnc, pNvDumpState, pEngineCallback) @ nv_debug_dump.c:274
Feb 02 16:44:36 staballoy kernel: NVRM: _issueRpcAndWait: rpcSendMessage failed with status 0x0000000f for fn 78!
Feb 02 16:44:36 staballoy kernel: NVRM: nvCheckOkFailedNoLog: Check failed: GPU lost from the bus [NV_ERR_GPU_IS_LOST] (0x0000000F) returned from nvdEngineDumpCallbackHelper(pGpu, pPrbEnc, pNvDumpState, pEngineCallback) @ nv_debug_dump.c:274
Feb 02 16:44:36 staballoy kernel: NVRM: RmLogGpuCrash: RmLogGpuCrash: failed to save GPU crash data
Feb 02 16:44:36 staballoy kernel: NVRM: _kgspLogRpcSanityCheckFailure: GPU0 sanity check failed 0xf waiting for RPC response from GSP. Expected function 76 (GSP_RM_CONTROL) (0x2080a7d7 0x2).
Feb 02 16:44:36 staballoy kernel: NVRM: GPU0 GSP RPC buffer contains function 78 (DUMP_PROTOBUF_COMPONENT) and data 0x0000000000000000 0x0000000000000000.
Feb 02 16:44:36 staballoy kernel: NVRM: GPU0 RPC history (CPU -> GSP):
Feb 02 16:44:36 staballoy kernel: NVRM: entry function data0 data1 ts_start ts_end duration actively_polling
Feb 02 16:44:36 staballoy kernel: NVRM: 0 76 GSP_RM_CONTROL 0x000000002080a7d7 0x0000000000000002 0x00062d2aa7262bdb 0x0000000000000000 y
Feb 02 16:44:36 staballoy kernel: NVRM: -1 76 GSP_RM_CONTROL 0x000000002080a7d7 0x0000000000000002 0x00062d2aa6d9defb 0x00062d2aa6d9e03a 319us
Feb 02 16:44:36 staballoy kernel: NVRM: -2 76 GSP_RM_CONTROL 0x000000002080a7d7 0x0000000000000002 0x00062d2aa68d9185 0x00062d2aa68d92dc 343us
Feb 02 16:44:36 staballoy kernel: NVRM: -3 76 GSP_RM_CONTROL 0x000000002080a7d7 0x0000000000000002 0x00062d2aa6414422 0x00062d2aa6414558 310us
Feb 02 16:44:36 staballoy kernel: NVRM: -4 76 GSP_RM_CONTROL 0x000000002080a7d7 0x0000000000000002 0x00062d2aa5f4f7a0 0x00062d2aa5f4f8c5 293us
Feb 02 16:44:36 staballoy kernel: NVRM: -5 76 GSP_RM_CONTROL 0x000000002080a7d7 0x0000000000000002 0x00062d2aa5a8ab01 0x00062d2aa5a8ac23 290us
Feb 02 16:44:36 staballoy kernel: NVRM: -6 76 GSP_RM_CONTROL 0x000000002080a7d7 0x0000000000000002 0x00062d2aa55c5e7c 0x00062d2aa55c5f9e 290us
Feb 02 16:44:36 staballoy kernel: NVRM: -7 76 GSP_RM_CONTROL 0x000000002080a7d7 0x0000000000000002 0x00062d2aa5101102 0x00062d2aa5101318 534us
Feb 02 16:44:36 staballoy kernel: NVRM: GPU0 RPC event history (CPU <- GSP):
Feb 02 16:44:36 staballoy kernel: NVRM: entry function data0 data1 ts_start ts_end duration during_incomplete_rpc
Feb 02 16:44:36 staballoy kernel: NVRM: 0 4099 POST_EVENT 0x0000000000000021 0x0000000000000008 0x00062d2aa6987fe5 0x00062d2aa6988001 28us
Feb 02 16:44:36 staballoy kernel: NVRM: -1 4099 POST_EVENT 0x0000000000000021 0x0000000000000020 0x00062d2aa691d80a 0x00062d2aa691d821 23us
Feb 02 16:44:36 staballoy kernel: NVRM: -2 4099 POST_EVENT 0x0000000000000021 0x0000000000000008 0x00062d2a00d45faa 0x00062d2a00d45faf 5us
Feb 02 16:44:36 staballoy kernel: NVRM: -3 4099 POST_EVENT 0x0000000000000021 0x0000000000000020 0x00062d2a00d154d4 0x00062d2a00d154da 6us
Feb 02 16:44:36 staballoy kernel: NVRM: -4 4099 POST_EVENT 0x0000000000000021 0x0000000000000008 0x00062d2a00b2d0d2 0x00062d2a00b2d0d8 6us
Feb 02 16:44:36 staballoy kernel: NVRM: -5 4099 POST_EVENT 0x0000000000000021 0x0000000000000020 0x00062d2a00afc544 0x00062d2a00afc54b 7us
Feb 02 16:44:36 staballoy kernel: NVRM: -6 4099 POST_EVENT 0x0000000000000021 0x0000000000000008 0x00062d2a00944ea2 0x00062d2a00944ea7 5us
Feb 02 16:44:36 staballoy kernel: NVRM: -7 4099 POST_EVENT 0x0000000000000021 0x0000000000000020 0x00062d2a009142a5 0x00062d2a009142aa 5us
Feb 02 16:44:36 staballoy kernel: NVRM: RmCheckForGcxSupportOnCurrentState: NVRM, Failed to get GCx pre-requisite, status=0xf
Feb 02 16:44:36 staballoy kernel: NVRM: Xid (PCI:0000:01:00): 154, pid='<unknown>', name=<unknown>, GPU recovery action changed from 0x0 (None) to 0x1 (GPU Reset Required)
Feb 02 16:44:39 staballoy kernel: [drm:__nv_drm_semsurf_wait_fence_work_cb [nvidia_drm]] *ERROR* [nvidia-drm] [GPU ID 0x00000100] Failed to register auto-value-update on pre-wait value for sync FD semaphore surface
Feb 02 16:44:41 staballoy kernel: NVRM: Error in service of callback
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertFailedNoLog: Assertion failed: status == NV_OK @ rs_client.c:843
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertFailedNoLog: Assertion failed: status == NV_OK @ rs_server.c:258
To Reproduce
- have your desktop environment run on Intel iGPU,
- run RimWorld through prime-run,
- play the game,
- change the active application couple of times,
- at some point after changing back to the game, game will freeze (-1 crash was either after alt-tabbing back from Firefox or just while playing, current one was from either Firefox or LibreOffice Calc),
- observe
GPU has fallen off the busandNV_ERR_GPU_IS_LOSTin system logs
Bug Incidence
Always
nvidia-bug-report.log.gz
More Info
I have these two set in my kernel cmdline: nvidia_drm.modeset=1 nvidia_drm.fbdev=1. My Xorg process running SDDM display manager runs on Nvidia dGPU according to nvidia-smi, and my KDE Plasma desktop environment is running on Intel iGPU.
# cat /etc/modprobe.d/nvidia-pm.conf
options nvidia "NVreg_DynamicPowerManagement=0x02"
# cat /etc/modprobe.d/nvidia.conf
options nvidia NVreg_PreserveVideoMemoryAllocations=0
options nvidia NVreg_TemporaryFilePath=/var/tmp
options nvidia NVreg_UsePageAttributeTable=1
During all of that, notebook Dell Precision 5770 is connected to the official 130W charger and external 4K 60Hz LG 32UD99 monitor through USB-C cable, notebook itself has lid closed and is set aside. It doesn't have any kind of thermal issues according to temperature monitoring tools.
I have issues with my GPU being limited to 30W (checked the maximum value with nvidia-smi) while running the game for a minute or so while it was previously listing maximum value as 60W.
I'm also using throttled to fix this notebook issues that limit my CPU frequencies after couple of seconds of high load or limit them to minimum after resuming from sleep. Its possible that its firmware has some BDPROCHOT clearing issues.
I will try reproducing the issue with and edit to add the details whether it also fails or not:
- closed source driver + throttled enabled, EDIT: Same issue,
GPU has fallen off the bus., happened during browsing with Firefox and having game in the background, noticed by runningnvidia-smithat failed withUnable to determine the device handle for GPU0000:01:00.0: Unknown Error - closed source driver + throttled disabled, EDIT: Cannot reproduce,
- (if pt. 2 also reproduces the issue) open source driver + throttled disabled, EDIT: Skipping.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No source file or test is identified. Start with the provided nvidia-bug-report.log.gz and the kernel Xid 79/NV_ERR_GPU_IS_LOST messages, then reproduce the Intel-iGPU/KDE, prime-run and alt-tab sequence while completing the proprietary-driver comparison. Done requires an identified open-driver cause and validation that the GPU no longer falls off the bus.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- c, linux
- Domain
- operating-systems
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100