NVIDIA / NVIDIA/open-gpu-kernel-modules

NV_ERR_GPU_IS_LOST while alt tabbing between the game and KDE Plasma environment

Open
#776 2 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
C
Stars
17.4k
Forks
1.9k
PR merge metrics
No merged PRs in 30d

Description

NVIDIA Open GPU Kernel Modules Version

565.77

Please confirm this issue does not happen with the proprietary driver (of the same version). This issue tracker is only for bugs specific to the open kernel driver.
  • I confirm that this does not happen with the proprietary driver package.
Operating System and Version

ArchLinux

Kernel Release

6.12.10-arch1-1

Please confirm you are running a stable release kernel (e.g. not a -rc). We do not accept bug reports for unreleased kernels.
  • I am running on a stable kernel release.
Hardware: GPU

NVIDIA RTX A2000 8GB Laptop GPU

Describe the bug

While running RimWorld through prime-run and alt-tabbing between the game and Firefox/LibreOffice Calc running on KDE Plasma environment running on Intel GPU after tens of such switches my game freezes and dGPU dies.

On Firefox I from time to time I also get some temporary browser freezes while having the game running in the background.

Feb 02 16:44:36 staballoy kernel: NVRM: GPU at PCI:0000:01:00: GPU-12c5ff8f-0e23-23a4-6e1e-8931aee48997
Feb 02 16:44:36 staballoy kernel: NVRM: Xid (PCI:0000:01:00): 79, pid='<unknown>', name=<unknown>, GPU has fallen off the bus.
Feb 02 16:44:36 staballoy kernel: NVRM: GPU 0000:01:00.0: GPU has fallen off the bus.
Feb 02 16:44:36 staballoy kernel: NVRM: kgspRcAndNotifyAllChannels_IMPL: RC all channels for critical error 79.
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _threadNodeCheckTimeout: API_GPU_ATTACHED_SANITY_CHECK failed!
Feb 02 16:44:36 staballoy kernel: NVRM: _issueRpcAndWait: rpcSendMessage failed with status 0x0000000f for fn 78!
Feb 02 16:44:36 staballoy kernel: NVRM: nvCheckOkFailedNoLog: Check failed: GPU lost from the bus [NV_ERR_GPU_IS_LOST] (0x0000000F) returned from nvdEngineDumpCallbackHelper(pGpu, pPrbEnc, pNvDumpState, pEngineCallback) @ nv_debug_dump.c:274
Feb 02 16:44:36 staballoy kernel: NVRM: _issueRpcAndWait: rpcSendMessage failed with status 0x0000000f for fn 78!
Feb 02 16:44:36 staballoy kernel: NVRM: nvCheckOkFailedNoLog: Check failed: GPU lost from the bus [NV_ERR_GPU_IS_LOST] (0x0000000F) returned from nvdEngineDumpCallbackHelper(pGpu, pPrbEnc, pNvDumpState, pEngineCallback) @ nv_debug_dump.c:274
Feb 02 16:44:36 staballoy kernel: NVRM: _issueRpcAndWait: rpcSendMessage failed with status 0x0000000f for fn 78!
Feb 02 16:44:36 staballoy kernel: NVRM: nvCheckOkFailedNoLog: Check failed: GPU lost from the bus [NV_ERR_GPU_IS_LOST] (0x0000000F) returned from nvdEngineDumpCallbackHelper(pGpu, pPrbEnc, pNvDumpState, pEngineCallback) @ nv_debug_dump.c:274
Feb 02 16:44:36 staballoy kernel: NVRM: RmLogGpuCrash: RmLogGpuCrash: failed to save GPU crash data
Feb 02 16:44:36 staballoy kernel: NVRM: _kgspLogRpcSanityCheckFailure: GPU0 sanity check failed 0xf waiting for RPC response from GSP. Expected function 76 (GSP_RM_CONTROL) (0x2080a7d7 0x2).
Feb 02 16:44:36 staballoy kernel: NVRM: GPU0 GSP RPC buffer contains function 78 (DUMP_PROTOBUF_COMPONENT) and data 0x0000000000000000 0x0000000000000000.
Feb 02 16:44:36 staballoy kernel: NVRM: GPU0 RPC history (CPU -> GSP):
Feb 02 16:44:36 staballoy kernel: NVRM:     entry function                   data0              data1              ts_start           ts_end             duration actively_polling
Feb 02 16:44:36 staballoy kernel: NVRM:      0    76   GSP_RM_CONTROL        0x000000002080a7d7 0x0000000000000002 0x00062d2aa7262bdb 0x0000000000000000          y
Feb 02 16:44:36 staballoy kernel: NVRM:     -1    76   GSP_RM_CONTROL        0x000000002080a7d7 0x0000000000000002 0x00062d2aa6d9defb 0x00062d2aa6d9e03a    319us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -2    76   GSP_RM_CONTROL        0x000000002080a7d7 0x0000000000000002 0x00062d2aa68d9185 0x00062d2aa68d92dc    343us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -3    76   GSP_RM_CONTROL        0x000000002080a7d7 0x0000000000000002 0x00062d2aa6414422 0x00062d2aa6414558    310us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -4    76   GSP_RM_CONTROL        0x000000002080a7d7 0x0000000000000002 0x00062d2aa5f4f7a0 0x00062d2aa5f4f8c5    293us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -5    76   GSP_RM_CONTROL        0x000000002080a7d7 0x0000000000000002 0x00062d2aa5a8ab01 0x00062d2aa5a8ac23    290us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -6    76   GSP_RM_CONTROL        0x000000002080a7d7 0x0000000000000002 0x00062d2aa55c5e7c 0x00062d2aa55c5f9e    290us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -7    76   GSP_RM_CONTROL        0x000000002080a7d7 0x0000000000000002 0x00062d2aa5101102 0x00062d2aa5101318    534us  
Feb 02 16:44:36 staballoy kernel: NVRM: GPU0 RPC event history (CPU <- GSP):
Feb 02 16:44:36 staballoy kernel: NVRM:     entry function                   data0              data1              ts_start           ts_end             duration during_incomplete_rpc
Feb 02 16:44:36 staballoy kernel: NVRM:      0    4099 POST_EVENT            0x0000000000000021 0x0000000000000008 0x00062d2aa6987fe5 0x00062d2aa6988001     28us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -1    4099 POST_EVENT            0x0000000000000021 0x0000000000000020 0x00062d2aa691d80a 0x00062d2aa691d821     23us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -2    4099 POST_EVENT            0x0000000000000021 0x0000000000000008 0x00062d2a00d45faa 0x00062d2a00d45faf      5us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -3    4099 POST_EVENT            0x0000000000000021 0x0000000000000020 0x00062d2a00d154d4 0x00062d2a00d154da      6us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -4    4099 POST_EVENT            0x0000000000000021 0x0000000000000008 0x00062d2a00b2d0d2 0x00062d2a00b2d0d8      6us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -5    4099 POST_EVENT            0x0000000000000021 0x0000000000000020 0x00062d2a00afc544 0x00062d2a00afc54b      7us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -6    4099 POST_EVENT            0x0000000000000021 0x0000000000000008 0x00062d2a00944ea2 0x00062d2a00944ea7      5us  
Feb 02 16:44:36 staballoy kernel: NVRM:     -7    4099 POST_EVENT            0x0000000000000021 0x0000000000000020 0x00062d2a009142a5 0x00062d2a009142aa      5us  
Feb 02 16:44:36 staballoy kernel: NVRM: RmCheckForGcxSupportOnCurrentState: NVRM, Failed to get GCx pre-requisite, status=0xf
Feb 02 16:44:36 staballoy kernel: NVRM: Xid (PCI:0000:01:00): 154, pid='<unknown>', name=<unknown>, GPU recovery action changed from 0x0 (None) to 0x1 (GPU Reset Required)
Feb 02 16:44:39 staballoy kernel: [drm:__nv_drm_semsurf_wait_fence_work_cb [nvidia_drm]] *ERROR* [nvidia-drm] [GPU ID 0x00000100] Failed to register auto-value-update on pre-wait value for sync FD semaphore surface
Feb 02 16:44:41 staballoy kernel: NVRM: Error in service of callback 
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertOkFailedNoLog: Assertion failed: Current device is not valid [NV_ERR_INVALID_DEVICE] (0x00000026) returned from rmDeviceGpuLocksAcquire(pGpu, GPUS_LOCK_FLAGS_NONE, RM_LOCK_MODULES_MEM_PMA) @ virtual_mem.c:115
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertFailedNoLog: Assertion failed: status == NV_OK @ rs_client.c:843
Feb 02 16:44:49 staballoy kernel: NVRM: nvAssertFailedNoLog: Assertion failed: status == NV_OK @ rs_server.c:258
To Reproduce
  • have your desktop environment run on Intel iGPU,
  • run RimWorld through prime-run,
  • play the game,
  • change the active application couple of times,
  • at some point after changing back to the game, game will freeze (-1 crash was either after alt-tabbing back from Firefox or just while playing, current one was from either Firefox or LibreOffice Calc),
  • observe GPU has fallen off the bus and NV_ERR_GPU_IS_LOST in system logs
Bug Incidence

Always

nvidia-bug-report.log.gz

nvidia-bug-report.log.gz

More Info

I have these two set in my kernel cmdline: nvidia_drm.modeset=1 nvidia_drm.fbdev=1. My Xorg process running SDDM display manager runs on Nvidia dGPU according to nvidia-smi, and my KDE Plasma desktop environment is running on Intel iGPU.

# cat /etc/modprobe.d/nvidia-pm.conf 
options nvidia "NVreg_DynamicPowerManagement=0x02"

# cat /etc/modprobe.d/nvidia.conf 
options nvidia NVreg_PreserveVideoMemoryAllocations=0
options nvidia NVreg_TemporaryFilePath=/var/tmp
options nvidia NVreg_UsePageAttributeTable=1

During all of that, notebook Dell Precision 5770 is connected to the official 130W charger and external 4K 60Hz LG 32UD99 monitor through USB-C cable, notebook itself has lid closed and is set aside. It doesn't have any kind of thermal issues according to temperature monitoring tools.

I have issues with my GPU being limited to 30W (checked the maximum value with nvidia-smi) while running the game for a minute or so while it was previously listing maximum value as 60W.

I'm also using throttled to fix this notebook issues that limit my CPU frequencies after couple of seconds of high load or limit them to minimum after resuming from sleep. Its possible that its firmware has some BDPROCHOT clearing issues.

I will try reproducing the issue with and edit to add the details whether it also fails or not:

  1. closed source driver + throttled enabled, EDIT: Same issue, GPU has fallen off the bus., happened during browsing with Firefox and having game in the background, noticed by running nvidia-smi that failed with Unable to determine the device handle for GPU0000:01:00.0: Unknown Error
  2. closed source driver + throttled disabled, EDIT: Cannot reproduce,
  3. (if pt. 2 also reproduces the issue) open source driver + throttled disabled, EDIT: Skipping.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No source file or test is identified. Start with the provided nvidia-bug-report.log.gz and the kernel Xid 79/NV_ERR_GPU_IS_LOST messages, then reproduce the Intel-iGPU/KDE, prime-run and alt-tab sequence while completing the proprietary-driver comparison. Done requires an identified open-driver cause and validation that the GPU no longer falls off the bus.

Written by the indexing model from the issue text.

Assessment

Tech stack
c, linux
Domain
operating-systems
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.