github / github/copilot-cli

Desktop app: UI thread hangs (AppHangB1) and Windows force-closes github.exe after repeated browser canvas use

未关闭
#4,872 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

triage
主要语言
Shell
星标
11.2k
派生
1.9k
平均合并
14 小时 16 分钟
30 天内合并 PR
6

描述

Describe the bug

The GitHub Copilot desktop app (github.exe 1.1.21.0) stopped responding and was force-closed by Windows. The evidence points at the browser preview / WebView2 subsystem blocking the UI thread, after repeatedly opening, navigating and reading browser canvases pointed at local HTTP servers.

This is a hang, not a crash — no access violation, no panic. Windows recorded:

Application Hang (Event ID 1002)
  The program github.exe version 1.1.21.0 stopped interacting with Windows and was closed.

Windows Error Reporting (Event ID 1001)
  Event Name:      AppHangB1
  P1: github.exe
  P2: 1.1.21.0
  P3: 6aa8d523
  Hang Signature:  6db1
  Hang Type:       134217728  (0x8000000)
  IsFatal:         1

The Rust backend was healthy right up to the moment it was killed. ~/.copilot/logs/github-app.<pid>.log (40 MB) ends ~5 seconds before the force-close, mid-way through entirely routine traffic:

11:18:33.703  DEBUG ws_handle_client_message kind="get_workspace_changes"
11:18:33.908  INFO  github_api_external_created_pr_lookup
11:18:33.909  DEBUG ...

No error, no panic, and no shutdown sequence — consistent with the process being killed from outside while the backend was still working normally. That pattern points at the UI thread specifically rather than at the backend.

Affected version
GitHub Copilot app 1.1.21.0 (build 2026/09/15:05:18:27!15938018)
Bundled CLI: copilot 0.0.395 (commit 4b4fe6e)
Steps to reproduce the behavior

Not reduced to a reliable repro, but the conditions present were:

  1. Open several browser canvases pointed at local servers (http://127.0.0.1:<port>/) started by a canvas extension.
  2. Repeatedly navigate_page / read_page / re-open_canvas against the same instance ids, with ~6 webviews live simultaneously.
  3. Reload the extension host, which replaces those local servers — the previously-opened canvases are now pointed at dead ports.
  4. Continue reading/navigating those canvases. Browser tool calls begin timing out.
  5. Shortly afterwards the UI stops responding entirely and Windows force-closes the app.

Step 3→4 (driving a browser canvas whose target port is no longer listening) looks like the most suspicious ingredient, since that is exactly when the failures started appearing.

Expected behavior

A browser preview pointed at a dead or unresponsive local port should fail that individual tool call and leave the application responsive. No preview operation should be able to block the UI thread to the point that Windows declares the process hung.

Additional context

Repeated profile cleanup failures preceded the hang:

Signal Count
WARN github_app::browser_preview: Failed to remove browser preview ephemeral profile directory 15
Orphaned profile dirs left in %TEMP%\github-app-browser-previews 3
Browser tool request timed out (in the ~2 min before the hang) 2
Webview ... was not found 1

The repeated cleanup failures suggest WebView2 processes lingering and holding file locks on their profile directories. A blocking WebView2/COM call on the UI thread would fit hang type 0x8000000 and matches the observed ordering: browser tool calls time out → a webview goes missing → UI stops responding.

Ruled out — websocket message volume. I checked whether a message storm was responsible. subscribe_workspace_changes per minute: 93 at 11:06, 66 at 11:09, 28 at 11:15, 17 at 11:16, and 26 in the final minute. Traffic at the end was below the earlier peak, so volume was not the trigger.

Ruled out — the canvas extension in use. A locally-developed canvas extension was under active development at the time, but it was not involved: every extension host reached === ready === with no exception, the host in use had no exit line at all (still running when the app died), it runs out-of-process under Node and cannot block the app's UI thread directly, and its one newly-added privileged code path was never invoked (zero log hits).

Not memory exhaustion. 13.7 GB free of 55.7 GB at the time. There was general pressure on the machine (Memory Compression at 4.8 GB, an unrelated WSL VM at 16 GB), but the app was not starved.

Suggested mitigations:

  • Bound browser preview tool calls with a timeout that cannot block the UI thread.
  • Tear down the WebView2 process and its ephemeral profile deterministically when a preview is replaced or its instance id is reused, rather than leaving both to a cleanup that is observably failing.
  • Treat repeated Failed to remove browser preview ephemeral profile directory as a signal to stop reusing the profile root.

Diagnostic artifacts available on request: the WER report (AppHang_github.exe_...), and the 40 MB backend log covering the hang.

Environment:

  • OS: Windows 11 Enterprise 10.0.26200
  • CPU architecture: AMD64
  • Memory: 55.7 GB total, 13.7 GB free at time of analysis

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

从 github_app::browser_preview 日志和报告的 browser tool 超时开始,然后追踪死掉的本地端口、WebView2 实例和临时配置文件清理是如何处理的。确定哪个操作可能阻塞 UI thread,以及重复使用的 preview instance IDs 是否会留下进程或配置文件。完成的标准是:死掉或无响应的 preview 能够独立失败,同时桌面应用仍保持响应,并且卡死场景可以复现,或其原因已被隔离。

由索引模型根据 Issue 内容生成。

评估

技术栈
rust
领域
desktop
Issue 类型
缺陷
难度
5/5
预计耗时
一周以上
活跃度
活跃
描述清晰度
需要澄清
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。