github / github/copilot-sdk

[.NET 1.0.6][FFI] CreateSessionAsync can hang after successful startup in Linux/Kubernetes

未关闭
#1,958 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
bug
主要语言
Java
星标
10.5k
派生
1.5k
平均合并
1 天 14 小时
30 天内合并 PR
129

描述

## Summary

With the .NET SDK 1.0.6 in-process transport (`RuntimeConnection.ForInProcess()`), `CopilotClient.StartAsync()` succeeds, but a later `CreateSessionAsync()` can remain incomplete indefinitely in a Linux container running on Kubernetes.

The same application/configuration works locally, and changing only the replacement transport from FFI to the default stdio child-process transport recovers session creation.

## Environment

- GitHub Copilot SDK for .NET: 1.0.6
- Runtime: .NET 8
- Host: Linux container on Kubernetes
- Client mode: `CopilotClientMode.Empty`
- Transport: `RuntimeConnection.ForInProcess()`
- Session uses a custom provider, tools/hooks, and `SessionFs`

## Observed behavior

1. Create the client with `RuntimeConnection.ForInProcess()`.
2. `StartAsync()` completes successfully and the transport reports setup complete.
3. Call `CreateSessionAsync(sessionConfig)`.
4. The returned task never completes or faults. The API has no cancellation token, so the caller cannot stop it.
5. Recreating another client with the same FFI transport can reproduce the wedge.
6. Recreating the client with `Connection = null` (stdio) allows the same session configuration to succeed.

Representative host-side timeout logs:

```text
CreateSessionAsync did not return within 120000ms; abandoning session creation. The SDK transport is likely wedged.

Create failed with a dead/wedged transport; recreating the Copilot client and retrying once.
```

We also observed that disposal of sessions/client associated with the wedged FFI transport may not complete promptly, so recovery must not depend on unbounded `DisposeAsync()` calls.

## Expected behavior

`CreateSessionAsync()` should either:

- complete successfully,
- fail with a transport/runtime exception, or
- accept cancellation/a timeout so the host can recover.

If the in-process runtime becomes unhealthy after startup, recreating the client should not recreate the same permanently wedged runtime state.

## Current workaround

We keep FFI as the preferred transport, but on the first post-start session-create/send transport timeout:

1. mark the FFI transport unhealthy,
2. detach sessions associated with that client generation,
3. bound/abandon cleanup of the wedged runtime,
4. recreate the client with stdio, and
5. retry once with a rebuilt session configuration.

## Related issue

#1934 tracks options/environment parity for the in-process transport. This appears distinct: startup succeeds, but the FFI transport later stops completing session operations.

Please let me know what additional native/runtime logging or a smaller container repro would be most useful.

贡献指南

打开贡献指南

调研方向

首先在 Linux/Kubernetes 容器中复现 .NET SDK 1.0.6 的情况,使用 RuntimeConnection.ForInProcess()、StartAsync() 和 CreateSessionAsync();将其与默认的 stdio 传输进行比较。跟踪启动后的 in-process 传输,并捕获会话创建和释放过程附近的 native/runtime 日志。完成的标准是操作能够完成或报告有界失败,并且有记录在案的恢复路径,且该路径不依赖于无界清理。

由索引模型根据 Issue 内容生成。

评估

技术栈
csharp, kubernetes, linux
领域
backend, infrastructure
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
冷清
描述清晰度
基本清楚
新手友好度
38/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。