OTel: parallel task dispatch emits successful chat spans without response identity or usage
还没有人认领这个 Issue。
- 主要语言
- Shell
- 星标
- 11.2k
- 派生
- 1.9k
- 平均合并
- 14 小时 16 分钟
- 30 天内合并 PR
- 6
描述
Describe the bug
With metadata-only OpenTelemetry file export enabled, a prompt-mode run that dispatches two named subagents emits two chat spans that have successful status and a stop finish reason but omit all of:
gen_ai.response.idgen_ai.response.modelgen_ai.usage.input_tokensgen_ai.usage.output_tokens
This also occurs when the parent model is explicitly pinned, so it is not limited to --model auto.
The remaining parent and subagent chat spans in the same trace contain complete request/response model and usage fields.
Affected version
GitHub Copilot CLI 1.0.83 on macOS.
Steps to reproduce
- Enable the file exporter and force content capture off.
- Run prompt mode with streaming enabled and an explicitly pinned parent model.
- In one synthetic prompt, use the
tasktool twice with two named agent types and explicit models. - Parse the resulting JSONL.
A representative trace contained one root invoke_agent, two execute_tool task spans, two nested named invoke_agent spans, and seven chat spans. Every completed-response chat had the expected provider, request model, response model, response ID, and usage. Two root-level chats had request model, streaming=true, status code 0, and finish reason stop, but no response identity or usage.
Expected behavior
Each logical completed inference should contain independently reported response identity and usage when available. An abandoned, cancelled, routing, or retry attempt should not look like a successful completed chat: it should carry an explicit attempt/outcome classification or a non-success terminal status. Consumers must not have to synthesize gen_ai.response.model from the requested model.
Privacy
The reproduction used synthetic prompts with OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=false. No prompt, response, repository, or credential content is included here.
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
首先,使用仅导出元数据的 OpenTelemetry、streaming、一个显式固定的父模型和两个命名的任务 agent,复现 prompt-mode trace。解析生成的 JSONL,并将两个 root-level 的不完整 chat span 与完整的 chat span 进行比较。完成的标准是:每个已完成的 inference 在可用时都报告 response identity 和 usage,而放弃的或未完成的尝试不会被标记为成功。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- shell
- 领域
- observability
- Issue 类型
- 缺陷
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 活跃
- 描述清晰度
- 基本清楚
- 新手友好度
- 52/100