aws / aws/bedrock-agentcore-sdk-python
feat: Evaluation Client — Lifecycle, Orchestration & Online Pipeline
- 主要语言
- Python
- 星标
- 761
- 派生
- 147
- 平均合并
- 1 天 23 小时
- 30 天内合并 PR
- 7
描述
## Problem
The SDK's `EvaluationClient` only exposes `run()`. On the control plane side, customers cannot programmatically create custom evaluators (LLM-as-a-judge configs), list available evaluators, update or delete evaluators, or manage online evaluation configs for continuous evaluation on live traffic — evaluator provisioning requires the console. On the data plane side, the starter toolkit's `EvaluationProcessor` provides significantly richer orchestration than `run()`: it fetches session data from CloudWatch independently, groups evaluators by level (SESSION vs TRACE), determines which spans to send based on evaluator level, and runs multiple evaluators with per-evaluator error handling. The toolkit also provides input validation, IAM role cleanup on delete, and typed config/result models.
## Acceptance Criteria
- [ ] Customers can create, get, list, update, and delete custom evaluators
- [ ] Customers can create, get, list, update, and delete online evaluation configs
- [ ] Online evaluation config supports enable/disable toggling and sampling rate adjustment
- [ ] Typed result models with error introspection (`has_error()`, `get_successful_results()`)
- [ ] Customers can fetch session trace data (spans + runtime logs) from CloudWatch for a given session and agent
- [ ] Customers can find the most recent session for an agent
- [ ] Multi-evaluator orchestration groups evaluators by level and selects appropriate spans per level
- [ ] Per-evaluator error handling — failures on one evaluator don't block others
- [ ] Online evaluation config deletion supports optional IAM execution role cleanup
- [ ] All functionality is verified via integration tests running in CI
## Relevant Links
- [`EvaluationControlPlaneClient`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/control_plane_client.py#L21)
- [`create_evaluator()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/control_plane_client.py#L107)
- [`create_online_evaluation_config()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/control_plane_client.py#L175)
- [`update_online_evaluation_config()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/control_plane_client.py#L312)
- [`EvaluationResult` / `EvaluationResults`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/models.py#L81)
- [`OnlineEvaluationConfig`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/models.py#L190)
- [`EvaluationProcessor`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/on_demand_processor.py#L27)
- [`evaluate_session()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/on_demand_processor.py#L418)
- [`fetch_session_data()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/on_demand_processor.py#L88)
- [`determine_spans_for_evaluator()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/on_demand_processor.py#L304)
- [`execute_evaluators()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/on_demand_processor.py#L343)
- [`EvaluationDataPlaneClient`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/data_plane_client.py#L20)
- [`delete_online_evaluation_config()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/online_processor.py#L200)
贡献指南
调研方向
先从 operations/evaluation/control_plane_client.py 和 models.py 开始,然后阅读 on_demand_processor.py 中的 EvaluationProcessor、EvaluationDataPlaneClient,以及 online_processor.py 中的 delete_online_evaluation_config()。首先追踪相关的入口点和现有的集成测试设置。完成意味着已覆盖所列出的 evaluator、在线配置、会话数据、编排、错误结果、清理以及 CI 集成测试的所有标准。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- aws, python
- 领域
- backend-api-design, cloud, testing
- Issue 类型
- 功能
- 难度
- 5/5
- 预计耗时
- 一周以上
- 活跃度
- 冷清
- 描述清晰度
- 描述清楚
- 新手友好度
- 35/100