aws / aws/bedrock-agentcore-sdk-python

feat: Evaluation Client — Lifecycle, Orchestration & Online Pipeline

Offen
#393 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
enhancement
Vorherrschende Sprache
Python
Sterne
761
Forks
147
Ø Merge
1 T. 23 Std.
Gemergte PRs (30 T.)
7

Beschreibung

## Problem

The SDK's `EvaluationClient` only exposes `run()`. On the control plane side, customers cannot programmatically create custom evaluators (LLM-as-a-judge configs), list available evaluators, update or delete evaluators, or manage online evaluation configs for continuous evaluation on live traffic — evaluator provisioning requires the console. On the data plane side, the starter toolkit's `EvaluationProcessor` provides significantly richer orchestration than `run()`: it fetches session data from CloudWatch independently, groups evaluators by level (SESSION vs TRACE), determines which spans to send based on evaluator level, and runs multiple evaluators with per-evaluator error handling. The toolkit also provides input validation, IAM role cleanup on delete, and typed config/result models.

## Acceptance Criteria

- [ ] Customers can create, get, list, update, and delete custom evaluators
- [ ] Customers can create, get, list, update, and delete online evaluation configs
- [ ] Online evaluation config supports enable/disable toggling and sampling rate adjustment
- [ ] Typed result models with error introspection (`has_error()`, `get_successful_results()`)
- [ ] Customers can fetch session trace data (spans + runtime logs) from CloudWatch for a given session and agent
- [ ] Customers can find the most recent session for an agent
- [ ] Multi-evaluator orchestration groups evaluators by level and selects appropriate spans per level
- [ ] Per-evaluator error handling — failures on one evaluator don't block others
- [ ] Online evaluation config deletion supports optional IAM execution role cleanup
- [ ] All functionality is verified via integration tests running in CI

## Relevant Links

- [`EvaluationControlPlaneClient`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/control_plane_client.py#L21)
- [`create_evaluator()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/control_plane_client.py#L107)
- [`create_online_evaluation_config()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/control_plane_client.py#L175)
- [`update_online_evaluation_config()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/control_plane_client.py#L312)
- [`EvaluationResult` / `EvaluationResults`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/models.py#L81)
- [`OnlineEvaluationConfig`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/models.py#L190)
- [`EvaluationProcessor`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/on_demand_processor.py#L27)
- [`evaluate_session()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/on_demand_processor.py#L418)
- [`fetch_session_data()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/on_demand_processor.py#L88)
- [`determine_spans_for_evaluator()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/on_demand_processor.py#L304)
- [`execute_evaluators()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/on_demand_processor.py#L343)
- [`EvaluationDataPlaneClient`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/data_plane_client.py#L20)
- [`delete_online_evaluation_config()`](https://github.com/aws/bedrock-agentcore-starter-toolkit/blob/4b9387f0d48cb6639633437b669fc6cc09ef07be/src/bedrock_agentcore_starter_toolkit/operations/evaluation/online_processor.py#L200)

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Beginne mit operations/evaluation/control_plane_client.py und models.py, lies anschließend EvaluationProcessor in on_demand_processor.py, EvaluationDataPlaneClient und delete_online_evaluation_config() in online_processor.py. Verfolge zuerst die verknüpften Einstiegspunkte und die vorhandene Einrichtung der Integrationstests. Erledigt ist die Aufgabe, wenn alle aufgeführten Kriterien für Evaluator, Online-Konfiguration, Sitzungsdaten, Orchestrierung, Fehlerergebnisse, Bereinigung und CI-Integrationstests abgedeckt sind.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
aws, python
Bereich
backend-api-design, cloud, testing
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Ruhig
Klarheit
Klar beschrieben
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.