feat(eval): expose maxTokens / temperature / topP flags for llm-as-a-judge evaluator (Bedrock)
- Dominant language
- TypeScript
- Stars
- 283
- Forks
- 95
- Avg merge
- 1d 2h
- Merged PRs (30d)
- 183
Description
### Summary
The `evaluator llm-as-a-judge create` handler only passes `modelId` to the Bedrock evaluator model config — there is no way to set inference parameters. CloudFormation's [`AWS::BedrockAgentCore::Evaluator InferenceConfiguration`](https://docs.aws.amazon.com/AWSCloudFormation/latest/TemplateReference/aws-properties-bedrockagentcore-evaluator-inferenceconfiguration.html) supports:
- `MaxTokens` — Integer, minimum `1`
- `Temperature` — Number, `0`–`1`
- `TopP` — Number, `0`–`1`
These map to `BedrockEvaluatorModelConfig.inferenceConfig` and apply to both the Bedrock and Bedrock **Mantle** evaluator paths.
### Current behavior
`src/handlers/eval/evaluator/llm-as-a-judge/create/index.tsx` builds the config with just the model id:
```ts
modelConfig: { bedrockEvaluatorModelConfig: { modelId: flags["model"] } },
```
There are no `--max-tokens`, `--temperature`, or `--top-p` flags, so users get whatever the service/construct defaults are and cannot tune scoring behavior.
For comparison, the **harness** handler already exposes these — see `src/core/project/schema/harness.ts` (`temperature`, `topP`, `maxTokens`) and `src/handlers/harness/parameterHelp.tsx`.
### Requested change
Add flags to `llm-as-a-judge create` (and any `update`/edit path):
- `--max-tokens ` (min 1)
- `--temperature ` (0.0–1.0)
- `--top-p ` (0.0–1.0)
Thread them into `bedrockEvaluatorModelConfig.inferenceConfig` (`maxTokens` / `temperature` / `topP`), omitting any unset value so existing defaults are preserved. Enforce the CFN-documented ranges above.
### Related
Companion construct/schema change: **aws/agentcore-l3-cdk-constructs#315** (the L3 currently hardcodes `{ temperature: 0, maxTokens: 4096 }` for Bedrock and rejects these fields unless the provider is OpenAI). This CLI change depends on that surface being available.
Contributor guide
Assessment
This issue has not been assessed yet.