googleapis / googleapis/python-aiplatform

response_evaluation_score (COHERENCE) fails with INVALID_ARGUMENT: Error parsing JSON in all regions — vertexai 1.153.1

Ouverte
#6,853 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
api: vertex-ai
Langage dominant
Python
Étoiles
905
Forks
465
Merge moyen
1 j 13 h
PR mergées (30 j)
44

Description

## Bug description

When using `PrebuiltMetric.COHERENCE` via `client.evals.evaluate()` (which backs `response_evaluation_score` in google-adk), the call consistently fails with:

```
400 INVALID_ARGUMENT: Error parsing JSON. Last error: Expecting property name
enclosed in double quotes: line 1 column 2 (char 1), Input: {* **STEP 1: ...}
```

The autorater LLM generates a free-text response starting with `{*` (bullet points), and the `:evaluateInstances` endpoint attempts to parse it as JSON and fails.

## Reproduction

```python
import vertexai
import pandas as pd

client = vertexai.Client(project="", location="us-central1")

df = pd.DataFrame([{
"prompt": "What slabs are in the model?",
"response": "The model has 17 C25 concrete slabs."
}])

result = client.evals.evaluate(
dataset=df,
metrics=[vertexai.types.PrebuiltMetric.COHERENCE],
)
```

## Environment

- `google-cloud-aiplatform` (vertexai): **1.153.1**
- Regions tested: `us-central1`, `europe-west1`, `europe-west3` — all fail
- The COHERENCE metric YAML (`gs://vertex-ai-generative-ai-eval-sdk-resources/metrics/coherence/v1.yaml`) produces free-text output, but the `:evaluateInstances` server endpoint attempts to parse the autorater response as JSON and fails

## Expected behavior

Metric evaluation completes and returns a coherence score.

## Actual behavior

`num_cases_error=1`, `num_cases_valid=0` on every call. The full error from the server:

```
google.genai.errors.ClientError: 400 INVALID_ARGUMENT.
Error parsing JSON. Last error: Expecting property name enclosed in double quotes:
line 1 column 2 (char 1), Input: {* **STEP 1: Identify the purpose and audience:...
```

The autorater generates a valid free-text evaluation but the server-side JSON parser rejects it.

Guide de contribution

Ouvrir le guide de contribution

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.