googleapis / googleapis/python-aiplatform

Shipped: vertexai-openeval-adapter — EvalPort import/export for vertexai.evaluation results

Đang mở
#7,078 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
api: vertex-ai
Ngôn ngữ chính
Python
Star
905
Fork
465
Merge trung bình
1 ngày 13 giờ
Pull request đã merge (30 ngày)
44

Mô tả

Built and shipped a standalone adapter that converts `vertexai.evaluation` metrics and `EvalResult`s to and from [EvalPort](https://github.com/adhabnr-ux/evalport) (Apache 2.0) — an open interchange format for portable LLM evaluation datasets (test cases, graders, suites, and result sets as plain JSON). It's already integrated with UK AISI's Inspect AI (PR merged) and has standalone adapter packages for a dozen+ eval/observability frameworks (Ragas, LangSmith, MLflow, Braintrust, DeepEval-adjacent tools aside — AutoGen, CrewAI, Langfuse, Evidently, TruLens, Opik, Giskard, Argilla), so a Vertex AI Gen AI Evaluation Service adapter puts it in company with the rest of that ecosystem.

**[`vertexai-openeval-adapter`](https://github.com/adhabnr-ux/evalport/tree/main/adapters/vertexai-openeval-adapter)**

```python
import pandas as pd
from vertexai.evaluation import EvalTask, PointwiseMetric, PointwiseMetricPromptTemplate
from vertexai_openeval_adapter import to_openeval, from_openeval, eval_result_to_openeval

dataset = pd.DataFrame({"prompt": ["What is the capital of France?"], "reference": ["Paris"]})
suite = to_openeval(dataset, input_column="prompt", expected_output_column="reference", suite_id="geo_quiz")

from openeval.validate import validate_suite
assert validate_suite(suite).valid

quality_metric = PointwiseMetric(
metric="quality",
metric_prompt_template=PointwiseMetricPromptTemplate(
criteria="Is the response factually correct?", metric_definition="Factual accuracy"
),
)
eval_task = EvalTask(dataset=dataset, metrics=[quality_metric])
result = eval_task.evaluate()

result_set = eval_result_to_openeval(result, suite_id="geo_quiz", run_id="run-1", started_at="2026-08-16T00:00:00Z")
assert validate_result_set(result_set).valid
```

The metric-mapping is the part I'd flag as genuinely interesting rather than routine: `PointwiseMetric` maps to EvalPort's `llm_judge` grader with the **actual rendered prompt template** preserved verbatim in the grader's `params.prompt_template` (read directly from `PointwiseMetricPromptTemplate`'s own rendering, not reconstructed or guessed) — so a suite exported from Vertex AI carries the real judge instructions, not a placeholder. `CustomMetric` and `PairwiseMetric` are exported as `custom`-typed graders (execution-only, not reconstructed on import) since both compute client-side per Vertex's own docstrings and have no portable representation. Raw string metric names (`"rouge_1"`, `"bleu"`, etc.) are explicitly rejected with a `TypeError` rather than silently guessed at, since their scoring logic isn't introspectable from the SDK's own objects. The adapter reads `EvalResult.metrics_table` using Vertex's own column convention (`f"{metric_name}/score"`), verified directly against `vertexai/evaluation/_evaluation.py` source rather than assumed.

Tested against the real `google-cloud-aiplatform[evaluation]` package (not mocks) and the real `openeval.validate.validate_suite()`/`validate_result_set()`. Full README with the complete mapping table and round-trip notes: https://github.com/adhabnr-ux/evalport/tree/main/adapters/vertexai-openeval-adapter#readme

No action needed here — this lives entirely as an external package (`pip install vertexai-openeval-adapter`), zero footprint on `google-cloud-aiplatform` itself. Flagging mainly so it's discoverable; happy to adjust the mapping if the evaluation module's public API shifts.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Hướng nghiên cứu

This issue describes the external vertexai-openeval-adapter package rather than a change to python-aiplatform. Start with the linked adapter README and package path to understand its Vertex AI and EvalPort mappings. There is no in-repository entry point, test target, or acceptance criterion for a contributor to complete.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
python
Lĩnh vực
machine-learning
Loại issue
Tính năng
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Ít trao đổi
Độ rõ ràng
Cần làm rõ
Mức phù hợp với người mới
10/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.