ProjectTech4DevAI / ProjectTech4DevAI/kaapi-frontend
Evaluation: Include judge cost tooltip
オープン
初心者向け
まだ誰も着手していません。
- 主要言語
- TypeScript
- スター
- 1
- フォーク
- 0
- 平均マージ
- 10時間 40分
- マージ済み PR(30日)
- 4
説明
Is your feature request related to a problem?
The cost tooltip on /evaluations only lists Response generation and omits the judge cost, leading to discrepancies in the cost breakdown compared to the total displayed. This causes confusion for users trying to understand the complete cost.
Describe the solution you'd like
- Update
EvalCostto includejudge?: EvalCostEntry. - Render a Judge scoring entry in the tooltip of
EvalRunCardwhenjob.cost.judgeis present. - Ensure all cost entries (response, judge, embedding) are shown in the tooltip to match
total_cost_usd.
Additional Context
Backend response
{
"id": 887,
"run_name": "assistant_v2_v1_0_ai_cohort_2_evals_demo_goldenqna_1788945623629",
"dataset_name": "ai_cohort_2_evals_demo_goldenqna",
"config_id": "dc576d3c-5e86-4eef-9b95-6d1e2194cce4",
"config_version": 1,
"dataset_id": 709,
"batch_job_id": 1805,
"embedding_batch_job_id": null,
"status": "completed",
"run_mode": "fast",
"object_store_url": null,
"score_trace_url": "s3://ai-platform-documents-staging/3ce7b9fe-2900-4f33-9a68-8162568a41be/evaluations/score/887/traces_887.json",
"total_items": 9,
"score": {
"overall": {
"verdict": "Needs Refinement",
"breakdown": [
{
"key": "ground_truth",
"name": "Adherence to Ground Truth",
"delta": -0.45,
"score": 3.44,
"weight": 0.71,
"verdict": "Needs Refinement"
},
{
"key": "prompt",
"name": "Adherence to Prompt",
"delta": 1.11,
"score": 5,
"weight": 0.29,
"verdict": "Good"
}
],
"ai_summary": "**Overall read:** The run is in generally good shape — most questions score 4–5 on ground truth and a clean 5 on prompt adherence, with no KB in play. The model answers are consistently substantive and well-structured; the main tension is between the model giving richer, modern-science answers and golden answers that expect specific, textbook-narrow responses.\n\n**Top 3 to check:**\n\n**Question 9** — Ground-truth score of 0: the golden answer expects a very specific socio-demographic list (sex, skin colour, caste, mother tongue, etc.) but the model answered from a biological/population-genetics frame; this looks like a golden-dataset framing issue more than a model failure, but needs a human call on which answer the use case actually wants.\n\n**Question 7** — Borderline ground-truth score (2): the model explicitly refuses to classify by skin colour and race, directly conflicting with the golden answer that includes skin colour; this is a values/alignment tension between the model's safety behaviour and the expected answer — worth deciding whether the golden answer or the model's stance is appropriate for this context.\n\n**Question 4** — Minor: the macrophage-as-viral-factory stage (a key step in the reference answer) is omitted; solid overall but worth a quick check if curriculum accuracy to the specific textbook is required, pointing at the model.\n\nThese are go-verify pointers — open each item, read the actual answer against the use case requirements, and decide based on what the deployment needs.",
"overall_score": 3.89
},
"summary_scores": [
{
"avg": 3.44,
"std": 1.42,
"name": "Adherence to Ground Truth",
"data_type": "NUMERIC",
"total_pairs": 9
},
{
"avg": 5,
"std": 0,
"name": "Adherence to Prompt",
"data_type": "NUMERIC",
"total_pairs": 9
}
]
},
"unscoreable": null,
"is_score_updated": true,
"is_judge_run": true,
"cost": {
"judge": {
"model": "gpt-5.6-luna",
"cost_usd": 0.002832,
"input_tokens": 16061,
"total_tokens": 18104,
"output_tokens": 2043
},
"response": {
"model": "gpt-5.6-luna",
"cost_usd": 0.001942,
"input_tokens": 276,
"total_tokens": 3466,
"output_tokens": 3190
},
"total_cost_usd": 0.004774
},
"error_message": null,
"organization_id": 1,
"project_id": 1,
"inserted_at": "2026-09-09T09:20:25.182635",
"updated_at": "2026-09-09T09:24:23.885105"
},
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
調査の方向性
/evaluations ページから開始し、EvalCost 型と EvalRunCard のツールチップのレンダリングを追跡します。job.cost.judge が存在する場合は、レスポンスコストおよび埋め込みコストと並べて判定コストの項目を追加し、ツールチップに利用可能なすべてのコストが表示され、total_cost_usd と一致することを確認します。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- typescript
- 領域
- frontend
- issue の種類
- 機能追加
- 難易度
- 2/5
- 見積もり時間
- 1〜3時間
- 活発さ
- 活発
- 明瞭さ
- 明確に書かれている
- 初心者へのやさしさ
- 78/100