lm-sys / lm-sys/FastChat

Inquiry Regarding Your Paper "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena"

Open
#3,590 0 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
39.5k
Forks
4.8k
PR merge metrics
No merged PRs in 30d

Description

Hello, let me some questions.

1. Is the evaluation in your paper based on a comparison between the human evaluation dataset provided by Hugging Face (https://huggingface.co/datasets/lmsys/mt_bench_human_judgments) and the outputs of LLMs? If not, would you kindly share the dataset used for the human evaluation in your paper?

2. If the answer to the first question is yes, I have an additional question. I noticed that the human evaluation was performed only on a subset of the questions, and that the number of evaluators varies across these questions. For example, "question_82_turn1 model1: claude-v1 model2: vicuna-13b-v1.2 judgment: No man", and "question_82_turn1 model1: vicuna-13b-v1.2 model2: gpt-3.5-turbo judgment: author_4, expert_2, expert_20, expert_24". Did you use the dataset as it is, or did you process it before calculating the evaluation scores?

3. The Hugging Face dataset lists "vicuna-13b-v1.2" as one of the evaluated models, but the models retrieved from the FastChat repository, include "vicuna-13b-v1.3" instead of "vicuna-13b-v1.2". Which model version was used in your evaluation? If both are correct, could you explain the reason for the version difference?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No project files, tests, or entry points are named. Start by reviewing the linked Hugging Face dataset and the paper's evaluation materials, then determine whether the dataset, preprocessing steps, and Vicuna model version are documented; done means providing authoritative answers or links for all three questions.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.