microsoft / microsoft/AVGen-Bench

Clarification on scorer-code versions and leaderboard scores

Open
#43 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
31
Forks
1
PR merge metrics
No merged PRs in 30d

Description

Hi AVGen-Bench authors,

I’m Mukesh from NVIDIA, and I’m evaluating AVGen-Bench results for comparison with the public leaderboard.

I am seeing differences between my reproduced scores and the public leaderboard values, particularly for LTX-2.3 (Outputs imported from here).

My understanding is that the public LTX-2.3 scores were generated in March, while the OSS Gemini scorer code received substantial updates in May, including the robustness change and the official-client change. It appears that the static leaderboard values may not have been recomputed after those updates.

For reference, below are examples of the differences I observe for LTX-2.3:

Metric My reproduced score Public leaderboard score
Text 43.1250 54.17
Music 2.08 1.38
Holistic 67.2638 65.22

Could you please confirm:

  • Which scorer-code version was used for the public LTX-2.3 leaderboard scores?
  • Are score differences expected when using the current OSS Gemini scorers?
  • Are there plans to recompute or version the public leaderboard results?

Thank you.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by comparing the linked LTX-2.3 outputs with the scorer changes in commits 75343b04ca49dfbb0ea014eb6c695bc66945fc4f and 57d8a37bbfb8c4c704f4780be76841c7e58219fd. Check the repository's scorer entry points and leaderboard data to identify the version used; done when the public-score provenance and any recomputation or versioning plan are documented.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
documentation
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.