microsoft / microsoft/AVGen-Bench
Clarification on scorer-code versions and leaderboard scores
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 31
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Description
Hi AVGen-Bench authors,
I’m Mukesh from NVIDIA, and I’m evaluating AVGen-Bench results for comparison with the public leaderboard.
I am seeing differences between my reproduced scores and the public leaderboard values, particularly for LTX-2.3 (Outputs imported from here).
My understanding is that the public LTX-2.3 scores were generated in March, while the OSS Gemini scorer code received substantial updates in May, including the robustness change and the official-client change. It appears that the static leaderboard values may not have been recomputed after those updates.
For reference, below are examples of the differences I observe for LTX-2.3:
| Metric | My reproduced score | Public leaderboard score |
|---|---|---|
| Text | 43.1250 | 54.17 |
| Music | 2.08 | 1.38 |
| Holistic | 67.2638 | 65.22 |
Could you please confirm:
- Which scorer-code version was used for the public LTX-2.3 leaderboard scores?
- Are score differences expected when using the current OSS Gemini scorers?
- Are there plans to recompute or version the public leaderboard results?
Thank you.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by comparing the linked LTX-2.3 outputs with the scorer changes in commits 75343b04ca49dfbb0ea014eb6c695bc66945fc4f and 57d8a37bbfb8c4c704f4780be76841c7e58219fd. Check the repository's scorer entry points and leaderboard data to identify the version used; done when the public-score provenance and any recomputation or versioning plan are documented.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- documentation
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100