microsoft / microsoft/SkillOpt

Several reported scores are incompatible with the released test split sizes

Open
#158 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
17.3k
Forks
1.6k
Avg merge
2d 7h
Merged PRs (30d)
17

Description

Hi, thanks for releasing SkillOpt. I found a possible mismatch between the paper results and the currently released dataset splits.

The released code computes the hard score as the mean of per-example scores. For LiveMath, the evaluator is binary exact match, and the released manifest contains 124 test items:

  • data/livemathematicianbench_id_split/split_manifest.json
  • skillopt/envs/livemathematicianbench/evaluator.py
  • skillopt/utils/scoring.py

However, several LiveMath results in Table 1 cannot be obtained from 124 binary examples after rounding to one decimal:

Score Implied result
22.4 28/125
27.2 34/125
52.0 65/125
31.2 39/125
29.6 37/125
41.6 52/125

These values all fall exactly on a 125-item grid, suggesting that some experiments may have used a 125-item test split rather than the released 124-item split.

I found a similar issue for DocVQA. Its released test split contains 374 items, but the reported score 87.6 is not attainable in a single binary run:

  • 327/374 = 87.4%
  • 328/374 = 87.7%

There are also some denominator-incompatible cells for OfficeQA and SpreadsheetBench under the released hard-score implementation.

Could you please clarify:

  1. Did the paper use an earlier version of the dataset splits?
  2. Were the reported values averaged across multiple runs?
  3. Were any failed or invalid examples removed before aggregation?
  4. Could the exact paper splits and aggregation protocol be released?

Publishing per-example results or the exact result manifests would make the reported table much easier to reproduce.

Related: #108

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by checking data/livemathematicianbench_id_split/split_manifest.json, skillopt/envs/livemathematicianbench/evaluator.py, and skillopt/utils/scoring.py to verify the released item counts and aggregation behavior. Compare those results with the reported table values and document whether alternate splits, repeated runs, or excluded examples explain the discrepancy; done means the exact paper splits and aggregation protocol are clarified or released.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data, testing-qa
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.