[CI-Work] Question about data availability
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 30/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Quiet
- Domain
- data
Research direction
Start by reading the README and checking the referenced data/seed/raw_seed.jsonl path, then verify where the 125 seeds and downstream artifacts are expected to come from. Done means documenting a data-release plan or making the requested benchmark files available so the evaluation can be reproduced.
Written by the indexing model from the issue text.
Description
Hi CI-Work authors,
Thanks for the great work and for releasing the code and prompt templates.
I was trying to run the evaluation pipeline and noticed that the data files (e.g., data/seed/raw_seed.jsonl) aren't in the repo.
The links in the README (https://aka.ms/ci-work and https://aka.ms/ci-work-code) both seem to point back to this code repo, and I couldn't find a separate data release.
Is there a plan to release the benchmark data? Ideally:
- The 125 task-oriented seeds (
raw_seed.jsonl) used in the paper. - Optionally, the downstream artifacts (enriched seeds, case episodes, formatted instructions) so people can run the trajectory evaluation directly without regenerating everything.
Either would be very helpful for reproducing the results.
Thanks!
- Dominant language
- Jupyter Notebook
- Stars
- 41
- Forks
- 8
- PR merge metrics
- No merged PRs in 30d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/ACV
-
Difficulty 2/5 1-3 hours Newbie friendliness 35/100
-
Difficulty 4/5 3-5 days Newbie friendliness 25/100
-
Difficulty 4/5 3-5 days Newbie friendliness 30/100
-
Difficulty 5/5 Over a week Newbie friendliness 20/100
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
-
Schema-level dtype cannot be serialized: to_yaml raises RepresenterError, to_json raises TypeError Open
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
unionai-oss/pandera#2511 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
unitaryfoundation/qldpc-challenge#1651 ·
-
Difficulty 1/5 Under an hour Newbie friendliness 90/100
statsmodels/statsmodels#10271 ·