OpenEuroLLM / OpenEuroLLM/Taskboard
Ensure that we do long context evaluation throughout the pipeline (at all stages)
Open
Nobody has claimed this yet.
4.6 post-training
T5.2 - dynamic evals
- Dominant language
- No language data
- Stars
- 3
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
Goal
We need to ensure that we do not lose long context capabilities at SFT, DP, RL, et.c.
Description
Deliverable scope
Include Long context tasks in post-training evals (NiH, et.c)
Dependencies
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by mapping the SFT, DP, and RL stages named in the issue and identifying how post-training evaluations are currently run. Add long-context tasks such as NiH to the post-training evaluations at each stage, then verify that long-context coverage is present throughout the pipeline.
Written by the indexing model from the issue text.
Assessment
- Domain
- machine-learning, testing-qa
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100