OpenEuroLLM / OpenEuroLLM/Taskboard

Ensure that we do long context evaluation throughout the pipeline (at all stages)

Open
#342 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

4.6 post-training T5.2 - dynamic evals
Dominant language
No language data
Stars
3
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Goal
We need to ensure that we do not lose long context capabilities at SFT, DP, RL, et.c.

Description

Deliverable scope
Include Long context tasks in post-training evals (NiH, et.c)

Dependencies

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by mapping the SFT, DP, and RL stages named in the issue and identifying how post-training evaluations are currently run. Add long-context tasks such as NiH to the post-training evaluations at each stage, then verify that long-context coverage is present throughout the pipeline.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning, testing-qa
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.