OpenEuroLLM / OpenEuroLLM/Taskboard

Curate post-training datasets with propella

Open
#318 2 comments 0 reactions 3 assignees View on GitHub

@geoalgo is already working on this.

Since Jun 18, 2026.

4.6 post-training research
Dominant language
No language data
Stars
3
Forks
0
PR merge metrics
No merged PRs in 30d

Description

We investigate the extent to which we can improve performance by curating post-training datasets using propella annotation.
We first focus on SFT data. The concept can later be extended to preference learning (DPO, RL).

Propella can be used to curate a dataset from sample-wise annotations, allowing us to filter and compose from various sources.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.