OpenEuroLLM / OpenEuroLLM/Taskboard

Annotate pre-training datasets for copyright

Open
#409 0 comments 0 reactions 1 assignee View on GitHub

@jsaizant is already working on this.

Since Sep 17, 2026.

T3.4 - regulatory compliance WP3
Dominant language
No language data
Stars
3
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Task framed within the WP3.4 Regulatory Compliance for detecting licenses and copyright-related content in text pre-training datasets, mainly coming from OCR processes, and for filtering out restrictive or copyrighted documents based on that detection.

There is an existing module developed and tested by the BSC, mainly consisting of a set of rules and heuristics for Spanish and English. The characteristics and limitations will be explained in detailed in this task, along with the repository with the code and test data which will be linked as well. The module originated from the necessity of filtering the FinePDFs dataset, which contains many OCR-based texts with copyright flags.

The expected outcome of this task is:

  • Publish copyright detection module on OELLM GitHub.
  • Apply copyright detection module to FinePDFs.
  • Analyze annotations and data retention in FInePDFS after filtering copyright.
  • Integrate filtered FinePDFs in the WP3 Training Catalogue.
  • Iterate over the copyright detection module.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.