OpenEuroLLM / OpenEuroLLM/Taskboard
Annotate pre-training datasets for copyright
@jsaizant is already working on this.
Since Sep 17, 2026.
- Dominant language
- No language data
- Stars
- 3
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
Task framed within the WP3.4 Regulatory Compliance for detecting licenses and copyright-related content in text pre-training datasets, mainly coming from OCR processes, and for filtering out restrictive or copyrighted documents based on that detection.
There is an existing module developed and tested by the BSC, mainly consisting of a set of rules and heuristics for Spanish and English. The characteristics and limitations will be explained in detailed in this task, along with the repository with the code and test data which will be linked as well. The module originated from the necessity of filtering the FinePDFs dataset, which contains many OCR-based texts with copyright flags.
The expected outcome of this task is:
- Publish copyright detection module on OELLM GitHub.
- Apply copyright detection module to FinePDFs.
- Analyze annotations and data retention in FInePDFS after filtering copyright.
- Integrate filtered FinePDFs in the WP3 Training Catalogue.
- Iterate over the copyright detection module.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.