bigscience-workshop / bigscience-workshop/biomedical

Proposal to add EMEA Parallel datasets

Open
#338 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
505
Forks
117
PR merge metrics
No merged PRs in 30d

Description

I wouldn't mind contributing translation pairs for the EMEA drug notices. They are already available here:

> This is a parallel corpus made out of PDF documents from the European Medicines Agency. All files are automatically converted from PDF to plain text using pdftotext.

> https://opus.nlpl.eu/EMEA.php

I have also found manually aligned data on the same dataset, which is not very well known, but I could contribute too:

> https://link.springer.com/chapter/10.1007/978-3-642-40802-1_32

If that dataset is not already being considered for span tagging, I could work on that too.

What are your thoughts?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.