bigscience-workshop / bigscience-workshop/biomedical
Proposal to add EMEA Parallel datasets
Open
- Dominant language
- Python
- Stars
- 505
- Forks
- 117
- PR merge metrics
- No merged PRs in 30d
Description
I wouldn't mind contributing translation pairs for the EMEA drug notices. They are already available here:
> This is a parallel corpus made out of PDF documents from the European Medicines Agency. All files are automatically converted from PDF to plain text using pdftotext.
> https://opus.nlpl.eu/EMEA.php
I have also found manually aligned data on the same dataset, which is not very well known, but I could contribute too:
> https://link.springer.com/chapter/10.1007/978-3-642-40802-1_32
If that dataset is not already being considered for span tagging, I could work on that too.
What are your thoughts?
Contributor guide
Assessment
This issue has not been assessed yet.