[Text Alignment] New data from MS73 is low quality
- Dominant language
- CSS
- Stars
- 46
- Forks
- 13
- PR merge metrics
- No merged PRs in 30d
Description
Recently we received an ocr data set for MS73 from a past lab member. Unfortunately this data is of a slightly lower quality than the data from Salzinnes and St Gallen. Firstly, the decorative capital letters are labelled with their letter value as opposed to ~. Also, many of the images have blur, are not centred, or scaled weirdly. Here are some examples.

Alleluia

confidentem Perpetua Gloria

rem et
To integrate this data into new models, the capital letters need to be re-annotated, and the messy images may need to be cleaned
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.