[Question] The images for MS234 on Cantus are in PDF format and aren't numbered by folio; what do we do?
- Dominant language
- Python
- Stars
- 0
- Forks
- 2
- Avg merge
- 10h 52m
- Merged PRs (30d)
- 42
Description
The link to the external images for MS234 on CantusDB brings the user to a PDF scan of the full manuscript. Since we're planning on running the manuscript through Mothra, there are two issues that I can perceive with this:
1. We need image formats of individual folios, so I guess we'll need to convert every page?
2. Since the images are in a PDF doc, the metadata numbers them by page, not by folio. But CantusDB numbers them by folio, which means that in Mothra, if I lock in the CantusID and want to select a folio or range of folios, the dropdown menu shows me folio numbers. Some of the folio images have folio numbers written in Roman numeral at the top, but some of them don't. So far I've dealt with this by figuring out the folio number myself, one folio at a time, but that won't work for the whole manuscript!
@kyrieb-ekat thoughts?
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the Mothra workflow that locks a CantusID and populates the folio or range dropdown. Compare CantusDB folio identifiers with the PDF's page metadata and determine how individual folio images should be supplied. Done means there is an agreed, repeatable workflow for importing MS234 and selecting folios correctly.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100