DDMAL / DDMAL/mothra

[Question] The images for MS234 on Cantus are in PDF format and aren't numbered by folio; what do we do?

Open
#324 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
0
Forks
2
Avg merge
10h 52m
Merged PRs (30d)
42

Description

The link to the external images for MS234 on CantusDB brings the user to a PDF scan of the full manuscript. Since we're planning on running the manuscript through Mothra, there are two issues that I can perceive with this:

1. We need image formats of individual folios, so I guess we'll need to convert every page?
2. Since the images are in a PDF doc, the metadata numbers them by page, not by folio. But CantusDB numbers them by folio, which means that in Mothra, if I lock in the CantusID and want to select a folio or range of folios, the dropdown menu shows me folio numbers. Some of the folio images have folio numbers written in Roman numeral at the top, but some of them don't. So far I've dealt with this by figuring out the folio number myself, one folio at a time, but that won't work for the whole manuscript!

@kyrieb-ekat thoughts?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the Mothra workflow that locks a CantusID and populates the folio or range dropdown. Compare CantusDB folio identifiers with the PDF's page metadata and determine how individual folio images should be supplied. Done means there is an agreed, repeatable workflow for importing MS234 and selecting folios correctly.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.