openlibhums / openlibhums/janeway

Scrub file metadata from .docx / .odt /pdf submissions

Open
#1,810 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

documentation new feature
Dominant language
Python
Stars
238
Forks
97
Avg merge
9d 1h
Merged PRs (30d)
8

Description

Is your feature request related to a problem? Please describe.
When handling .docx files during the peer-review authors, editors and reviewers often forget to clear any document metadata that might betray their identities. Ideally, Janeway could clear this metadata when uploading/downloading documents during submission and peer reveiew

Files to be scrubbed includes:

  • Original manuscript
  • Annotated files handed over as peer-review comments to the author after a peer-review round is completed.

Describe the solution you'd like
As an editor, when an author submits a paper in .docx or .odt or pdf format, I would like it for the system to scrub the metadata associated with the document.

Describe alternatives you've considered
To keep this as a manual process

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by tracing the submission upload and download handling for .docx, .odt, and PDF files. Check how original manuscripts and annotated peer-review files move through the system; done means metadata is scrubbed from both file categories before the relevant author or reviewer access.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend, security
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.