internetarchive / internetarchive/openlibrary

Maintenance dashboard UI for collaborative data quality tasks, reports for monitoring progress

Open
#7,630 4 comments 0 reactions 0 assignees View on GitHub
Affects: Librarians Lead: @jimchamp Module: Integrated Librarian Environment (ILE Bar) Needs: Community Discussion Priority: 3 Type: Feature Request
Dominant language
Python
Stars
6.7k
Forks
2k
Avg merge
2d 19h
Merged PRs (30d)
138

Description

There is much potential in the data needs for OpenLibrary to pair up existing tools as well as bot and humans for content curation and data quality review. Currently most of this either occur in specialized tools not linked or easily discoverable from the OpenLibrary UI, documentation, or navigation and these are isolated from the content for the issues they are addressing as the issues are not exposed on OpenLibrary to prompt users. This could potentially be called "FullyBooked" or some library pun regarding curation and/or reaching high data quality, but names aside the function is what matters.

### Describe the problem that you'd like solved
Primarily helping the community coordinate on large scale tasks that need human in the loop resolution where bots alone cannot resolve.

Create a UI for aggregation of crowdsourcing review and data cleanup tasks to be shared and process per each task to be laid out in a clear format and workflow to afford users of any skill levels to join in reviewing content.

Critical to the design consideration is focusing on defined tasks needing human review, clear resolution workflows, and protocols to assess tasks were resolved correctly instead of creating large scale data quality issues.
1. Defining known issues and listing them for task design - identify how the data is being addressed (fully automated if so, note the bot/script/quality process in place and who is the contact monitoring this vs task needing human in the loop) For human in the loop tasks who is the contact and what is the scale and need (one time vs ongoing) and if large
2. Define a workflow for each task and how the issue would be considered fix including defining how to flag false positives and notifying the task maintainer to improve the flagging or exempt from future tasks.
3. Designing an interface to discover one issue to check how the user marks complete to move on to the next tasked item to check.
4. Design to enable review tools for tracking what has been fixed by which users in case of large scale data issues present it could be tracked down to users of this tool.

The goal is for OpenLibrary to make the process of finding tasks more straightforward and enable more users to join in fixing of large scale issues by creating more "paint by number" oriented tasks thank needing to have internalized all the ways to fix issues. As well as creating a common area for those working on data quality to be able to add new quality review tools in a commons area for others to be aware of who and how they issues are being addressed.

(optional enhancement) Bring together existing tools as part of the clearing house architecture that are applicable to curation activities to make discovery of the tools more accessible from the site navigation and not buried in documentation.

### Proposal & Constraints
A possible integrated guiding concept for crowdsourcing task exposure in the OpenLibrary UI and complementary platform is the OpenFoodFacts Hunger Games site https://github.com/openfoodfacts/hunger-games https://hunger.openfoodfacts.org. Their interfaces pairs computationally flagged (image recognition tasks) as possible additions as suggested could be added to OpenFoodFacts both in the stand alone UI as well as on OpenFoodFacts product listing. Similarly OpenLibrary could use OCR and Image Recognition, Natural Language Processing, and basic regular expressions (like requested in #7629) to flag content to be added to fields on works/editions/authors, and build matching and data verification processes as tasks where human in the loop is critical to determine the difference.

This approach could be similarly applied to OpenLibrary in combination with feature request #7627 to rely on flags generated by existing and new bots as well as humans to identify what content should be reviewed, for what reason, and what is the process to resolve such data quality issue. ie document the cases and how to resolve them in a workflow and to identify edge cases not addressed.

Alternative crowdsourcing task design could look toward Zooniverse.org as a broad design concept of crowdsourcing and guided task design, but the scope of Hunger Games of OpenFoodFacts seems more applicable in the tighter integration with a specific set of data management components. The most important aspect of Zooniverse is the ability to discuss specific tasks if resolution it is unclear or conflated with additional issues.

A tool for OpenStreetMap may also be another tool for inspiration - Maproulette https://maproulette.org/ which functions as a collaborative QA task manager for which users can create and join in the task design and processing to resolve the issue or flag as "too hard" or "already fixed".

Designing of a UI framework that is flexible to display a variety of task types and workflows to process issue flags or labels for different content review. The task design should include a description of the problem, how to resolve, ways to flag edge cases that are not addressed in the "ways to resolve" to improve the documentation, examples of false positives, and workflow of resolving the task. Metrics for activities should be considered such as "what qualifies as a task" how to handle reverting users who don't understand the task or the task wasn't clear to track what changes might need to be reviewed that were performed via the task.

Design of task and maintenance dashboard as a component of the progress tracking. For each task a scope of the issue to the best of the ability could be determined report the number of items queued and number completed at the minimum.

Ensure edits made from the UI banner and from the tool are tagged in the changelog to identify the maintenance tool as the interface and/or process to track in the recent changes as a facet and in review of changes diff https://openlibrary.org/recentchanges

### Additional context
Hunger Games
![image](https://user-images.githubusercontent.com/717866/224364533-18934180-99e7-480f-a7c1-1db1890f166d.png)

Maproulette
![image](https://user-images.githubusercontent.com/717866/224364103-a79508fd-af45-49a1-9134-b84de3b988fb.png)

Wikivoyage maintenance dashboard (as well as many other open collaborative projects list their tasks outstanding to encourage people to jump in an assist with the maintenance and data. Example: https://en.wikivoyage.org/wiki/Wikivoyage:Maintenance_panel as well as some WikiProjects have dashboards for task - https://en.wikipedia.org/wiki/Wikipedia:WikiProject_Lakes/backlog

### Stakeholders

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.