INCATools / INCATools/boomer

Supporting Mapping QC workflow

Open
#333 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Scala
Stars
37
Forks
2
PR merge metrics
No merged PRs in 30d

Description

The mapping QC workflow is about reviewing the existing mappings on an ongoing basis. The idea is to review the bottom N clusters once per month and thereby implement an ongoing cycle of ever improving mappings.

Note, there is _no_ mappings being generated by this workflow. This is part of another issue.

### Workflow:

- Input Ontology O
- Input M: existing mappings separated into two levels of confidence
- Reviewed: 0.99 %
- Not reviewed 0.95 %
- Key: No new mappings are added
- PT=sssom-py:ptable(M)
- {results.json, |cluster-X.png|, |cluster-X.md}, =boomer(PT, O)
- {BOTTOM_10_CLUSTERS, LEAST_PROBABLE_MAPPINGS} = oak:boomerang(results.json, N)
- GitHub Action: make issues for BOTTOM_10_CLUSTERS, including `cluster-X.png and cluster-X.md`
- The reviewer now checks each cluster and _adds a `semapv:MappingReview` justification, which is separately curated from the existing mapping. If need be the existing mapping will be changed as well. This will be used to generate confidence scores for input M. There should never be more than 10 issues open. _Ideally we can somehow recognise for a given cluster that an issue already exists_ (by parsing its title for the hashcode boomer provides).

![image](https://user-images.githubusercontent.com/7070631/215748242-ba10d572-7fda-4fb9-a37d-1a2c22fcd441.png)

### New boomer requirements

- [ ] Output report results.json contains probability scores that enable us to select cliques which should be reviewed.
- [ ] results.json should conform to the new OAK cluster data model
- [ ] cluster-X.md files should be on a by-clique basis rather than one huge file and ideally already contain the image tag which can be assumed to be in the same directory (not sure how this will work with posting a github issue though - maybe you know how this could be automated)

### Comments

- "joint posterior prop most likely of clique / prop next most likely - how interesting is this cluster?" @cmungall

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by tracing the boomer workflow that produces results.json, cluster-X.md, and cluster-X.png, then inspect how the GitHub Action could consume those outputs. Done means the workflow reviews existing mappings without generating new ones, produces the required cluster reports, creates no more than 10 relevant issues, and supports the listed boomer requirements.

Written by the indexing model from the issue text.

Assessment

Tech stack
github-actions, scala
Domain
ci-cd, data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.