obophenotype / obophenotype/uberon
Uberon BOT module is huge; recommendations to shrink?
Nobody has claimed this yet.
- Dominant language
- Emacs Lisp
- Stars
- 163
- Forks
- 43
- Avg merge
- 1d 17h
- Merged PRs (30d)
- 5
Description
Embedding UBERON in species specific phenotype ontologies is exceedingly hard because even small signatures lead to enormous bottom modules. For example, the import of UBERON for hp (33 MB) is larger than hp itself (27 MB). This is a problem; it makes building and QC processes long and merged artefacts huge. Does anyone have an idea, ever so hacky, to deal with this? Obviously UBERON is highly connected with many relations connecting all sorts of branches; we cant change that, and we need those for querying, inference and all. Still feel something should be done.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the HP import and comparing its reported 33 MB size with the 27 MB HP artefact. Then trace how the UBERON bottom module is produced and affects build and QC processes. Done would require an agreed, concrete approach to reduce artefact size without losing required querying and inference relations.
Written by the indexing model from the issue text.
Assessment
- Domain
- build-system, data
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100