obophenotype / obophenotype/uberon

Uberon BOT module is huge; recommendations to shrink?

Open
#1,478 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Emacs Lisp
Stars
163
Forks
43
Avg merge
1d 17h
Merged PRs (30d)
5

Description

Embedding UBERON in species specific phenotype ontologies is exceedingly hard because even small signatures lead to enormous bottom modules. For example, the import of UBERON for hp (33 MB) is larger than hp itself (27 MB). This is a problem; it makes building and QC processes long and merged artefacts huge. Does anyone have an idea, ever so hacky, to deal with this? Obviously UBERON is highly connected with many relations connecting all sorts of branches; we cant change that, and we need those for querying, inference and all. Still feel something should be done.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the HP import and comparing its reported 33 MB size with the 27 MB HP artefact. Then trace how the UBERON bottom module is produced and affects build and QC processes. Done would require an agreed, concrete approach to reduce artefact size without losing required querying and inference relations.

Written by the indexing model from the issue text.

Assessment

Domain
build-system, data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.