intermine / intermine/pombemine
Don't load "comments " from UniProt
- Dominant language
- Java
- Stars
- 0
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Description
The comments are the free text annotation from UniProt. So they are really just "unstructured" annotations.

However these annotations are not (currently) attached to entities, so you can't really do anything with this data.
Although the comments are "typed", they are still problematic because as free text they are not standardized, and they do not include the associated metadata and prrovenence (sometimes the source is within the comment, but not often)
(176) Phenotypes -> covered by FYPO (PomBase: 60,000)
(32750 Functions -& 500 pathway> covered by GO MF & BP (PomBase:25,000)
(1780) Subunit -> covered by GO complex annotations(Pombase:5000)
(464) Subcellular location - Covered by PomBase GO component annotations(Pombase:10,000)
(220) PTM -> covered by PRO annotation (PomBase:602210 PTMs (not yet imported-https://github.com/intermine/pombemine/issues/12
These annotation aren't really "mineable" because the text is a mixture of concepts i.e. phenotypes/ penetrance
The phenotypes are not associated with specific alleles (even if the connection was imported it is only to the protein, not the allele)
We have standardised assignments to PRO ontology terms and captured modified residues and modifying entities in a standardized way
etc
On balance I don't think. it is very useful to include these annotation in a data mining tool because the inconsistencies. and incmpleteness could make them more misleading than useful.
ACTION: don't import data in "comments" at least for the time being
discuss with
@manulera
@kimrutherford
for another opinion,
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.