waldronlab / waldronlab/agent-protocols
Protocol: semantic similarity between signatures at mixed taxonomic levels
Nobody has claimed this yet.
- Dominant language
- No language data
- Stars
- 0
- Forks
- 1
- Avg merge
- 1h 2m
- Merged PRs (30d)
- 9
Description
Tier: C (BugSigDB paper) · Type: atomic · Category: Statistical Analysis
What
Compare two microbial signatures whose taxa are reported at different taxonomic ranks, using an information-content similarity measure (Lin's) over the NCBI taxonomy, with lowest-common-ancestor information content, combined pairwise into a signature-level score by best-match average.
Why it matters
Jaccard forces everything down to one rank and throws away the rest. Semantic similarity is what let the BugSigDB paper compare signatures as curated. It is the single most reusable method in that paper and currently exists only as vignette code that downloads ncbitaxon.obo from a URL in a commented-out line.
Source material
waldronlab/BugSigDBPaper—vignettes/Figure2.Rmd("Semantic similarity on mixed taxonomic levels", "Heatmap comparison" against the genus-level Jaccard matrix)- Paper: Geistlinger et al., Nat Biotechnol 2023, 10.1038/s41587-023-01872-y
Scope
In: which taxonomy version and how it is pinned; information content estimation and what corpus defines it; Lin's measure; the best-match-average combination rule and its asymmetry; handling of taxa absent from the taxonomy; the comparison against genus-level Jaccard as a sanity check.
Out: clustering and display of the resulting matrix — that is signature-similarity-clustering, which this protocol feeds on equal footing with jaccard-signature-similarity; metasignature pooling (separate).
Frontmatter starting point
type: "atomic"
category: "Statistical Analysis"
citation: "" # Lin's similarity measure, Lin 1998 (ICML) — no DOI; establish the correct
# primary reference. Resnik 1995 sits behind it for information content. NOT the
# BugSigDB paper, which applies the measure rather than proposing it.
tags: [bugsigdb, semantic-similarity, lin, ncbi-taxonomy, ontology]
Acceptance criteria
- Cites Lin's original paper, not the BugSigDB paper
- Output matrix layout matches what
signature-similarity-clusteringexpects - Taxonomy version pinning is a required step (the vignette's live OBO download is the failure mode to fix)
- The best-match-average rule is written out, including what happens with unequal signature lengths
- States the behavior for taxa with no taxonomy match rather than silently dropping them
Cite the method's origin, not its users
PROTOCOL_STANDARD.md is explicit: an atomic protocol carries "strictly 1 citation... corresponding
to the primary literature where the method was originally published." Find the paper that proposed
the method. Do not cite a paper that merely applied it — including the BugSigDB and curatedMetagenomicData
papers, which are the source of the analysis these protocols were extracted from but almost never the
source of the method.
Tracing a method back to its first publication is real work, and it is part of the task, not a
formality. Three things to expect:
- Candidate DOIs in this issue are leads, not answers. Anything marked VERIFY has not been checked.
- Some methods predate modern citation practice or have no single identifiable origin. If that is
genuinely the case, say so in the pull request rather than reaching for a convenient recent paper.
Raise it as an issue inwaldronlab/agent-protocol-standard— the standard may need a way to express
"classical method, no primary source". - If you cannot name one paper that proposed everything the protocol does, it is more than one
protocol. That test has now split four protocols out of this batch: enrichment into three methods,
filtering from transformation, LODO from random forest, and PERMANOVA from ANOSIM.
Where the lab's own paper genuinely did propose the method — the oral-to-gut score, and LODO
cross-validation in Pasolli et al. 2016 — citing it is correct. That is the exception, not the pattern.
Before you start
Read CONTRIBUTING.md and
PROTOCOL_STANDARD.md.
The format is defined in the standard repo, not this one. Protocols are prose, not code: they say what to do
and why, precisely enough that two people — or two agents, in two languages — get the same answer. The existing
independent-filtering-variance
protocol is the model to imitate for tone and level of detail.
Validate locally before opening the PR:
git clone https://github.com/waldronlab/agent-protocol-standard.git
Rscript agent-protocol-standard/scripts/validate-protocol.R protocols
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Read CONTRIBUTING.md and PROTOCOL_STANDARD.md first, then use protocols/independent-filtering-variance/protocol.md as the format model. Trace the method and output expectations from BugSigDBPaper/vignettes/Figure2.Rmd, establish the primary citation and taxonomy pinning, and document the acceptance criteria. Validate the finished protocol with the provided Rscript command.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- r
- Domain
- bioinformatics, documentation
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Clearly specified
- Newbie friendliness
- 68/100