waldronlab / waldronlab/agent-protocols
Protocol: CBEA: competitive balances for taxonomic enrichment analysis against BugSigDB signatures
Nobody has claimed this yet.
- Dominant language
- No language data
- Stars
- 0
- Forks
- 1
- Avg merge
- 1h 2m
- Merged PRs (30d)
- 9
Description
Tier: C (BugSigDB paper) · Type: atomic · Category: Statistical Analysis
What
Treat BugSigDB signatures as microbe sets and test them for enrichment using CBEA, which forms an isometric log-ratio balance contrasting the microbes in a set against those outside it, giving a per-sample set score. Input: CLR-appropriate compositional abundances.
Why it matters
"Apply BugSigDB to my own data" is the most externally requested capability in the repository, and this is one of three genuinely different ways to do it.
Why this is its own protocol. The original draft bundled ORA, PADOG and CBEA into one issue with the method as a user choice. They have three different primary citations, three different null hypotheses, and three different input requirements — and PROTOCOL_STANDARD.md requires an atomic protocol to carry exactly one citation naming where the method was originally published. One protocol per method; a separate composite benchmarks them against each other, which is what Figure3.Rmd actually does.
Source material
waldronlab/BugSigDBPaper—vignettes/Figure3.Rmd("Obtain microbe signatures from BugSigDB", "Harmonize taxonomic level", "Enrichment analysis")- Applied in: Geistlinger et al., Nat Biotechnol 2023, 10.1038/s41587-023-01872-y — cite this only as an application, never as the method's source
Scope
In: taxonomic-level harmonization of query and sets before anything else; minimum and maximum set size; multiple testing across sets; what a significant result licenses you to conclude. Method-specific: That CBEA is compositional by construction and therefore interacts with the choice of transformation upstream; that it produces a per-sample score, which the other two do not, and what that enables; the pooled-versus-per-dataset distinction the vignette draws.
Out: producing the differential-abundance input; comparing against the other two methods (that is the benchmark composite).
Frontmatter starting point
type: "atomic"
category: "Statistical Analysis"
citation: "" # Nguyen et al. — establish the correct primary reference; do not guess a DOI
tags: [enrichment, microbe-sets, cbea, bugsigdb]
Acceptance criteria
- Cites the paper that proposed the method, not the BugSigDB paper
- Rank harmonization is a required step with defined behavior for unmatched taxa
- The universe/background and set-size bounds are explicit
- States what the test's null hypothesis actually is, in words
Cite the method's origin, not its users
PROTOCOL_STANDARD.md is explicit: an atomic protocol carries "strictly 1 citation... corresponding
to the primary literature where the method was originally published." Find the paper that proposed
the method. Do not cite a paper that merely applied it — including the BugSigDB and curatedMetagenomicData
papers, which are the source of the analysis these protocols were extracted from but almost never the
source of the method.
Tracing a method back to its first publication is real work, and it is part of the task, not a
formality. Three things to expect:
- Candidate DOIs in this issue are leads, not answers. Anything marked VERIFY has not been checked.
- Some methods predate modern citation practice or have no single identifiable origin. If that is
genuinely the case, say so in the pull request rather than reaching for a convenient recent paper.
Raise it as an issue inwaldronlab/agent-protocol-standard— the standard may need a way to express
"classical method, no primary source". - If you cannot name one paper that proposed everything the protocol does, it is more than one
protocol. That test has now split four protocols out of this batch: enrichment into three methods,
filtering from transformation, LODO from random forest, and PERMANOVA from ANOSIM.
Where the lab's own paper genuinely did propose the method — the oral-to-gut score, and LODO
cross-validation in Pasolli et al. 2016 — citing it is correct. That is the exception, not the pattern.
Before you start
Read CONTRIBUTING.md and
PROTOCOL_STANDARD.md.
The format is defined in the standard repo, not this one. Protocols are prose, not code: they say what to do
and why, precisely enough that two people — or two agents, in two languages — get the same answer. The existing
independent-filtering-variance
protocol is the model to imitate for tone and level of detail.
Validate locally before opening the PR:
git clone https://github.com/waldronlab/agent-protocol-standard.git
Rscript agent-protocol-standard/scripts/validate-protocol.R protocols
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Read CONTRIBUTING.md and PROTOCOL_STANDARD.md first, then inspect vignettes/Figure3.Rmd for the BugSigDB enrichment context. Identify and verify the primary paper proposing CBEA, then write the atomic protocol with the stated harmonization, background, set-size, null-hypothesis, and citation requirements. Validate with the provided Rscript command before opening a PR.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- r
- Domain
- data, documentation
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Clearly specified
- Newbie friendliness
- 64/100