waldronlab / waldronlab/agent-protocols
Composite protocol: generate a hypothesis from BugSigDB, test it in individual-participant data
Nobody has claimed this yet.
- Dominant language
- No language data
- Stars
- 0
- Forks
- 1
- Avg merge
- 1h 2m
- Merged PRs (30d)
- 9
Description
Tier: B (bridge, capstone-scale) · Type: composite · Category: Statistical Analysis
What
The two-stage design: use the frequency of reported taxa in BugSigDB to nominate a small number of candidate taxa for a condition, then test those specific candidates in individual-participant data from curatedMetagenomicData using an appropriate regression model — with the candidate list fixed before the second stage begins.
Why it matters
This is the most methodologically interesting workflow in the whole set, and the one most vulnerable to being done wrong. Its validity rests entirely on the hypothesis being fixed before the individual-participant data is touched; a protocol can make that a required, recorded step instead of an informal promise. It also connects the two lab resources, which is the thing neither paper's code does cleanly.
Source material
waldronlab/bugSigSimple—vignettes/fieldworkanalysis_samara.Rmd(endometriosis; linear regression, t-test, Poisson, and zero-inflated negative binomial models, plus an explicit "Comparing mean to variance of the outcome variable" step to choose among them)
Scope
In: the nomination rule and how many candidates may be carried forward; the pre-registration of the candidate list; the count-model selection procedure (the mean/variance comparison that picks Poisson vs. negative binomial vs. zero-inflated); covariate adjustment; multiple testing across the fixed candidate set; and what to report when the second stage does not replicate.
Out: re-deriving the frequency table or the cohort assembly — those are component protocols.
Frontmatter starting point
type: "composite"
category: "Statistical Analysis"
citation: "10.1038/s41587-023-01872-y"
protocols_used:
- name: "bugsigdb-signature-subset"
- name: "taxon-frequency-direction-binomial"
- name: "cmd-cohort-assembly"
- name: "clr-transformation"
Acceptance criteria
- Requires the candidate list to be recorded before stage two, and says so as a step
- The model-selection rule is deterministic, not "inspect the plot"
- Non-replication is a reportable outcome with a defined write-up, not a reason to loop back to stage one
Cite the method's origin, not its users
PROTOCOL_STANDARD.md is explicit: an atomic protocol carries "strictly 1 citation... corresponding
to the primary literature where the method was originally published." Find the paper that proposed
the method. Do not cite a paper that merely applied it — including the BugSigDB and curatedMetagenomicData
papers, which are the source of the analysis these protocols were extracted from but almost never the
source of the method.
Tracing a method back to its first publication is real work, and it is part of the task, not a
formality. Three things to expect:
- Candidate DOIs in this issue are leads, not answers. Anything marked VERIFY has not been checked.
- Some methods predate modern citation practice or have no single identifiable origin. If that is
genuinely the case, say so in the pull request rather than reaching for a convenient recent paper.
Raise it as an issue inwaldronlab/agent-protocol-standard— the standard may need a way to express
"classical method, no primary source". - If you cannot name one paper that proposed everything the protocol does, it is more than one
protocol. That test has now split four protocols out of this batch: enrichment into three methods,
filtering from transformation, LODO from random forest, and PERMANOVA from ANOSIM.
Where the lab's own paper genuinely did propose the method — the oral-to-gut score, and LODO
cross-validation in Pasolli et al. 2016 — citing it is correct. That is the exception, not the pattern.
Before you start
Read CONTRIBUTING.md and
PROTOCOL_STANDARD.md.
The format is defined in the standard repo, not this one. Protocols are prose, not code: they say what to do
and why, precisely enough that two people — or two agents, in two languages — get the same answer. The existing
independent-filtering-variance
protocol is the model to imitate for tone and level of detail.
Validate locally before opening the PR:
git clone https://github.com/waldronlab/agent-protocol-standard.git
Rscript agent-protocol-standard/scripts/validate-protocol.R protocols
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Read CONTRIBUTING.md and PROTOCOL_STANDARD.md first, then use protocols/independent-filtering-variance/protocol.md as the style model and inspect vignettes/fieldworkanalysis_samara.Rmd for the source workflow. Validate the new protocol with Rscript agent-protocol-standard/scripts/validate-protocol.R protocols; it is done when the candidate list is fixed before stage two, model selection is deterministic, non-replication is reportable, and the primary method citation is verified.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- r
- Domain
- data, documentation
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Clearly specified
- Newbie friendliness
- 38/100