waldronlab / waldronlab/agent-protocols
Protocol: Monte Carlo critical frequency for recurrent taxa across signatures
Nobody has claimed this yet.
- Dominant language
- No language data
- Stars
- 0
- Forks
- 1
- Avg merge
- 1h 2m
- Merged PRs (30d)
- 9
Description
Tier: A (fieldwork, slightly harder) · Type: atomic · Category: Statistical Analysis
What
Estimate, by simulation, how often the most frequently reported taxon would be expected to recur across a set of signatures under a null in which taxa are drawn from a relevant background at their background frequencies, matching the number and length of the observed signatures. Report the 1 − alpha quantile as a critical threshold, with a bootstrap CI.
Why it matters
The naive "top 10 most frequent taxa" table has no null. This is what makes a recurrence claim testable, and it correctly accounts for the fact that common taxa recur because they are common. Getting the background right — which signatures define the universe — is the substantive decision and the part a protocol needs to pin down.
Source material
waldronlab/bugSigSimple—R/montecarlo.R(frequencySigs,simulateSignatures,.countBug,getCriticalN,plotCriticalN); tests intests/testthat/test-getCriticalN.R,test-simulateSignatures.Rvignettes/capstoneanalysis_clare.rmd— "Monte-Carlo simulation for increased abundance taxa"vignettes/goldstandard_vignette_peace.Rmd— "Monte Carlo critical N"
Scope
In: choosing the background signature set; matching simulated signature count and lengths to observed; sampling with/without replacement within a signature; number of simulations and the bootstrap CI; the comparison rule (note: >=, not > — the vignette carries a comment fixing exactly this off-by-one).
Out: the direction test (separate protocol).
Frontmatter starting point
type: "atomic"
category: "Statistical Analysis"
citation: "" # The method is a Monte Carlo / randomization test. Candidates for the primary
# source: Metropolis & Ulam 1949 (10.1080/01621459.1949.10483310) for Monte Carlo;
# Fisher 1935 or Pitman 1937 for randomization testing. Establish which this
# protocol actually implements and cite that, NOT the BugSigDB paper.
tags: [bugsigdb, monte-carlo, permutation, recurrence, null-distribution]
Acceptance criteria
- The background-set choice is a required, recorded parameter with stated guidance
- The threshold comparison is stated as
>=and justified - Default
nsim/alpha/CI settings given, with a note on when to raisensim
Cite the method's origin, not its users
PROTOCOL_STANDARD.md is explicit: an atomic protocol carries "strictly 1 citation... corresponding
to the primary literature where the method was originally published." Find the paper that proposed
the method. Do not cite a paper that merely applied it — including the BugSigDB and curatedMetagenomicData
papers, which are the source of the analysis these protocols were extracted from but almost never the
source of the method.
Tracing a method back to its first publication is real work, and it is part of the task, not a
formality. Three things to expect:
- Candidate DOIs in this issue are leads, not answers. Anything marked VERIFY has not been checked.
- Some methods predate modern citation practice or have no single identifiable origin. If that is
genuinely the case, say so in the pull request rather than reaching for a convenient recent paper.
Raise it as an issue inwaldronlab/agent-protocol-standard— the standard may need a way to express
"classical method, no primary source". - If you cannot name one paper that proposed everything the protocol does, it is more than one
protocol. That test has now split four protocols out of this batch: enrichment into three methods,
filtering from transformation, LODO from random forest, and PERMANOVA from ANOSIM.
Where the lab's own paper genuinely did propose the method — the oral-to-gut score, and LODO
cross-validation in Pasolli et al. 2016 — citing it is correct. That is the exception, not the pattern.
Before you start
Read CONTRIBUTING.md and
PROTOCOL_STANDARD.md.
The format is defined in the standard repo, not this one. Protocols are prose, not code: they say what to do
and why, precisely enough that two people — or two agents, in two languages — get the same answer. The existing
independent-filtering-variance
protocol is the model to imitate for tone and level of detail.
Validate locally before opening the PR:
git clone https://github.com/waldronlab/agent-protocol-standard.git
Rscript agent-protocol-standard/scripts/validate-protocol.R protocols
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Read CONTRIBUTING.md and PROTOCOL_STANDARD.md first, then inspect the cited R functions, tests, and Monte Carlo vignettes. Use protocols/independent-filtering-variance/protocol.md as the style model and validate with the provided Rscript command. Done means the protocol records the background choice, comparison rule, simulation and CI defaults, and one verified primary citation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- r
- Domain
- documentation
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Clearly specified
- Newbie friendliness
- 55/100