waldronlab / waldronlab/agent-protocols
Protocol: derive a reference species set for a body site or sample group by prevalence
Nobody has claimed this yet.
- Dominant language
- No language data
- Stars
- 0
- Forks
- 1
- Avg merge
- 1h 2m
- Merged PRs (30d)
- 9
Description
Tier: D (cMD paper) · Type: atomic · Category: Curation & Data Preparation
What
Given any labelled group of samples — a body site, a population, a disease cohort, an environment — select the species prevalent enough within that group to serve as its reference set, and publish the result as a versioned, citable list.
Why it matters
Generalized from the oral-specific draft, per review. The original version derived the oral species list and nothing else, which made it useful exactly once. The procedure is not oral-specific: pick a reference sample group, apply a prevalence rule, publish the list. Stated generally it supports a skin set, a vaginal set, a healthy-adult-gut set, a population-specific set — and the 305-species oral list becomes one worked example rather than the whole point.
It is also the generator for a family of "enrichment score" analyses. Once a reference set for group X exists, the scoring protocol can measure X-typical organisms in samples from elsewhere. Oral-to-gut is the instance the paper published; there is no reason it should be the only one.
Source material
waldronlab/curatedMetagenomicDataAnalyses—cMD3_paper_analyses/oral_enrichment/01-estimate_oral_enrichment.py,python_tools/oral_introgression_score.py(evaluate_oral_species)- Worked example: 10.1038/s41467-025-66888-1 — 305 species prevalent in ≥ 1% of oral samples
Scope
In: how the reference sample group is defined and what makes it adequate (size, diversity of source studies, whether a single cohort can ever be sufficient); the prevalence rule and threshold, and its relationship to the detection floor of the profiler used; the profiler and reference-database dependency — a list derived under one taxonomic backbone is not directly usable under another, and the protocol must say how that is recorded; the versioning and publication scheme, so that scoring protocols can pin a specific list; what must be re-validated when a list is regenerated.
Out: computing scores from the list (that is oral-to-gut-enrichment-score and its future siblings).
Frontmatter starting point
type: "atomic"
category: "Curation & Data Preparation"
citation: "10.12688/f1000research.8986.1" # prevalence as the selection criterion; but see note
protocols_used:
- name: "prevalence-filtering"
tags: [reference-set, prevalence, body-site, species-list, reference-data]
Citation needs thought. The prevalence criterion itself traces to Callahan et al. 2016, but using prevalence within a reference group to define a marker set for a different group may be a distinct contribution — plausibly the cMD paper's own. Establish which and say so.
Acceptance criteria
- Stated for an arbitrary sample group; oral appears only as a worked example
- Rerunning it on the paper's oral reference set reproduces the 305-species list
- The profiler / reference-database dependency is recorded in the published list
- Defines the versioning scheme that scoring protocols pin against
Cite the method's origin, not its users
PROTOCOL_STANDARD.md is explicit: an atomic protocol carries "strictly 1 citation... corresponding
to the primary literature where the method was originally published." Find the paper that proposed
the method. Do not cite a paper that merely applied it — including the BugSigDB and curatedMetagenomicData
papers, which are the source of the analysis these protocols were extracted from but almost never the
source of the method.
Tracing a method back to its first publication is real work, and it is part of the task, not a
formality. Three things to expect:
- Candidate DOIs in this issue are leads, not answers. Anything marked VERIFY has not been checked.
- Some methods predate modern citation practice or have no single identifiable origin. If that is
genuinely the case, say so in the pull request rather than reaching for a convenient recent paper.
Raise it as an issue inwaldronlab/agent-protocol-standard— the standard may need a way to express
"classical method, no primary source". - If you cannot name one paper that proposed everything the protocol does, it is more than one
protocol. That test has now split four protocols out of this batch: enrichment into three methods,
filtering from transformation, LODO from random forest, and PERMANOVA from ANOSIM.
Where the lab's own paper genuinely did propose the method — the oral-to-gut score, and LODO
cross-validation in Pasolli et al. 2016 — citing it is correct. That is the exception, not the pattern.
Before you start
Read CONTRIBUTING.md and
PROTOCOL_STANDARD.md.
The format is defined in the standard repo, not this one. Protocols are prose, not code: they say what to do
and why, precisely enough that two people — or two agents, in two languages — get the same answer. The existing
independent-filtering-variance
protocol is the model to imitate for tone and level of detail.
Validate locally before opening the PR:
git clone https://github.com/waldronlab/agent-protocol-standard.git
Rscript agent-protocol-standard/scripts/validate-protocol.R protocols
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Read CONTRIBUTING.md, PROTOCOL_STANDARD.md, and protocols/independent-filtering-variance/protocol.md first. Review the cited oral-enrichment scripts and trace the primary literature for the prevalence method, then write the general protocol with the profiler/database, versioning, and regeneration requirements. Validate it with Rscript agent-protocol-standard/scripts/validate-protocol.R protocols and confirm the oral example reproduces the 305-species list.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, r
- Domain
- bioinformatics, data, documentation
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Clearly specified
- Newbie friendliness
- 65/100