waldronlab / waldronlab/agent-protocols

Protocol: derive a reference species set for a body site or sample group by prevalence

Open
#29 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

atomic-protocol data-curation microbiome
Dominant language
No language data
Stars
0
Forks
1
Avg merge
1h 2m
Merged PRs (30d)
9

Description

Tier: D (cMD paper) · Type: atomic · Category: Curation & Data Preparation

What

Given any labelled group of samples — a body site, a population, a disease cohort, an environment — select the species prevalent enough within that group to serve as its reference set, and publish the result as a versioned, citable list.

Why it matters

Generalized from the oral-specific draft, per review. The original version derived the oral species list and nothing else, which made it useful exactly once. The procedure is not oral-specific: pick a reference sample group, apply a prevalence rule, publish the list. Stated generally it supports a skin set, a vaginal set, a healthy-adult-gut set, a population-specific set — and the 305-species oral list becomes one worked example rather than the whole point.

It is also the generator for a family of "enrichment score" analyses. Once a reference set for group X exists, the scoring protocol can measure X-typical organisms in samples from elsewhere. Oral-to-gut is the instance the paper published; there is no reason it should be the only one.

Source material

  • waldronlab/curatedMetagenomicDataAnalysescMD3_paper_analyses/oral_enrichment/01-estimate_oral_enrichment.py, python_tools/oral_introgression_score.py (evaluate_oral_species)
  • Worked example: 10.1038/s41467-025-66888-1 — 305 species prevalent in ≥ 1% of oral samples

Scope

In: how the reference sample group is defined and what makes it adequate (size, diversity of source studies, whether a single cohort can ever be sufficient); the prevalence rule and threshold, and its relationship to the detection floor of the profiler used; the profiler and reference-database dependency — a list derived under one taxonomic backbone is not directly usable under another, and the protocol must say how that is recorded; the versioning and publication scheme, so that scoring protocols can pin a specific list; what must be re-validated when a list is regenerated.

Out: computing scores from the list (that is oral-to-gut-enrichment-score and its future siblings).

Frontmatter starting point

type: "atomic"
category: "Curation & Data Preparation"
citation: "10.12688/f1000research.8986.1"   # prevalence as the selection criterion; but see note
protocols_used:
  - name: "prevalence-filtering"
tags: [reference-set, prevalence, body-site, species-list, reference-data]

Citation needs thought. The prevalence criterion itself traces to Callahan et al. 2016, but using prevalence within a reference group to define a marker set for a different group may be a distinct contribution — plausibly the cMD paper's own. Establish which and say so.

Acceptance criteria

  • Stated for an arbitrary sample group; oral appears only as a worked example
  • Rerunning it on the paper's oral reference set reproduces the 305-species list
  • The profiler / reference-database dependency is recorded in the published list
  • Defines the versioning scheme that scoring protocols pin against

Cite the method's origin, not its users

PROTOCOL_STANDARD.md is explicit: an atomic protocol carries "strictly 1 citation... corresponding
to the primary literature where the method was originally published."
Find the paper that proposed
the method. Do not cite a paper that merely applied it — including the BugSigDB and curatedMetagenomicData
papers, which are the source of the analysis these protocols were extracted from but almost never the
source of the method.

Tracing a method back to its first publication is real work, and it is part of the task, not a
formality. Three things to expect:

  • Candidate DOIs in this issue are leads, not answers. Anything marked VERIFY has not been checked.
  • Some methods predate modern citation practice or have no single identifiable origin. If that is
    genuinely the case, say so in the pull request rather than reaching for a convenient recent paper.
    Raise it as an issue in waldronlab/agent-protocol-standard — the standard may need a way to express
    "classical method, no primary source".
  • If you cannot name one paper that proposed everything the protocol does, it is more than one
    protocol.
    That test has now split four protocols out of this batch: enrichment into three methods,
    filtering from transformation, LODO from random forest, and PERMANOVA from ANOSIM.

Where the lab's own paper genuinely did propose the method — the oral-to-gut score, and LODO
cross-validation in Pasolli et al. 2016 — citing it is correct. That is the exception, not the pattern.

Before you start

Read CONTRIBUTING.md and
PROTOCOL_STANDARD.md.
The format is defined in the standard repo, not this one. Protocols are prose, not code: they say what to do
and why, precisely enough that two people — or two agents, in two languages — get the same answer. The existing
independent-filtering-variance
protocol is the model to imitate for tone and level of detail.

Validate locally before opening the PR:

git clone https://github.com/waldronlab/agent-protocol-standard.git
Rscript agent-protocol-standard/scripts/validate-protocol.R protocols

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Read CONTRIBUTING.md, PROTOCOL_STANDARD.md, and protocols/independent-filtering-variance/protocol.md first. Review the cited oral-enrichment scripts and trace the primary literature for the prevalence method, then write the general protocol with the profiler/database, versioning, and regeneration requirements. Validate it with Rscript agent-protocol-standard/scripts/validate-protocol.R protocols and confirm the oral example reproduces the 305-species list.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, r
Domain
bioinformatics, data, documentation
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
65/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.