galaxyproject / galaxyproject/brc-analytics
Summer 2026 Roadmap
- Dominant language
- TypeScript
- Stars
- 7
- Forks
- 11
- Avg merge
- 2d 12h
- Merged PRs (30d)
- 16
Description
# BRC.analytics summer roadmap — May–August 2026
Claude is great, but it is overly wordy. So the first section of this document is written by human.
What needs to be done ASAP
1. Pangenome reconstruction for several key species. We will start with P. vivax. This will involve the following steps: find high quality assemblies and compute their divergence estimates, create pangenome, map Pv4 dataset reads against the pangenome graph, create variant projection for every vivax assembly. Will require creating of separate tracks for each UCSC browser involved, and also alignmnet projections fro .PAF generated by PGGB. @nekrut
2. Massive paper replication. Use Orbit in conjunction with BRC-MCP and Galaxy-MCP to reproduce results from manuscripts listed below @nekrut / @scottcain
3. LexicMap / Logan: test existing deployment. Redeploy original interface of TACC resources. @d-callan / @kbeavers / @NoopDog
4. AI agent - WE NEED TO GET THIS DONE. It cannot be a prototype anymore. polish. Transfer effort to @CC (@dannon / @NoopDog )
5. Scraping VeuPathDb for as much as we can get out of it (@d-callan , @maximilianh )
6. E2E polish and ability to configure inputs from AI agent (@mvdbeek / @NoopDog )
7. Deploy @jdavcs SRA metadata DB on BRC
## 1. Pangenome reconstruction + comparative genomics
### Context
Tried PGGB and minigraph-cactus head-to-head on **simulated** 1 Mb haploids diverged from a reference at 1/2/5/10/15 % (six haplotypes, two chromosomes each). Both graphs recovered the expected collinearity; PGGB's GFA was richer (2.2M lines) than MC's (1.7M lines) for the same input.
### Linked issues
- **#1153** *Comparative Genomics Workflow in BRC Analytics* — user-facing entry point; first-iteration tracker.
- **#1160** *Scale Genome Assembly Page for Comparative Genomics*.
- **#991** *integrate host-pathogen interaction workflows* — stepper picks two assemblies from different organisms.
- **#1168** *[GA2] VGP Phase 1 Publication Readiness* — sibling assembly track.
- **#736** *Gene models and community annotations to UCSC*.
### Pipeline
```
High-quality genomes ─▶ sourmash (k=21, scaled=1000) pairwise distance
─▶ pick haplotypes with reasonable distances (mash d < 0.10 by default)
─▶ PGGB (-p 90 -s 5000 -n N -t 18)
─▶ vg giraffe (index → short-read map)
─▶ vg deconstruct / vg surject for variant projection back to a chosen reference
```
PGGB is the default constructor — richer graph, better preserved structural variants in the simulated runs. MC is the fallback when PGGB times out or when the input is closely related (PGGB on hg002.mat vs hg002.pat ran a week with kegAlign, per the alignments-channel thread; expect similar pain on Pf strains).
### Target organisms — summer
- *P. vivax* — assemblies "actually quite fragmented" per the 2026-05-01 pangenome thread
- *P. falciparum* — tighter strains; longdust + masking pre-pass likely needed
- *Candida auris* — five clades; pangenome surfaces clade-specific resistance loci
- *Aspergillus fumigatus* — azole resistance hotspots; clinical relevance
### Sub-tasks
1. **Genome curation** — pull the existing high-quality assemblies per organism, run sourmash, build the distance heatmap, pick haplotypes. Re-use `udt1/sourmash_distances.py` and `udt1/heatmap_distances.py`.
2. **PGGB build** — Galaxy workflow once exact details settle.
3. **Variant projection** — `vg deconstruct` to produce a VCF on each haplotype path. Land in BRC.analytics as a "pangenome-projected variants" track.
4. **RNA-seq mapping onto the graph** — `vg giraffe` with paired-end RNA reads, then `vg surject` to a reference for downstream `featureCounts`/`salmon`. Worth a short benchmark vs STAR-on-linear baseline. Open question: which read sets — start with the MalariaGEN Pf RNA-seq subset.
5. **Pangenome view in BRC.analytics** — new entity type `pangenomes` in `site-config/`; new route `/data/pangenomes/[id]`; viewer slot reuses Vega-Lite for the synteny strip + an `odgi viz` PNG embed for the graph. UCSC hub gets a "haplotype track" set so external users can liftOver against the linear reference.
6. **Comparative Genomics page (#1153)** — UCSC-style two-assembly form; submission triggers the Galaxy alignment workflow.
7. **Documentation + Galaxy training material**
8.
---
## 2. Paper replication across BRC phylogenetic branches
### Context
`microbs/REPORT.md` curates 17 high-impact 2023–2025 studies with open SRA/ENA data and clean experimental designs — 5 bacteriology, 4 virology, 4 mycology, 4 parasitology. Each one is a candidate for end-to-end reproduction inside BRC.analytics → Galaxy, demonstrating the platform's breadth.
### Linked issues
- **#740** *C. auris RNASeq datasets to process* — 4 transcriptome sets already linked from FungiDB JBrowse; ready inputs for mycology rep + C. auris pangenome.
- **#809** *Prioritize a workflow: Fungal RNAseq*.
- **#747** *Integration of Influenza analysis workflow*.
- **#989** *integrate metagenomic workflows* — fits Li / Guo / Odenwald.
- **#991** *integrate host-pathogen interaction workflows*.
- **#1059** *Add E2E differential expression workflow config to BRC* — see also §6.
### Papers
**Bacteriology (5)**
- *Li, Cell Host Microbe 2024* — gut *Faecalibacterium*/*Eubacterium* tryptophan catabolism, PRJNA1095001. Small WGS dataset, clean metabolomics pedigree.
- *Guo, Cell Host Microbe 2023* — ME/CFS dysbiosis metagenomics + plasma metabolomics, PRJNA751448 (197 samples).
- *Müller, mBio 2023* — *Acinetobacter baumannii* carbapenem resistance, PRJEB27899 (314 isolates, OXA-23). Companion snakemake repo.
- *Lutz, Nat Microbiol 2023* — AMF + soil pathogen amplicon, PRJEB53587 / PRJEB56590; ML yield prediction.
- *Odenwald, Nat Microbiol 2023* — *Bifidobacterium* lactulose fermentation in cirrhosis, PRJNA912122 / PRJNA838648 (919 runs).
**Virology (4)**
- *Taylor & Starr, PLOS Pathog 2023* — SARS-CoV-2 RBD deep mutational scanning, PRJNA770094 (858 runs). Full snakemake on GitHub.
- *Wang Y, mBio 2023* — SARS-CoV-2 in NYC rats, PRJNA924317 (26 amplicon-tiled runs).
- *Zhou, mBio 2023* — SARS-CoV-2 defective viral genomes, PRJNA690577 + 3 companions (218 runs, short+long read).
- *Min, PLOS Pathog 2023* — IAV circVAMP3 restriction, PRJNA750806 (4 RNase R + RNA-seq runs).
**Mycology (4)**
- *Lankiewicz, Nat Microbiol 2023* — anaerobic *Neocallimastigomycota* transcriptomics, SRP288871–SRP288885 (15 RNA-seq runs).
- *Wang X, mBio 2023* — *C. albicans* + oral carcinogenesis, PRJNA944155 / PRJNA944176 (15 RNA-seq runs).
- *Krespach, Nat Microbiol 2023* — *Streptomyces*–*Aspergillus*/*Penicillium* cross-kingdom BGC induction, PRJNA830323 (2 WGS runs + chemistry).
- *Sauters, PLOS Pathog 2023* — *Cryptococcus* population genomics + amoeba predation, PRJNA932005 (384 isolates).
**Parasitology (4)**
- *Díaz-Viraqué, Nat Microbiol 2023* — *T. cruzi* chromatin MNase-seq, PRJNA665060 (6 runs).
- *Tulloch, PLOS Pathog 2024* — *L. donovani* CYP51-inhibitor resistance, PRJNA994719 (9 BGISEQ runs).
- *Bracha, Nat Microbiol 2024* — *T. gondii* engineered protein delivery to neurons, PRJNA934842 (6 RNA-seq runs).
- *Beesley, PLOS Pathog 2023* — *F. hepatica* triclabendazole resistance, PRJEB50899 (29 WGS + bulk-segregant).
### Sub-tasks per paper
1. Pull SRA/ENA reads via the existing `fetch raw/primary data information from sra` machinery in BRC.analytics (CHANGELOG v0.6.0).
2. Wire the closest matching Galaxy IWC workflow (update `workflows.yml` and `workflow-assembly-mappings.json`).
3. Run end-to-end, validate against the published headline number.
4. Land a one-page reproducibility note in `microbs/replication/_.md`.
---
## 3. LexicMap / Logan
### Context
TACC-hosted Logan-style k-mer search for BRC pathogens — both as a public clone of logan-search.org and as the backend for the "Logan tracks" pilot on Pv mitochondrion (2025-04-17 thread; ask was for SRA-accession mouseovers and possibly k-mer-mutation-driven tracks). Three-part task per **#803**: indices, Galaxy wrap, anonymous frontend.
### Linked issues
- **#803** *Integrate LexicMap / Logan* — three sub-tasks (indices, Galaxy wrap via tools-iuc#7106, anonymous frontend).
- **#804** *Integrate LoganSearch*.
### Sub-tasks
1. **Index deployment on Vista** — fetch LexicMap / Logan indexes for the BRC priority taxa (Pv, Pf, C. auris, A. fumigatus, mpox, IAV) and stage on Vista.
2. **logan-search.org clone** — fork the front-end, repoint at TACC indexes.
3. **Hantavirus test** — small-genome end-to-end run as the canary.
4. **Track integration** — Logan-hit BED tracks land in the UCSC BRC hub with SRA-accession + biosample_title mouseovers.
---
## 4. AI / agentic tooling
### Context
ASTMH 2026 talk slot 2 is "Using agentic AI tools in pathogen research." Anchored on the in-flight analysis-assistant work (PR #1212) and the Q1 priority line "Starting to test BRC Agent on JetStream." Skill collections — not MCPs — are the right surface for our biologist users (per the 2026-05-06 thread).
### Linked issues / PRs
- **PR #1212** *analysis assistant stepper handoff and many UX improvements*.
- **PR #1262** *pydantic-evals harness for AI-enabled services*.
- **PR #1265** *prompt-injection hardening (audit follow-up to #1212)*.
- **PR #1269** *docs: add MCP server page to Learn section*.
- **#765** *LLM prompt on the home page*.
- **#766** *LLM prompt for SRA picker*.
- **#943** *Light-touch exploration of Python AI frameworks*.
### Sub-tasks
1. **Agent on JetStream** — staging deployment; eval harness (PR #1262) wired to CI.
2. **Prompt-injection hardening** — close PR #1265; threat-model write-up.
3. **Skills catalogue for BRC** — port relevant items from https://github.com/aipoch/ collection (per 2026-05-06 thread); make them the recommended surface over MCPs for researchers.
4. **SRA-picker LLM prompt** — close #766; user-test with a biologist before merge.
5. **Home-page assistant prompt** — close #765.
### Deliverables (end of July)
- BRC Agent on JetStream with eval harness gating each release
- Skills catalogue page in `/learn/`
- Two LLM-driven UX touchpoints (home, SRA picker) shipped
- ASTMH talk draft
## 5. Differential expression + organism-scoped workflow UX
### Context
Differential expression is the most-asked-for analysis pattern and the closest thing to a "demo workflow" for VEuPathDB refugees. The organism-scoped workflow stepper (Epic **#1200**) is the user-facing form that drives it; the DE-specific bits ride on top.
### Linked issues
- **#1200** *Epic: Support Organism-Scoped Workflows with Collection Inputs* — anchor.
- **#1105** *differential expression e2e*.
- **#1103** *make DE stepper play nice with SRA browser*.
- **#1102** *support intersection in primary contrasts*.
- **#1059** *Add E2E differential expression workflow config to BRC*.
- **#1007** *rna-seq workflow with optional DE in iwc*.
- **#1143** *DE workflow stepper improvements*.
- **#1213** *provide template samplesheet for DE*.
- **#1101** *help users classify columns*.
- **#1100** *md5 support in samplesheets*.
- **#1244 / #1243 / #1261 / #1263 / #1267** — organism-scoped stepper UI/nav.
- **#1230 / #1229** — LMLS docs + validation.
- **#1145** *Workflows List Page*.
- **#723 / #726 / #733 / #762 / #712** — ENA picker work.
- **#996** *Allow user to select any sequencing runs for a workflow*.
- **#724** *enable more annotations workflows*.
### Sub-tasks
1. **DE e2e** — wire end-to-end DE on at least one fungal RNA-seq set (#809, #740 C. auris) and one protist set (Pf).
2. **Stepper polish** — collection inputs, samplesheet templates, intersection contrasts, column classification.
3. **ENA picker** — study-prominence restructure, text search, large-bioproject handling.
4. **Workflows List page** — consolidated view; ties into #1145 + #810 reports.
### Deliverables (end of July)
- DE e2e working for C. auris + Pf, gallery linked from §3
- Organism-scoped stepper Epic closed
- Workflows List page shipped
## 6. VEuPathDB scraping — preempt closure, beat the UI
### Context
EuPathDB is "rather weird to use" and we can attract biologist users by surfacing better gene-centric summary pages now.
### Linked issues
- **#1131** *Compare ViewPath GTFs with existing UCSC browser tracks* — ~400 assemblies, ~20-min script; related #1057.
- **#859** *veupath gene models as workflow inputs* — grab GFF from VEuPathDB downloads on a cadence; host at UCSC; restricted to the ~600 assemblies identical to NCBI.
- **#860** *format workflow results uploadable as veupath user datasets* — bidirectional bridge.
- **#990** *introduce gene sets* — natural follow-on to the new gene page.
- **#1192** *Epic: Database Backing the Catalog Pipeline* — see §5; the substrate that holds the scraped data.
### Scope
We prioritise **PlasmoDB (Pf, Pv), FungiDB (C. auris, Aspergillus), and ToxoDB** for the summer. Other portals (CryptoDB, GiardiaDB, TriTrypDB, AmoebaDB, MicrosporidiaDB, PiroplasmaDB, TrichDB, HostDB, VectorBase) follow in the next iteration.
### Plan
1. **Schema reconnaissance** — produce a one-page ERD covering the gene, transcript, expression, SNP, and ortholog regions. Source repo: https://github.com/VEuPathDB/GusSchema. Reference the full-VM mirror recipe at https://github.com/VEuPathDB/offprint as the deployable baseline.
2. **Flat-file mirror** — wget the `Current_Release/` tree for the priority portals (PlasmoDB, FungiDB, ToxoDB) and stash the canonical FASTA, GFF, GAF (GO annotations), CodonUsage, GeneAliases, and the per-isolate VCF/CNV subdirs. Compare against the archive.org VectorBase 68 snapshot (https://archive.org/details/vector-base-68) as a known-good baseline.
3. **WDK REST harvest** — for each priority organism, hit `POST /record-types/gene/searches/.../reports/standard` to pull derived attributes the flat files miss (functional summary, ortholog group IDs, phenotype links). Cache responses in our backend. OpenAPI spec: https://veupathdb.org/service-api.html.
4. **Gene-page surface in BRC.analytics** — new `genes` entity type in `site-config/brc-analytics/`, new catalog JSON in `catalog/output/genes.json`, new route `/data/genes/[geneId]`. Mock-up sections: overview (synonyms, location, product), expression heatmap (Vega-Lite, already in `package.json`), variation summary, orthologs, "launch in Galaxy" hooks. Beat the VEuPathDB layout by collapsing redundant tabs and putting the high-traffic items (expression heatmap, function, paralogs, drug-resistance flag) above the fold.
5. **UCSC track import** — pipe per-gene resequencing VCFs and expression bigWigs into the UCSC BRC assembly hubs (https://hgdownload.gi.ucsc.edu/hubs/BRC/). Tracks the JBrowse2 mirroring tracked in galaxyproject/brc-analytics#129.
6. **Galaxy bridge** — surface "Re-run upstream analysis in Galaxy" buttons on each gene page pointing at the VEuPathDB-style RNA-seq pipelines we already host (https://galaxyproject.org/use/veupathdb/).
## 7. Deploy SRA metadata DB on BRC
### To do
1. Expand field subset (60 -> 100+)
1. Improve indexing and document analysis: per-field tokenization rules, stoplists, etc.
1. Decide on update cadence; automate process
1. Integrate with repo; secure, harden deployment
1. Polish code & documentation
Ref:
https://github.com/jdavcs/solr-playbook
https://github.com/jdavcs/ena-metadata-util
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.