NVIDIA-BioNeMo / NVIDIA-BioNeMo/BioNeMo-Inference-Runtime
protenix-v2 has no end-to-end pipeline
Nessuno ha ancora preso questa issue.
- Lingua principale
- Python
- Stelle
- 60
- Fork
- 4
- Merge medio
- 1g 17h
- PR unite (30g)
- 1
Descrizione
Summary
protenix-v2 is listed as a supported model, but there is no end-to-end pipeline for it: build_processor(model_source="protenix-v2") cannot fold from sequence input the way Boltz-2 / OpenFold3 / AlphaFold2 can. Only the optimized model forward is wired; the surrounding stages are stubs.
What I found
Using the same EngineProcessorConfig → build_processor path that folds Boltz-2 end-to-end, Protenix-v2 does not run: ProtenixFactory's tokenizer / feature_factory / postprocessor raise NotImplementedError (the model/trunk+diffusion is the only piece implemented). By contrast Boltz-2, OpenFold3 and AF2 have all stages wired and fold from a sequence request out of the box.
If you instead try to drive the optimized Protenix model directly, the forward expects an input_feature_dict on a different schema than a naive OSS ByteDance Protenix feature dump produces — e.g. it reads keys like d_lm / v_lm and drops profile / deletion_mean. So even the forward-only path needs an (undocumented) feature-schema conversion, which makes it hard to reproduce the reported Protenix speedups end-to-end.
Reproduce
from bionemo_ir.registry import register_all_factories; register_all_factories()
from bionemo_ir.pipeline.processor.engine_proc import EngineProcessorConfig, build_processor
cfg = EngineProcessorConfig(
model_source="protenix-v2",
runtime_args={"diffusion_samples": 1, "num_sampling_steps": 200, "recycling_steps": 10},
writer_stage={"output_path": "/out", "format": "cif"},
)
proc = build_processor(cfg) # protenix-v2: tokenizer/feature_factory/postprocessor NotImplementedError
(The identical pattern with model_source="boltz-2" folds fine.)
Ask
- Is an end-to-end Protenix-v2 pipeline planned (tokenizer + featurizer + postprocessor wired into the stage framework, like Boltz-2)?
- If not, would a PR porting the OSS ByteDance Protenix featurization into the stage framework be welcome? Happy to contribute if the direction is wanted.
- Separately, could the expected
input_feature_dictschema for the optimized Protenix forward (thed_lm/v_lmkeys) be documented, so the forward can be exercised against an OSS feature dump in the meantime?
Environment
bionemo-ir 0.1.0, Python 3.12, CUDA 13.2, driver 580.95.05, NVIDIA L40S (Modal). Boltz-2 end-to-end works in the same environment.
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Direzione di ricerca
Inizia dal percorso riprodotto EngineProcessorConfig → build_processor e ispeziona ProtenixFactory insieme alla pipeline Boltz-2 funzionante. Conferma l’ambito con i maintainers prima di modificare tokenizer, feature_factory o postprocessor, poiché l’issue chiede se questa direzione è desiderata. Il lavoro sarebbe considerato completato quando una sequence request viene eseguita end-to-end; separatamente, lo schema di input previsto per d_lm/v_lm verrebbe documentato se questo lavoro viene accettato.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python
- Ambito
- machine-learning
- Tipo di issue
- Funzionalità
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Attiva
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 35/100