NVIDIA-BioNeMo / NVIDIA-BioNeMo/BioNeMo-Inference-Runtime
protenix-v2 has no end-to-end pipeline
まだ誰も着手していません。
- 主要言語
- Python
- スター
- 60
- フォーク
- 4
- 平均マージ
- 1日 17時間
- マージ済み PR(30日)
- 1
説明
Summary
protenix-v2 is listed as a supported model, but there is no end-to-end pipeline for it: build_processor(model_source="protenix-v2") cannot fold from sequence input the way Boltz-2 / OpenFold3 / AlphaFold2 can. Only the optimized model forward is wired; the surrounding stages are stubs.
What I found
Using the same EngineProcessorConfig → build_processor path that folds Boltz-2 end-to-end, Protenix-v2 does not run: ProtenixFactory's tokenizer / feature_factory / postprocessor raise NotImplementedError (the model/trunk+diffusion is the only piece implemented). By contrast Boltz-2, OpenFold3 and AF2 have all stages wired and fold from a sequence request out of the box.
If you instead try to drive the optimized Protenix model directly, the forward expects an input_feature_dict on a different schema than a naive OSS ByteDance Protenix feature dump produces — e.g. it reads keys like d_lm / v_lm and drops profile / deletion_mean. So even the forward-only path needs an (undocumented) feature-schema conversion, which makes it hard to reproduce the reported Protenix speedups end-to-end.
Reproduce
from bionemo_ir.registry import register_all_factories; register_all_factories()
from bionemo_ir.pipeline.processor.engine_proc import EngineProcessorConfig, build_processor
cfg = EngineProcessorConfig(
model_source="protenix-v2",
runtime_args={"diffusion_samples": 1, "num_sampling_steps": 200, "recycling_steps": 10},
writer_stage={"output_path": "/out", "format": "cif"},
)
proc = build_processor(cfg) # protenix-v2: tokenizer/feature_factory/postprocessor NotImplementedError
(The identical pattern with model_source="boltz-2" folds fine.)
Ask
- Is an end-to-end Protenix-v2 pipeline planned (tokenizer + featurizer + postprocessor wired into the stage framework, like Boltz-2)?
- If not, would a PR porting the OSS ByteDance Protenix featurization into the stage framework be welcome? Happy to contribute if the direction is wanted.
- Separately, could the expected
input_feature_dictschema for the optimized Protenix forward (thed_lm/v_lmkeys) be documented, so the forward can be exercised against an OSS feature dump in the meantime?
Environment
bionemo-ir 0.1.0, Python 3.12, CUDA 13.2, driver 580.95.05, NVIDIA L40S (Modal). Boltz-2 end-to-end works in the same environment.
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
調査の方向性
再現された EngineProcessorConfig → build_processor のパスから始め、動作している Boltz-2 パイプラインと併せて ProtenixFactory を調査します。tokenizer、feature_factory、または postprocessor を変更する前に、maintainers とスコープを確認してください。issue ではこの方向性が望まれているかどうかを尋ねているためです。完了とは、sequence request がエンドツーエンドで実行されることを意味します。別途、この作業が受け入れられた場合は、想定される d_lm/v_lm 入力スキーマを文書化します。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python
- 領域
- machine-learning
- issue の種類
- 機能追加
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 活発さ
- 活発
- 明瞭さ
- 説明が足りない
- 初心者へのやさしさ
- 35/100