NVIDIA-BioNeMo / NVIDIA-BioNeMo/BioNeMo-Inference-Runtime

protenix-v2 has no end-to-end pipeline

Open
#5 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
60
Forks
4
Avg merge
1d 17h
Merged PRs (30d)
1

Description

Summary

protenix-v2 is listed as a supported model, but there is no end-to-end pipeline for it: build_processor(model_source="protenix-v2") cannot fold from sequence input the way Boltz-2 / OpenFold3 / AlphaFold2 can. Only the optimized model forward is wired; the surrounding stages are stubs.

What I found

Using the same EngineProcessorConfigbuild_processor path that folds Boltz-2 end-to-end, Protenix-v2 does not run: ProtenixFactory's tokenizer / feature_factory / postprocessor raise NotImplementedError (the model/trunk+diffusion is the only piece implemented). By contrast Boltz-2, OpenFold3 and AF2 have all stages wired and fold from a sequence request out of the box.

If you instead try to drive the optimized Protenix model directly, the forward expects an input_feature_dict on a different schema than a naive OSS ByteDance Protenix feature dump produces — e.g. it reads keys like d_lm / v_lm and drops profile / deletion_mean. So even the forward-only path needs an (undocumented) feature-schema conversion, which makes it hard to reproduce the reported Protenix speedups end-to-end.

Reproduce
from bionemo_ir.registry import register_all_factories; register_all_factories()
from bionemo_ir.pipeline.processor.engine_proc import EngineProcessorConfig, build_processor

cfg = EngineProcessorConfig(
    model_source="protenix-v2",
    runtime_args={"diffusion_samples": 1, "num_sampling_steps": 200, "recycling_steps": 10},
    writer_stage={"output_path": "/out", "format": "cif"},
)
proc = build_processor(cfg)   # protenix-v2: tokenizer/feature_factory/postprocessor NotImplementedError

(The identical pattern with model_source="boltz-2" folds fine.)

Ask
  1. Is an end-to-end Protenix-v2 pipeline planned (tokenizer + featurizer + postprocessor wired into the stage framework, like Boltz-2)?
  2. If not, would a PR porting the OSS ByteDance Protenix featurization into the stage framework be welcome? Happy to contribute if the direction is wanted.
  3. Separately, could the expected input_feature_dict schema for the optimized Protenix forward (the d_lm/v_lm keys) be documented, so the forward can be exercised against an OSS feature dump in the meantime?
Environment

bionemo-ir 0.1.0, Python 3.12, CUDA 13.2, driver 580.95.05, NVIDIA L40S (Modal). Boltz-2 end-to-end works in the same environment.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the reproduced EngineProcessorConfig → build_processor path and inspect ProtenixFactory alongside the working Boltz-2 pipeline. Confirm the scope with maintainers before changing the tokenizer, feature_factory, or postprocessor, since the issue asks whether this direction is wanted. Done would mean a sequence request runs end to end; separately, the expected d_lm/v_lm input schema would be documented if that work is accepted.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.