sokrypton / sokrypton/ColabFold

Runtime error since update: "PTX was compiled with an unsupported toolchain"

Open
#131 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Jupyter Notebook
Stars
2.9k
Forks
747
PR merge metrics
No merged PRs in 30d

Description

Hi,
I used colabfold on my local computer successfully (Thank you guys!)
unfortunately, after I ran the update script, I have an error and I can't run it anymore:

jax._src.traceback_util.UnfilteredStackTrace: RuntimeError: INTERNAL: CustomCall failed: jaxlib/cuda_prng_kernels.cc:30: operation cudaGetLastError() failed: the provided PTX was compiled with an unsupported toolchain.

I tried using cuda 11.1 with cudnn 8.0.5
and tried cuda 11.3 with cudnn 8.2.1
and bothe does not work.

I installed jax by the instructions an I tried to uninstall and install it again with the specific cuda and cudnn versions - but it didn't help.

I tried to build the colabfold-batch from pip on new enviroment - but it acts the same too.

I run on linux, dgx-40 nvidia machine with GPU A100 (with 24 GB memory)

error message from colabfold_batch:

If you require more MSAs, please host your own API and pass it to --host-url
2021-12-22 15:32:30,368 non-fasta/a3m file in input directory: slurm-767542.out
2021-12-22 15:32:30,403 Running colabfold 1.2.0 (633d9cae284c4c6a0e881e011d4b318193e22ba9)
2021-12-22 15:32:30,416 Found 5 citations for tools or databases
2021-12-22 15:32:52,008 Query 1/2: second (length 354)
2021-12-22 15:32:53,358 Running model_3
2021-12-22 15:36:38,277 Could not predict second. Not Enough GPU memory? INTERNAL: CustomCall failed: jaxlib/cuda_prng_kernels.cc:30: operation cudaGetLastError() failed: the provided PTX was compiled with an unsupported toolchain.
2021-12-22 15:36:38,279 Query 2/2: first (length 837)
2021-12-22 15:36:38,979 Running model_3
2021-12-22 15:39:38,164 Could not predict first. Not Enough GPU memory? INTERNAL: CustomCall failed: jaxlib/cuda_prng_kernels.cc:30: operation cudaGetLastError() failed: the provided PTX was compiled with an unsupported toolchain.
2021-12-22 15:39:38,165 Done

error message from local installation:

File "/cs/labs/dina/seanco/xl_parser/alphafold/colabfold/sean_runner.py", line 686, in run
model_runner.predict(processed_feature_dict, random_seed=seed), "cpu")
File "/cs/labs/dina/seanco/xl_parser/alphafold/colabfold/alphafold/model/model.py", line 134, in predict
result, recycles = self.apply(self.params, jax.random.PRNGKey(random_seed), feat)
File "/cs/labs/dina/seanco/xl_parser/xl_db_parser_venv/lib/python3.7/site-packages/jax/_src/dispatch.py", line 436, in _execute_compiled
out_bufs = compiled.execute(input_bufs)
RuntimeError: INTERNAL: CustomCall failed: jaxlib/cuda_prng_kernels.cc:30: operation cudaGetLastError() failed: the provided PTX was compiled with an unsupported toolchain.

Any idea what went wrong?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing colabfold_batch on the reported Linux A100 setup, using the stated CUDA, cuDNN, JAX, and jaxlib versions. Inspect the update script and the JAX CUDA runtime compatibility around the reported cuda_prng_kernels.cc failure. Done means identifying a reproducible compatibility boundary or documenting the supported environment and workaround.

Written by the indexing model from the issue text.

Assessment

Tech stack
linux, python
Domain
bioinformatics, machine-learning, operating-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.