NVIDIA / NVIDIA/CUDALibrarySamples

Trying to compress LLM models using nvcomp

Open
#216 3 comments 0 reactions 1 assignee View on GitHub

@naveenaero is already working on this.

Since Jun 27, 2025.

nvCOMP
Dominant language
Cuda
Stars
2.5k
Forks
478
PR merge metrics
No merged PRs in 30d

Description

Hi all, I've been trying to compress Large language models using nvcomp but can't succeed. I only managed to compress the tokenizer.json and config.json files of the model but was unable to compress the .safetensors or .gguf model files.

Does nvcomp currently support this? Can I ask how can i do so?

Much appreciated

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.