abetlen / abetlen/llama-cpp-python

`Llama.from_pretrained` should work with `HF_HUB_OFFLINE=1`

Aberta
#1,801 0 comentários 0 reações 0 responsáveis Ver no GitHub
Linguagem predominante
Python
Estrelas
10.6k
Forks
1.4k
Métricas de merge de PRs
Métricas de PR pendentes

Descrição

**Is your feature request related to a problem? Please describe.**
Even with a model downloaded, the package attempts a call to HF HUB, which increases the load time.

From a quick scan of the logic [here](https://github.com/abetlen/llama-cpp-python/blob/7c4aead82d349469bbbe7d8c0f4678825873c039/llama_cpp/llama.py#L2268-L2299), it seems that the code just wants to check that the filename provided is in the repo provided.

**Describe the solution you'd like**
If you skipped that check and just assumed that the file existed and called `hf_hub_download`, that function would handle the case of errors if it couldn't find the file in the given repo.

The error may not be quite as focused, but init would run in a third the time.

On my machine:
- loading from cache takes 400ms
- loading from cache with this additional check of available files in the repo takes 1,200ms

**Describe alternatives you've considered**
The workaround is to use `from_pretrained` to download the appropriate file (if I want to do it all in Python), then get the cached file location and pass that as `model_path` to `Llama` without using `from_pretrained`.

**Additional context**
For work with HF models, I have `HF_HUB_OFFLINE=1` set by default, only turning it off when I need a new model (because a few HF operations like to make checks for model info that require network requests, even with cache primed). It would be great if this was compatible with `llama-cpp-python`.

Side note: I just started using this today and was delighted with how easy it was to install, with CUDA support, from a single pip command. Nice work.

Guia de contribuição

Abrir o guia de contribuição

Avaliação

Esta issue ainda não foi avaliada.

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.