abetlen / abetlen/llama-cpp-python

setting n_gpu_layers to 0 or -1 still tries to use the gpu llamaindex

Open
#1,158 6 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
10.6k
Forks
1.4k
PR merge metrics
PR metrics pending

Description

torch.cuda.OutOfMemoryError: HIP out of memory. Tried to allocate 224.00 MiB. GPU 0 has a total capacty of 23.98 GiB of which 44.00 MiB is free. Of the allocated memory 23.68 GiB is allocated by PyTorch, and 1.14 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting max_split_size_mb to avoid fragmentation. See documentation for Memory Management and PYTORCH_HIP_ALLOC_CONF

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.