abetlen / abetlen/llama-cpp-python

meteor lake support not functionnal

Đang mở
#1,231 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
Python
Star
10.6k
Fork
1.4k
Chỉ số merge pull request
Chỉ số pull request đang chờ

Mô tả

i tried running python llamacpp on meteor lake core ultra 9 device on wsl with ubuntu 22.04 but observing some strange behavior

Expected behavior:
- build with sycl works correctly
- all devices are detected by llama cpp (Cpu, Npu and GPU as opencl, devices + GPU as sycl device)
- I am able to load gguf models (tinyllama Q4) on all devices
- run on cpu works well (around 58 tokens/s)

Odd behaviour:
- when loading the model on gpu, it suddently become very slow, with around 1.5 tokens/s

Any clue on what makes gpu run so slow?

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.