abetlen / abetlen/llama-cpp-python

Generate answer from embedding vectors

未關閉
#1,897 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
主要語言
Python
星號
10.6k
分支
1.4k
PR 合併指標
PR 指標待擷取

描述

Hi, I'm not familiar with llama-cpp-python (actually not familiar with cpp) but I have to use gguf model for my project.

I want to generate answer from pre-computed embedding vectors(torch.Tensor) with size (1, n_tokens, 4096), not from query text. Here I mean the embedding vectors are text embeddings that generated from torch.nn.Embedding()
(Just like inputs_embeds argument of generate() function of transformers model)

What I want to do is just skip process 1 and 2:
1. tokenize input string
2. make text embeddings from tokens
3. model inference
4. get output token
5. detokenize

Is this feature already implemented? If not, please anyone help me where should I begin.

貢獻指南

開啟貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。