abetlen / abetlen/llama-cpp-python

How can I extract data from documents in JSON Output Format?

Open
#942 0 comments 0 reactions 0 assignees View on GitHub
question
Dominant language
Python
Stars
10.6k
Forks
1.4k
PR merge metrics
PR metrics pending

Description

Hello,

I cannot find how can I extract data from documents in a JSON Output Format?

Just like https://hackernoon.com/unlocking-structured-json-data-with-langchain-and-gpt-a-step-by-step-tutorial for OpenAI.

I have tried to prompt through the following lines:

```
document_query = "Crear un esquema basado en este texto: " + document[0].page_content
_input = prompt.format_prompt(question=document_query)
output = chat_model(_input.to_string())
print(output)
parsed = parser.parse(output)
texts_json.append(json.dumps(parsed.dict()).encode('utf-8').decode('unicode-escape'))
```

However I cannot get just the JSON output. I get long unparsed conversations (for example, human-assistant dialogues if I use vicuna ggml-vicuna-13b-4bit.bin).

Could somebody tell me if Llamacpp can already be used for those goals?

Thanks in advance.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.