abetlen / abetlen/llama-cpp-python

openai API `max_completion_tokens` argument is ignored

Open
#1,907 0 comments 3 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
10.6k
Forks
1.4k
PR merge metrics
PR metrics pending

Description

# Prerequisites

Please answer the following questions for yourself before submitting an issue.

- [X] I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- [X] I carefully followed the [README.md](https://github.com/abetlen/llama-cpp-python/blob/main/README.md).
- [X] I [searched using keywords relevant to my issue](https://docs.github.com/en/issues/tracking-your-work-with-issues/filtering-and-searching-issues-and-pull-requests) to make sure that I am creating a new issue that is not already open (or closed).
- [X] I reviewed the [Discussions](https://github.com/abetlen/llama-cpp-python/discussions), and have a new bug or useful enhancement to share.

# Current Behavior

I'm running llama-server with following command:
```
python3 -m llama_cpp.server --model models/mys/ggml_llava-v1.5-13b/ggml-model-q4_k.gguf --clip_model_path models/mys/ggml_llava-v1.5-13b/mmproj-model-f16.gguf --model_alias llava-v1.5-13b-q4_k --chat_format llava-1-5 --port 10322
```

(models downloaded from https://huggingface.co/mys/ggml_llava-v1.5-13b/tree/main)

When I call the server using openai python package:
```python
from openai import OpenAI

client = OpenAI(
base_url="http://localhost:10322/v1", # "http://:port"
api_key = "sk-no-key-required"
)

chat_completion = client.chat.completions.create(
model="models/mys/ggml_llava-v1.5-13b/ggml-model-q4_k.gguf",
messages=[
{"role": "user", "content": "Write a limerick about python exceptions"}
],
max_tokens=3,
)
print(chat_completion.usage.completion_tokens) # returns 3, ok.
print(chat_completion.choices[0].finish_reason) # returns "length", ok.

chat_completion = client.chat.completions.create(
model="models/mys/ggml_llava-v1.5-13b/ggml-model-q4_k.gguf",
messages=[
{"role": "user", "content": "Write a limerick about python exceptions"}
],
max_completion_tokens=3,
)
print(chat_completion.usage.completion_tokens) # returns much more than 3 (complete answer).
print(chat_completion.choices[0].finish_reason) # returns "stop".
```

According to OpenAI API, `max_completion_tokens` argument is replacing the deprecated `max_tokens` argument.
It's seems that only `max_tokens` is not ignored by the server.

# Environment and Context

llama_cpp installed with `pip install llama-cpp-python[server]`
`print(llama_cpp.__version__)`: 0.3.6
`print(openai.__version__)`: 1.59.7

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.