OpenAI o1 models require `max_completion_tokens` instead of `max_tokens`
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 12.5k
- Forks
- 998
- Avg merge
- 3d 13h
- Merged PRs (30d)
- 10
Description
Problem
OpenAI o1 models require max_completion_tokens instead of max_tokens.
When the llm command is used with an OpenAI o1 series model like o1-preview, the -o max_tokens N option returns an error from OpenAI that max_completion_tokens should be used instead.
When -o max_completion_tokens N is used, llm generates an error instead of passing it to the OpenAI API.
Model Provider Documentation
The OpenAI docs explain that max_tokens is deprecated and is already not compatible with the o1 models:
https://platform.openai.com/docs/api-reference/chat/create
max_tokens Deprecated integer or null
Optional
The maximum number of tokens that can be generated in the chat completion. This value can be used to control costs for text generated via API.This value is now deprecated in favor of max_completion_tokens, and is not compatible with o1 series models.
max_completion_tokens integer or null
Optional
An upper bound for the number of tokens that can be generated for a completion, including visible output tokens and reasoning tokens.
Examples
These commands demonstrate the problem:
$ LLM_OPENAI_SHOW_RESPONSES=1 llm --no-stream -m o1-preview -o max_completion_tokens 32000 "The response to this is a creative novel"
Error: max_completion_tokens
Extra inputs are not permitted
$ LLM_OPENAI_SHOW_RESPONSES=1 llm --no-stream -m o1-preview -o max_tokens 32000 "The response to this is a creative novel"
Request: POST https://api.openai.com/v1/chat/completions
Headers:
host: api.openai.com
connection: keep-alive
accept: application/json
content-type: application/json
user-agent: OpenAI/Python 1.60.1
x-stainless-lang: python
x-stainless-package-version: 1.60.1
x-stainless-os: Linux
x-stainless-arch: x64
x-stainless-runtime: CPython
x-stainless-runtime-version: 3.10.12
authorization: [...]
x-stainless-async: false
x-stainless-retry-count: 0
content-length: 138
Body:
{
"messages": [
{
"role": "user",
"content": "The response to this is a creative novel"
}
],
"model": "o1-preview",
"max_tokens": 32000,
"stream": false
}
Response: status_code=400
Headers:
date: Tue, 28 Jan 2025 05:46:32 GMT
content-type: application/json
content-length: 245
connection: keep-alive
access-control-expose-headers: X-Request-ID
openai-organization: [...]
openai-processing-ms: 20
openai-version: 2020-10-01
x-ratelimit-limit-requests: 10000
x-ratelimit-limit-tokens: 30000000
x-ratelimit-remaining-requests: 9999
x-ratelimit-remaining-tokens: 29995904
x-ratelimit-reset-requests: 6ms
x-ratelimit-reset-tokens: 8ms
x-request-id: req_[...]
strict-transport-security: max-age=31536000; includeSubDomains; preload
cf-cache-status: DYNAMIC
set-cookie: __cf_bm=...
x-content-type-options: nosniff
server: cloudflare
cf-ray: [...]
alt-svc: h3=":443"; ma=86400
Body:
{
"error": {
"message": "Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead.",
"type": "invalid_request_error"
,
"param": "max_tokens",
"code": "unsupported_parameter"
}
}
Error: Error code: 400 - {'error': {'message': "Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead.", 'type': 'invalid_request_error', 'param': 'max_tokens', 'code': 'unsupported_parameter'}}
$ llm --version
llm, version 0.20
$ lsb_release -a
No LSB modules are available.
Distributor ID: Ubuntu
Description: Ubuntu 22.04.5 LTS
Release: 22.04
Codename: jammy
$ uname -a
Linux [...] 6.8.0-52-generic #53~22.04.1-Ubuntu SMP PREEMPT_DYNAMIC Wed Jan 15 19:18:46 UTC 2 x86_64 x86_64 x86_64 GNU/Linux
llm Documentation
This documentation can also be updated along with fixing the code, as it lists max_tokens instead of max_completion_tokens for the o1 models:
https://github.com/simonw/llm/blob/main/docs/contributing.md
OpenAI Chat: o1
Options:
temperature: float
max_tokens: int
[...]
OpenAI Chat: o1-2024-12-17
Options:
temperature: float
max_tokens: int
[...]
OpenAI Chat: o1-preview
Options:
temperature: float
max_tokens: int
[...]
OpenAI Chat: o1-mini
Options:
temperature: float
max_tokens: int
[...]
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the OpenAI option handling and the examples in docs/contributing.md, which currently list max_tokens for the o1 models. Verify that both max_tokens and max_completion_tokens behave correctly for o1 requests, update the documentation, and confirm the demonstrated commands no longer produce the reported errors.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, cli
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100