lm-sys / lm-sys/FastChat

Hardcoded host localhost and port 9090 for a rate monitor

Open
#3,531 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
39.5k
Forks
4.8k
PR merge metrics
No merged PRs in 30d

Description

In gradio_web_server.py, there is a hardcoded host and port for a supposed monitor daemon, but no daemon is around:
https://github.com/lm-sys/FastChat/blob/a04072e35d0893e64169c8e3ea153312eb0fe9ac/fastchat/serve/gradio_web_server.py#L391

This, in turn, breaks the usage of openAI endpoints with the --register to a json file.

So, for example, this openai_compatible_server.json can't run:

```
{
"Llama 405": {
"model_name": "llama3.1:405b",
"api_type": "openai",
"api_base": "http://localhost:11434/v1",
"api_key": "",
"anony_only": false,
"recommended_config": {
"temperature": 0.7,
"top_p": 1.0
}
}
}
```

Because it will always fail with a `CONNECTION REFUSED` since we have no such monitor daemon:

```
2024-09-19 21:31:47 | INFO | gradio_web_server | bot_response. ip: 127.0.0.1
2024-09-19 21:31:47 | INFO | gradio_web_server | monitor error: HTTPConnectionPool(host='localhost', port=9090): Max retries exceeded with url: /is_limit_reached?model=Llama%20405%20on%20WestAI&user_id=127.0.0.1 (Caused by NewConnectionError(': Failed to establish a new connection: [Errno 61] Connection refused'))
```

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in fastchat/serve/gradio_web_server.py at the linked line around 391 and inspect how the rate-monitor request is made. Reproduce the failure with the provided openai_compatible_server.json configuration, then verify that the configuration no longer fails because of the unavailable localhost:9090 monitor daemon.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.