microsoft / microsoft/BitNet

Feature request: document/support llama-server HTTP endpoint for OpenAI-compatible serving

Open
#432 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
40.3k
Forks
3.7k
PR merge metrics
No merged PRs in 30d

Description

Background

setup_env.py already builds llama-server as part of the cmake step (it lives at build/bin/llama-server after a successful build). This binary provides a fully OpenAI-compatible HTTP API (/v1/chat/completions, /v1/completions, /v1/models) — the same interface as llama.cpp's server.

The README currently only documents run_inference.py for inference. The server binary is silently present but undiscovered by most users.

What this unlocks

  • Drop-in replacement for OpenAI API in downstream tools (LangChain, Open WebUI, custom apps) without code changes
  • Persistent model loading (no 2-3s cold-start per request)
  • Integration with job queues or proxy layers that speak OpenAI protocol

Minimal usage (after build)

./build/bin/llama-server     --model models/BitNet-b1.58-2B-4T-gguf/ggml-model-i2_s.gguf     --host 127.0.0.1     --port 8080     --parallel 1     --ctx-size 4096

# Then:
curl http://127.0.0.1:8080/v1/chat/completions   -d '{"model":"bitnet","messages":[{"role":"user","content":"Hello"}]}'

Question

Would the team be interested in a PR that:

  1. Documents this capability in the README (a single section — no code changes)
  2. Optionally adds a minimal Python wrapper script (consistent with the repo's Python-first style) to make the invocation discoverable

Happy to contribute either or both if there's interest. Flagging as a question first rather than opening a cold PR.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the README section covering run_inference.py and inspect setup_env.py to confirm where build/bin/llama-server is produced. Document the shown server command and OpenAI-compatible endpoints, and clarify whether the Python wrapper is in scope. Done means a newcomer can discover and launch the HTTP server from the README.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, python
Domain
api, documentation
Issue type
Documentation
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
58/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.