microsoft / microsoft/BitNet

Llama server prediction

Open
#299 1 comment 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
40.3k
Forks
3.7k
PR merge metrics
No merged PRs in 30d

Description

Hi,
When I am using cnv mode and it is correctly giving the exact answer.

But when I am using llama server mode and triggering endpoint with question,
It returns unwanted things as well with the correct answer. (N_predict is set to 20).it actually returns words until 20 tokens.

1.Why it is not behaving like cnv mode.
2.it means can't we use server mode.
3.any solution for this?

Thankyou

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the reported difference between cnv mode and llama server mode by calling the server endpoint with n_predict set to 20. Compare the response with the expected answer and inspect how each mode handles the prediction limit. Done means documenting the cause of the extra output and confirming whether a fix or configuration change is needed.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
ai
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.