Difference between the prompts of v0 and v1
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 39.5k
- Forks
- 4.8k
- PR merge metrics
- No merged PRs in 30d
Description
Thanks for your great job! I have checked the code of v0 and v1, and found some differences between the prompts:
```python
# Vicuna-v0
sep = "###"
train_prompt = "system.### Human: xxx.\n### Assistant: yyy.\n### Human: xxx.\n### Assistant: yyy.\n"
target = " ### ### Assistant: yyy.\n### ### Assistant: yyy.\n"
eval_prompt = "system.###Human: xxx.###Assistant:"
# Vicuna-v1
sep = " "
sep2 = ""
train_prompt = "system. Human: xxx. Assistant: yyy.Human: xxx. Assistant: yyy."
target = " yyy. yyy."
eval_prompt = "system. Human: xxx. Assistant:"
```
Is there a bug in the eval_prompt of v0? It seems the " " after "###" and "\n" are ignored.
Besides, is it essential to predict the "###" before the " Human" when training?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the v0 and v1 prompt code shown in the issue, focusing on sep, sep2, train_prompt, target, and eval_prompt. Confirm the prompt and token differences, then determine whether the v0 spacing is intentional and whether predicting ### before Human is required; no file or test is named in the report.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100