deepseek-ai / deepseek-ai/DeepSeek-Coder
疑惑:为什么 base 模型的 tokenizer 词表中也有类似 <|Assistant|> 这样多用于 chat 模型的 special tokens?
Open
- Dominant language
- Python
- Stars
- 24.3k
- Forks
- 2.9k
- PR merge metrics
- No merged PRs in 30d
Description
按照论文所说,预训练阶段没有加入 SFT 数据,那这部分 token 是否有未被充分训练的风险呢?
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue names no files, tests, or entry points. Start by examining the base model's tokenizer vocabulary and the documented pretraining setup, then compare how the special tokens are represented and trained; done means a documented, evidence-based answer about whether they are undertrained.
Written by the indexing model from the issue text.
Assessment
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100