deepseek-ai / deepseek-ai/DeepSeek-Coder

疑惑:为什么 base 模型的 tokenizer 词表中也有类似 <|Assistant|> 这样多用于 chat 模型的 special tokens?

Open
#165 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
24.3k
Forks
2.9k
PR merge metrics
No merged PRs in 30d

Description

按照论文所说,预训练阶段没有加入 SFT 数据,那这部分 token 是否有未被充分训练的风险呢?

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no files, tests, or entry points. Start by examining the base model's tokenizer vocabulary and the documented pretraining setup, then compare how the special tokens are represented and trained; done means a documented, evidence-based answer about whether they are undertrained.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.