deepseek-ai / deepseek-ai/DeepSeek-Coder
Construction of the FIM training data
- Dominant language
- Python
- Stars
- 24.3k
- Forks
- 2.9k
- PR merge metrics
- No merged PRs in 30d
Description
Hi, dear:
Thank you very much for your open source. Will the code of FIM dataset construction and training be made public? such as the number of lines or length of the code for Prefix, suffix, and middle.
We would like to build on your model and fine-tune it on our own code data warehouse, especially to improve the FIM performance of our internal code.
Thx.
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue asks whether FIM dataset construction and training code will be released, including Prefix, suffix, and middle lengths, but names no files, tests, or entry points. Start by reviewing the repository documentation and training materials; the issue is complete only when the requested implementation or a clear release decision is available.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100