deepseek-ai / deepseek-ai/DeepSeek-Coder

Construction of the FIM training data

Open
#107 4 comments 2 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
24.3k
Forks
2.9k
PR merge metrics
No merged PRs in 30d

Description

Hi, dear:

Thank you very much for your open source. Will the code of FIM dataset construction and training be made public? such as the number of lines or length of the code for Prefix, suffix, and middle.
We would like to build on your model and fine-tune it on our own code data warehouse, especially to improve the FIM performance of our internal code.

Thx.

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue asks whether FIM dataset construction and training code will be released, including Prefix, suffix, and middle lengths, but names no files, tests, or entry points. Start by reviewing the repository documentation and training materials; the issue is complete only when the requested implementation or a clear release decision is available.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.