ByteDance-Seed / ByteDance-Seed/Seed-Coder

Some Questions about Quality Model~

Open
#18 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
No language data
Stars
759
Forks
59
PR merge metrics
No merged PRs in 30d

Description

Hi there! Thanks for your great work. I have a few questions regarding your custom-trained code quality scorer model.
The paper mentions that you adopted a Llama-series pretrained model as the backbone. However, in Appendix A2.1 about the evaluation prompt, it states:

"It remains consistent throughout the entire pipeline, from collecting ground-truth data to training the quality scorer and applying it across all GitHub data during inference."

I would like to confirm:
Does this mean you appended the fixed prompt to code samples during the training phase of the quality model?
Additionally, since the base Llama model is only pretrained and has not undergone instruction tuning, is it necessary to feed a task-specific prompt together with code inputs for quality model training?
Thanks a lot for your clarification!

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.