Clouditera / Clouditera/SecGPT
咨询一下预训练数据集中论文的部分
Open
- Dominant language
- Python
- Stars
- 3.1k
- Forks
- 370
- PR merge metrics
- No merged PRs in 30d
Description
看到数据集类型统计中,论文占了 51%。
比较好奇,有两个问题不知道作者能否解答:
1. 和普通网页、书籍类的知识相比,论文这种形式的数据,不同比例影响大么?
2. 是否有论文的收集清洗方案可以开源出来,复用于其他类似领域?
多谢。
Contributor guide
No contributing guide indexed for this repository
Research direction
Review the dataset type statistics referenced in the issue and any existing project documentation about pretraining data. Clarify whether the project documents the effect of paper-data proportions and whether a reusable collection and cleaning scheme is available. Done means providing answers or links for both questions.
Written by the indexing model from the issue text.
Assessment
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100