aboutcode-org / aboutcode-org/vulnerablecode

Improve PoC collection using GitHub archive data

未关闭
#2,429 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
702
派生
328
平均合并
3 天 8 小时
30 天内合并 PR
3

描述

We are currently collecting PoCs primarily from this repository:

* https://github.com/nomi-sec/PoC-in-GitHub

From my understanding, this repository is generated by an automated bot.

There is another project doing something very similar:

* https://github.com/ycdxsb/PocOrExp_in_Github

The general approach seems to be running a CI job that uses the GitHub API to search for repositories containing CVE IDs and then collecting the results.

However, I think we could use a cleaner and potentially more reliable approach for collecting PoCs.

Instead of repeatedly querying the GitHub API for every CVE ID, we could use the GitHub hourly archive data:

* https://github.com/giant-hourly-archive/giant-hourly-archive-2011
* ...
* https://github.com/giant-hourly-archive/giant-hourly-archive-2026

The idea would be to process the archive data locally and search for CVE IDs across newly indexed GitHub content. This could significantly reduce the number of GitHub API requests and give us a more reproducible dataset.

We could then add an extra validation layer to determine whether a discovered repository is actually a valid PoC/Exploit repository. For example, this could involve:

* Manual review by contributors for higher-confidence results.
* An LLM-based validation step that reads the repository metadata/content and determines whether it actually contains a PoC or exploit related to the identified CVE.
* Potentially combining both approaches to assign a confidence level to each result.

As a proof of concept, I’ve tested this approach with:
* https://github.com/ziadhany/ExploitArchive

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。