aboutcode-org / aboutcode-org/vulnerablecode

Improve PoC collection using GitHub archive data

未關閉
#2,429 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
主要語言
Python
星號
702
分支
328
平均合併
3 天 8 小時
30 天內合併 PR
3

描述

We are currently collecting PoCs primarily from this repository:

* https://github.com/nomi-sec/PoC-in-GitHub

From my understanding, this repository is generated by an automated bot.

There is another project doing something very similar:

* https://github.com/ycdxsb/PocOrExp_in_Github

The general approach seems to be running a CI job that uses the GitHub API to search for repositories containing CVE IDs and then collecting the results.

However, I think we could use a cleaner and potentially more reliable approach for collecting PoCs.

Instead of repeatedly querying the GitHub API for every CVE ID, we could use the GitHub hourly archive data:

* https://github.com/giant-hourly-archive/giant-hourly-archive-2011
* ...
* https://github.com/giant-hourly-archive/giant-hourly-archive-2026

The idea would be to process the archive data locally and search for CVE IDs across newly indexed GitHub content. This could significantly reduce the number of GitHub API requests and give us a more reproducible dataset.

We could then add an extra validation layer to determine whether a discovered repository is actually a valid PoC/Exploit repository. For example, this could involve:

* Manual review by contributors for higher-confidence results.
* An LLM-based validation step that reads the repository metadata/content and determines whether it actually contains a PoC or exploit related to the identified CVE.
* Potentially combining both approaches to assign a confidence level to each result.

As a proof of concept, I’ve tested this approach with:
* https://github.com/ziadhany/ExploitArchive

貢獻指南

這個儲存庫沒有索引到貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。