github / github/codeql

How to make a database of a very large codebase very quickly?

未关闭
#8,342 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
question
主要语言
CodeQL
星标
10.1k
派生
2.1k
平均合并
2 天 15 小时
30 天内合并 PR
141

描述

**Description of the issue**
I have close 10K small projects. All these projects have individual CSS/HTML/JS files. Often many times I want to scan the entire catalog for some query. The catalog also keeps changing with new addition/modification to packages

My requirements are to scan the entire catalog on demand in the most efficient manner possible.

I have explored 2 approaches

1. Upload all packages as individual databases to LGTM and scan from there. The challenge here is that it is very slow to scan entire catalog. I have 9 query workers and it takes 7 hours to scan.

2. Create a big LGTM database and scan over this only on demand. The challenge here is that combined all these packages are around 27GB. How do I create a LGTM database in the minimum time possible?

I am using a Standard_F32s_v2 VM on azure and it is still taking lot of time (at least 5 hours so far) to run. This is a compute optimized VM.
What should be my VM configuration ? Is more power better? Should I go for a memory optimized or compute optimized ?
Are there any more command parameters than I can specify to speed this up. I currently use `threads=0` while creating a database though I don't know its significance.

Any other inputs would be highly appreciated.

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。