ACM-VIT/scrag
View on GitHubA flexible web scraper that intelligently adapts to different website structures using multiple extraction strategies (newspaper3k, readability-lxml, BeautifulSoup, and optional headless rendering). It outputs clean, structured data for RAG pipelines or local LLMs, with an optional extension to automatically build RAG indexes from web queries.
- Stars
- 4
- Forks
- 15
- Open beginner issues
- 0
- Indexed issues
- 4
- Dominant language
- Python
- License
- MIT
- Last GitHub push
- Nov 12, 2025
- Latest indexed
- Sep 14, 2026
- Contributing guide
- Contributing guide
- Code of conduct
- Code of conduct
- Beginner labels
- hacktoberfest good first issue help wanted
- PR merge metrics
- No merged PRs in 30d
-
enhancement good first issue hacktoberfest help wanted
-
enhancement hacktoberfest
-
enhancement good first issue hacktoberfest help wanted