ACM-VIT

ACM-VIT/scrag

View on GitHub

A flexible web scraper that intelligently adapts to different website structures using multiple extraction strategies (newspaper3k, readability-lxml, BeautifulSoup, and optional headless rendering). It outputs clean, structured data for RAG pipelines or local LLMs, with an optional extension to automatically build RAG indexes from web queries.

Stars
4
Forks
15
Open beginner issues
0
Indexed issues
4
Dominant language
Python
License
MIT
Last GitHub push
Nov 12, 2025
Latest indexed
Sep 14, 2026
Contributing guide
Contributing guide
Code of conduct
Code of conduct
Beginner labels
hacktoberfest good first issue help wanted
PR merge metrics
No merged PRs in 30d
4 open issues indexed Loading issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.