internetarchive / internetarchive/dweb-gateway
Handle robots.txt
Open
- Dominant language
- Python
- Stars
- 20
- Forks
- 5
- PR merge metrics
- No merged PRs in 30d
Description
Seeing queries to /arc/archive.org/robots.txt which probably means they are going to dweb.archive.org/robots.txt
* Decide on policy - probably block it.
* Implement on dweb.archive.org and probably on dweb.me/arc/archive.org/robots.txt and dweb.me/robots.txt
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by tracing how the gateway handles requests for /arc/archive.org/robots.txt, dweb.archive.org/robots.txt, dweb.me/arc/archive.org/robots.txt, and dweb.me/robots.txt. Resolve the intended blocking policy first, then verify that the selected paths consistently apply it.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend, web-dev
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100