elastic / elastic/crawler

Validate domains before running crawl

Open
#345 0 comments 0 reactions 0 assignees View on GitHub
enhancement team:extract-and-transform
Dominant language
Ruby
Stars
224
Forks
48
Avg merge
23h 16m
Merged PRs (30d)
18

Description

Currently we don't validate domains before crawling. This can only be done through the `validate` command.
It is this way because we previously had a UI that would validate domain inputs, so invalid domains would be impossible (theoretically) to configure. However, that is no longer the case, and users can try to crawl invalid domains, which causes all sorts of weird errors.

Quick example

- `https://elastic.co` is invalid (it redirects)
- `https://www.elastic.co` is the correct domain

A user doesn't necessarily understand that the first URL is invalid based on our current error logging.
If we validate the domain at crawl time, we can avoid this problem.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.