internetarchive / internetarchive/heritrix3

Avoid speculative links extraction for meta fields known not to contain links

Open
#225 5 comments 0 reactions 0 assignees View on GitHub
bug pull request welcome
Dominant language
Java
Stars
3.3k
Forks
793
Avg merge
1d 8h
Merged PRs (30d)
8

Description

Following [this report](https://groups.yahoo.com/neo/groups/archive-crawler/conversations/topics/8952;_ylc=X3oDMTM0b2wwNG0wBF9TAzk3MzU5NzE0BGdycElkAzg3NTk4NjcEZ3Jwc3BJZAMxNzA1MDA0OTI0BG1zZ0lkAzg5NTQEc2VjA2Z0cgRzbGsDdnRwYwRzdGltZQMxNTQ2NTE0NjUwBHRwY0lkAzg5NTI-) of a URL being constructed from `` elements:

> I'm using heritrix 3.3.0-SNAPSHOT and see some strange behavior in the link extraction. This is one example in crawl.log:
>
> 2018-12-21T04:07:03.874Z 404 7161 https://stitch-maps.com/news/2018/10/twofer/Stitch-Maps.com RLX https://stitch-maps.com/news/2018/10/twofer/ text/html #116 20181219040702090+1782 sha1:K7HLTQ7SFI4KAQN3NVAO4OJ4UBYT3FGE - -
>
> There isn't any link to the crawled url on the given src page, so it seems like the facebook tags on the src page have something to do with it:
>
>
>
>
> Isn't it a bug, that heritrix combined these two urls to https://stitch-maps.com/news/2018/10/twofer/Stitch-Maps.com?
```

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.