elastic / elastic/issue-crawler

Feature Request: Allow PDF & office-suite documents extraction in Web Crawler

Open
#3 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
JavaScript
Stars
1
Forks
5
Avg merge
1d 12h
Merged PRs (30d)
2

Description

Hi there,

We're looking to migrate from Swiftype to Elastic Cloud however on of the key features of the Swiftype crawler, the abilityt o index PDF and office documents, appears to be missing from the Elastic Cloud crawler solution.

We see here this is a known issue:

https://www.elastic.co/guide/en/app-search/current/crawl-web-content.html

```The web crawler does not extract and index non-HTML content (e.g. JavaScript, PDF).```

We also enabled the **Ingest Attachment Processor Plugin** however it looks as though the crawler won't automatically support this.

What would be involved in getting this feature implemented?

We're mindful that all Swiftype instances will have to be migrated to Elastic Cloud eventually however at this stage it's not a viable option if PDF/office document indexing is not supported.

Thanks,
Sam

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.