elastic / elastic/issue-crawler
Feature Request: Allow PDF & office-suite documents extraction in Web Crawler
- Dominant language
- JavaScript
- Stars
- 1
- Forks
- 5
- Avg merge
- 1d 12h
- Merged PRs (30d)
- 2
Description
Hi there,
We're looking to migrate from Swiftype to Elastic Cloud however on of the key features of the Swiftype crawler, the abilityt o index PDF and office documents, appears to be missing from the Elastic Cloud crawler solution.
We see here this is a known issue:
https://www.elastic.co/guide/en/app-search/current/crawl-web-content.html
```The web crawler does not extract and index non-HTML content (e.g. JavaScript, PDF).```
We also enabled the **Ingest Attachment Processor Plugin** however it looks as though the crawler won't automatically support this.
What would be involved in getting this feature implemented?
We're mindful that all Swiftype instances will have to be migrated to Elastic Cloud eventually however at this stage it's not a viable option if PDF/office document indexing is not supported.
Thanks,
Sam
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.