akosbalasko / akosbalasko/yarle

Re-OCR PDFs

未关闭
#82 5 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
enhancement
主要语言
TypeScript
星标
1.8k
派生
111
PR 合并指标
30 天内没有已合并 PR

描述

This might be out of scope, but it might be interesting: OCR PDFs (maybe through [OCRmyPDF](https://ocrmypdf.readthedocs.io/en/latest/)? Is there something more native to JS/TS?) that are included as attachments. Evernote describes their image recognition [here](https://evernote.com/blog/how-evernotes-image-recognition-works/):

>If you export a note containing a PDF that has been processed by the OCR system, there will be two nodes in the document: data and alternate-data . The data node contains a base–64 encoded version of the original PDF and the alternative-data represents the searchable version of the same PDF.

In [the DevonThink Forums](https://discourse.devontechnologies.com/t/how-to-batch-ocr-15-000-evernote-notes-many-of-which-contain-pdfs/59050/16?u=matti) one person from that company said that the OCR-Data is not available, but the above statement makes it sound it would be included.

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。