akosbalasko / akosbalasko/yarle

Re-OCR PDFs

未關閉
#82 5 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
enhancement
主要語言
TypeScript
星號
1.8k
分支
111
PR 合併指標
30 天內沒有已合併 PR

描述

This might be out of scope, but it might be interesting: OCR PDFs (maybe through [OCRmyPDF](https://ocrmypdf.readthedocs.io/en/latest/)? Is there something more native to JS/TS?) that are included as attachments. Evernote describes their image recognition [here](https://evernote.com/blog/how-evernotes-image-recognition-works/):

>If you export a note containing a PDF that has been processed by the OCR system, there will be two nodes in the document: data and alternate-data . The data node contains a base–64 encoded version of the original PDF and the alternative-data represents the searchable version of the same PDF.

In [the DevonThink Forums](https://discourse.devontechnologies.com/t/how-to-batch-ocr-15-000-evernote-notes-many-of-which-contain-pdfs/59050/16?u=matti) one person from that company said that the OCR-Data is not available, but the above statement makes it sound it would be included.

貢獻指南

這個儲存庫沒有索引到貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。