alphapapa / alphapapa/org-protocol-capture-html

FR: `capture-html` , but with `eww-readable`

Đang mở
#53 1 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
Emacs Lisp
Star
461
Fork
41
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

As always, thanks for the useful package!

### Feature Request / Proposal

I propose to add a capture protocol that behaves like `capture-html` currently does, but (if nothing is selected), sends the whole page's html (i.e., `document.body`) to emacs and uses `eww-readable` to extract the interesting html (before converting it with pandoc). To basically a mix of the two existing protocols.

### Use case

- Let's say I want to capture a webpage into org. If I only want to capture selected part(s) of the website, then `capture-html` works perfectly.
- On the other hand, if I want to capture the full webpage, and want to use the power of `eww-readable`, but want to capture the content I have already open, this is currently not possible. (Or I did misunderstand something.)

### Background

What I mean: If I understand things correctly,

`org-protocol-capture-html--with-pandoc` (used by the `capture-html` protocol)
1) user selects some text in browser, clicks on bookmarklet
2) selected text is send to emacs via org-protocol
3) function converts the html (coming from the browser) to org

`org-protocol-capture-html--capture-eww-readable` (used by the `capture-eww-readable` protocol`)
1) user clicks on bookmarklet (which does not send selected text to emacs)
2) function uses `org-protocol-capture-html--url-html` to _download html_ directly
3) function then converts the html (downloaded by curl) to org

The difference being that in `capture-html` the html is retrieved by the browser, while in `capture-eww-readable` it is retrieved by emacs/curl.

This makes a difference when the content being captured is, e.g., a page behind a paywall, or dynamically generated content. It is also a difference when the text is inserted into the DOM by some javascript which loads the content after the page itself is loaded.

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.