python / python/cpython

Add flush() to HTMLParser to fill gap after implementation change

未關閉
#157,003 1 則留言 0 個 reaction 已指派 1 人 在 GitHub 檢視

@serhiy-storchaka 已經在處理了。

開始於 2026年9月5日。

3.16 stdlib type-feature
主要語言
Python
星號
77.2k
分支
36k
PR 合併指標
PR 指標待擷取

描述

Feature or enhancement

Proposal:

With recent changes it isn't clear when the the parser actually does the parsing/processing anymore on the data given to feed().

Changes were made in #153030 , (seen in 3.14.7), which leave this gap in the HTMLParser.

Propose adding a flush method as seen below ( I reset the parse threshold to 1 which differs from my proposal on discuss, seems appropriate?):

class PrototypeHTMLParser(HTMLParser):

    def flush(self):
        """
        Process all data fed insofar as it contains complete elements any remaining data remains buffered.
        """
        if self._pending:
            self.rawdata += ''.join(self._pending)
            self._pending.clear()
            self._pending_len = 0
            self._parse_threshold = 1
        # Perform incremental parsing but its not EOF.
        self.goahead(0)
Example
from html.parser import HTMLParser


class FlushableHTMLParser(HTMLParser):
    """
    Parser with flush implemented.

    This would be meant to be added to the existing HTMLParser.
    """
    def flush(self):
        if self._pending:
            self.rawdata += ''.join(self._pending)
            self._pending.clear()
            self._pending_len = 0
            self._parse_threshold = 1
        self.goahead(0)


class DemoHTMLParser(FlushableHTMLParser):

    def handle_starttag(self, tag, attrs):
        print (f'Found starttag {tag=} {len(attrs)=}')

    def handle_endtag(self, tag):
        print (f'Found endtag {tag=}')


def parse(html_str_parts: list[str]):
    parser = DemoHTMLParser()
    for part in html_str_parts:
        print (f'Feed {len(part)}')
        parser.feed(part)
    print ('')
    print ('Calling flush()')
    parser.flush()
    print ('Calling close()')
    parser.close()


def demo_attrs(attrs_len):
    print (f'\nDemo with {attrs_len=}')
    print ('='*20)
    attrs = [f" a{i}='1'" for i in range(attrs_len)]
    parse(["<div", *attrs, ">", "content", "</div>"])


if __name__ == '__main__':
    for attrs_len in (2, 8):
        demo_attrs(attrs_len)

Demo with attrs_len=2
====================
Feed 4
Feed 7
Feed 7
Feed 1
Feed 7
Found starttag tag='div' len(attrs)=2
Feed 6
Found endtag tag='div'

Calling flush()
Calling close()

Demo with attrs_len=8
====================
Feed 4
Feed 7
Feed 7
Feed 7
Feed 7
Feed 7
Feed 7
Feed 7
Feed 7
Feed 1
Feed 7
Feed 6

Calling flush()
Found starttag tag='div' len(attrs)=8
Found endtag tag='div'
Calling close()
Other Doc Changes

I think the feed() doc string also needs a cleanup to match these changes. Should that also be part of this issue?

Handling EOF

There is a related issue about how can a stdlib user determine what should be done when close() is called and data is discarded. Specifically from the living spec, eof-in-tag error case. I'm going to link it but include the text here as well:

This error occurs if the parser encounters the end of the input stream in a start tag or an end tag (e.g., <div id=). Such a tag is ignored.

Should I start that in discuss in the same thread as before or make another issue about that? Maybe it is beyond the scope of the python stdlib parser? It seems like a hard problem but it kind of makes the parser into partial solution in certain use cases.

Has this already been discussed elsewhere?

I have already discussed this feature proposal on Discourse

Links to previous discussion of this feature:

https://discuss.python.org/t/proposal-add-new-method-to-htmlparser-flush-self-str/108855

Linked PRs
  • gh-157012

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。