python / python/cpython

Add flush() to HTMLParser to fill gap after implementation change

Open
#157,003 1 comment 0 reactions 1 assignee View on GitHub

@serhiy-storchaka is already working on this.

Since Sep 5, 2026.

3.16 stdlib type-feature
Dominant language
Python
Stars
77.2k
Forks
35.9k
PR merge metrics
PR metrics pending

Description

Feature or enhancement

Proposal:

With recent changes it isn't clear when the the parser actually does the parsing/processing anymore on the data given to feed().

Changes were made in #153030 , (seen in 3.14.7), which leave this gap in the HTMLParser.

Propose adding a flush method as seen below ( I reset the parse threshold to 1 which differs from my proposal on discuss, seems appropriate?):

class PrototypeHTMLParser(HTMLParser):

    def flush(self):
        """
        Process all data fed insofar as it contains complete elements any remaining data remains buffered.
        """
        if self._pending:
            self.rawdata += ''.join(self._pending)
            self._pending.clear()
            self._pending_len = 0
            self._parse_threshold = 1
        # Perform incremental parsing but its not EOF.
        self.goahead(0)
Example
from html.parser import HTMLParser


class FlushableHTMLParser(HTMLParser):
    """
    Parser with flush implemented.

    This would be meant to be added to the existing HTMLParser.
    """
    def flush(self):
        if self._pending:
            self.rawdata += ''.join(self._pending)
            self._pending.clear()
            self._pending_len = 0
            self._parse_threshold = 1
        self.goahead(0)


class DemoHTMLParser(FlushableHTMLParser):

    def handle_starttag(self, tag, attrs):
        print (f'Found starttag {tag=} {len(attrs)=}')

    def handle_endtag(self, tag):
        print (f'Found endtag {tag=}')


def parse(html_str_parts: list[str]):
    parser = DemoHTMLParser()
    for part in html_str_parts:
        print (f'Feed {len(part)}')
        parser.feed(part)
    print ('')
    print ('Calling flush()')
    parser.flush()
    print ('Calling close()')
    parser.close()


def demo_attrs(attrs_len):
    print (f'\nDemo with {attrs_len=}')
    print ('='*20)
    attrs = [f" a{i}='1'" for i in range(attrs_len)]
    parse(["<div", *attrs, ">", "content", "</div>"])


if __name__ == '__main__':
    for attrs_len in (2, 8):
        demo_attrs(attrs_len)

Demo with attrs_len=2
====================
Feed 4
Feed 7
Feed 7
Feed 1
Feed 7
Found starttag tag='div' len(attrs)=2
Feed 6
Found endtag tag='div'

Calling flush()
Calling close()

Demo with attrs_len=8
====================
Feed 4
Feed 7
Feed 7
Feed 7
Feed 7
Feed 7
Feed 7
Feed 7
Feed 7
Feed 1
Feed 7
Feed 6

Calling flush()
Found starttag tag='div' len(attrs)=8
Found endtag tag='div'
Calling close()
Other Doc Changes

I think the feed() doc string also needs a cleanup to match these changes. Should that also be part of this issue?

Handling EOF

There is a related issue about how can a stdlib user determine what should be done when close() is called and data is discarded. Specifically from the living spec, eof-in-tag error case. I'm going to link it but include the text here as well:

This error occurs if the parser encounters the end of the input stream in a start tag or an end tag (e.g., <div id=). Such a tag is ignored.

Should I start that in discuss in the same thread as before or make another issue about that? Maybe it is beyond the scope of the python stdlib parser? It seems like a hard problem but it kind of makes the parser into partial solution in certain use cases.

Has this already been discussed elsewhere?

I have already discussed this feature proposal on Discourse

Links to previous discussion of this feature:

https://discuss.python.org/t/proposal-add-new-method-to-htmlparser-flush-self-str/108855

Linked PRs
  • gh-157012

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.