ArthurHub / ArthurHub/HTML-Renderer
Stack Overflow reading huge old HTML
- Dominant language
- C#
- Stars
- 1.4k
- Forks
- 550
- Avg merge
- 14d 9h
- Merged PRs (30d)
- 1
Description
I had a performance problem displaying huge HTML , then I was looking for huge simple HTML in english.
I found [http://www.gutenberg.org/files/1661/1661-h/1661-h.htm](http://www.gutenberg.org/files/1661/1661-h/1661-h.htm) and I tried and got StackOverflow error.
I knew this HTML was so old therefore it might be cause error on parsing.
I suggest using [SGMLReader](https://github.com/MindTouch/SGMLReader) for parsing any HTML.
I'm using it for many scraping projects and results are so reasonable.
Contributor guide
No contributing guide indexed for this repository
Research direction
Reproduce the StackOverflow error while rendering the Gutenberg HTML at http://www.gutenberg.org/files/1661/1661-h/1661-h.htm, then trace the HTML parsing and rendering entry points. Compare the current behavior with the proposed SGMLReader approach; done means the large document renders without a stack overflow and its existing output remains reasonable.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- csharp
- Domain
- frontend, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100