Implement faster tokenizer versions for `DOMParser`/`innerHTML`
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 2.6k
- Forks
- 288
- Avg merge
- 2d 22h
- Merged PRs (30d)
- 8
Description
The main tokenizer needs to count line numbers, which introduces significant overhead, especially in the SIMD case.
These line numbers are not required in some situations and we should add an option (perhaps with a bool const generic?) to disable them.
Gecko has something similar in https://searchfox.org/firefox-main/source/parser/html/nsHtml5TokenizerLoopPoliciesSIMD.h.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by tracing the main tokenizer and its SIMD path for DOMParser and innerHTML, then compare the Gecko tokenizer loop policy linked in the issue. Determine how an option can disable line-number counting for these cases while preserving it where required; done means the faster paths work without line-number overhead and existing parser behavior remains intact.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust
- Domain
- web-dev
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100