servo / servo/html5ever

Implement faster tokenizer versions for `DOMParser`/`innerHTML`

Open
#703 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

experiment performance
Dominant language
Rust
Stars
2.6k
Forks
288
Avg merge
2d 22h
Merged PRs (30d)
8

Description

The main tokenizer needs to count line numbers, which introduces significant overhead, especially in the SIMD case.

These line numbers are not required in some situations and we should add an option (perhaps with a bool const generic?) to disable them.

Gecko has something similar in https://searchfox.org/firefox-main/source/parser/html/nsHtml5TokenizerLoopPoliciesSIMD.h.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by tracing the main tokenizer and its SIMD path for DOMParser and innerHTML, then compare the Gecko tokenizer loop policy linked in the issue. Determine how an option can disable line-number counting for these cases while preserving it where required; done means the faster paths work without line-number overhead and existing parser behavior remains intact.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
web-dev
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.