eclipse-ee4j / eclipse-ee4j/parsson

Improve performance of org.glassfish.json.JsonParserImpl

Open
#15 1 comment 1 reaction 0 assignees View on GitHub
Dominant language
Java
Stars
17
Forks
25
PR merge metrics
No merged PRs in 30d

Description

As more and more products use the JSON-P reference implementation in production, it is critical that parsing performance is good.

I think there is an opportunity to improve the performance of JsonParserImpl. Right now, the underlying tokenizer operates on a Java character string and is completely unaware of the underlying byte representation. In many cases, JSON is persisted as UTF8 - from rfc8259:

> JSON text exchanged between systems that are not part of a closed
> ecosystem MUST be encoded using UTF-8 [RFC3629].

Java characters are represented in UTF-16 and conversion from UTF-8 to UTF-16 is often expensive.

I suggest making a special purpose tokenizer that operates directly on the UTF8 byte stream. Other encodings can continue to use the current code path as they will be less common. A special-case UTF8 tokenizer would provide the following benefits:

(1) Markup characters in the ascii range (curly braces, brackets, string delimiters, white space, etc) can be scanned with byte comparisons and never converted to UTF-16.

(2) JSON numbers, true, false, and null don't need to be converted to UTF-16

(3) Strings (keys and values) can be converted to UTF-16 lazily so that if they are never consumed by an application, they need not be converted.

(4) Skip methods (like skipArray() and skipObject()) could avoid any character set conversion of the skipped item.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.