apache / apache/parquet-java

byte array has better performance than ByteBuffer

Open
#2,713 1 comment 0 reactions 0 assignees View on GitHub
Component: Parquet Priority: Major Type: enhancement
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
3d 12h
Merged PRs (30d)
33

Description

Currently the The abstract class BytePacker has the following method

@Deprecated
public void unpack8Values(final byte[] input, final int inPos, final int[] output, final int outPos) {
    unpack8Values(ByteBuffer.wrap(input), inPos, output, outPos);

}

I don’t know why to use ByteBuffer wrap byte[], ByteBuffer has poor performance.

 

I suggest using  

public abstract void unpack8Values(final byte[]input, final int inPos, final int[] output, final int outPos);

to replace

@Deprecated
public void unpack8Values(final byte[] input, final int inPos, final int[] output, final int outPos) {
    unpack8Values(ByteBuffer.wrap(input), inPos, output, outPos);

}

 

 

Tested by me the byte array api has better performance than ByteBuffer api, 

My test result is:

[Unpack8ValuesByteArray spent time] 80 ms
[Unpack8ValuesByteBuffer spent time] 133 ms

 

My test code is:

package org.apache.parquet.column.values.bitpacking;

import java.nio.ByteBuffer;

public class ByteBufferTest {
  private static final BytePacker bytePacker = Packer.LITTLE_ENDIAN.newBytePacker(7);

  private static final int COUNT = 100000;

  public static void main(String[] args) {
    byte  [] in  = new byte[1008];
    int [] out = new int[1152];
    int [] out1 = new int[1152];
    int [] out2 = new int[1152];

    int res = 0;

    for(int i = 0; i < in.length; i++) {
      in[i] = (byte) i;
    }

    for(int i = 0; i < COUNT; i++) {
      res += unpack8ValuesBytes(in, out, i % out.length);
    }

    res = 0;
    long t1 = System.currentTimeMillis();
    for(int i = 0; i < COUNT; i++) {
      res += unpack8ValuesBytes(in, out1, i % out.length);
    }
    long t2 = System.currentTimeMillis();
    System.out.println("[Unpack8ValuesByteArray spent time] " + (t2-t1) + " ms");

    ByteBuffer byteBuffer = ByteBuffer.wrap(in);

    for(int i = 0; i < COUNT; i++) {
      res += unpack8ValuesByteBuffer(byteBuffer, out, i % out.length);
    }

    res = 0;
    long t3 = System.currentTimeMillis();
    for(int i = 0; i < COUNT; i++) {
      res += unpack8ValuesByteBuffer(byteBuffer, out2, i % out.length);
    }
    long t4 = System.currentTimeMillis();
    System.out.println("[Unpack8ValuesByteBuffer spent time] " + (t4-t3) + " ms");

    for (int i=0; i**Note**: *This issue was originally created as [PARQUET-2189](https://issues.apache.org/jira/browse/PARQUET-2189). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Contributor guide

No contributing guide indexed for this repository

Research direction

Start at the abstract BytePacker.unpack8Values overload and review the supplied ByteBufferTest entry point, then identify the implementations affected by changing this API. Compare byte-array and ByteBuffer behavior and benchmark results; done means the byte-array path preserves the unpacked values and its performance benefit is validated.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data, performance
Issue type
Refactor
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.