G-Research / G-Research/ParquetSharp.DataFrame

Need to be able to write IEnumerable<DataFrame> to disk

Open
#4 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
C#
Stars
26
Forks
12
Avg merge
33m
Merged PRs (30d)
1

Description

Thanks for creating this project!

Implement writing multiuple dataframes into the same file

Current implementation supports reading of a large file into several dataframes, however, there is no way to write a large amount of data to a file, that might span several dataframes. Assuming all the dataframes have the same schema, how can we write them to the same Apache Parquet file?

Why is this useful?

Whenever you have a large amount of data that is partitionable, it should be processed independently for each partition. Then you want to be able to parallelise the computation and avoid blowing up memory. Yet you might want to be able save output of each partition to the same file.

Scenario:
  • Open Large Parquet file.
  • Parallel.For each rowgroup
    • read it into a data frame
    • process dataframe
    • Write each dataframe consecutively into disk.

Sample Usage of the new Interface:

// get data
IEnumerable<DataFrame> dataframes = GetDataFrames();
dataframes = ProcessOneDataFrameAtATime(dataframes);

// Extension Method for IEnumerable<DataFrame>
using var propertiesBuilder = new WriterPropertiesBuilder();
propertiesBuilder.Compression(Compression.Snappy);
using var properties = propertiesBuilder.Build();

// Extension for IEnumerable<DataFrame>
// can assume all frames have same schema.
dataframes.ToParquet(parquet_file_path, properties);

// Iterative interface is useful you when need to write one at a time. In general,
// if you have only one implementation -- do this one as it is more general.
using var parquetWriter = new ParquetFileWriter(parquet_file_path);
foreach(var dataframe in dataframes)
{
       parquetWriter.Write(dataframe)
}

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating the existing DataFrame Parquet read/write entry points and inspect the APIs represented by ToParquet, ParquetFileWriter, and WriterPropertiesBuilder. Implement support for writing an IEnumerable with a shared schema, both through the extension method and iterative writer, then verify that multiple frames are written consecutively to one Apache Parquet file.

Written by the indexing model from the issue text.

Assessment

Tech stack
csharp
Domain
data
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.