apache / apache/parquet-java

[C++] Data set integrity tool

Offen
#2,251 2 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Component: C++ Component: Java Component: Parquet Priority: Major Type: enhancement
Vorherrschende Sprache
Java
Sterne
3.1k
Forks
1.6k
Ø Merge
3 T. 12 Std.
Gemergte PRs (30 T.)
33

Beschreibung

Parquet encryption protects integrity of individual files. However, data sets (such as tables) are often written as a collection of files, say

"/path/to/dataset"/part0.parquet.encrypted

..

"/path/to/dataset"/partN.parquet.encrypted

 

In an untrusted storage, removal of one or more files will go unnoticed. Replacement of one file contents with another will go unnoticed, unless a user has provided unique AAD prefixes for each file.

 

The data set integrity tool solves these problems. While it doesn't necessarily belong in Parquet functionality (that is focused on individual files (?)) - it will assist higher level frameworks that use Parquet, to cryptographically protect integrity of data sets comprised of multiple files.

The use of this tool is not obligatory, as frameworks can use other means to verify table (file collection) integrity.

 

The tool works by creating a small file, that can be stored as say

"/path/to/dataset"/.dataset.signature

 

that contains the dataset unique name (URI) and the number of files. It can also contain an explicit list of file names (with or without full path). The file contents is either encrypted with AES-GCM  (authenticated, encrypted) - or hashed and signed (authenticated, plaintext). 

 

On the writer side, the tools creates AAD prefixes for every data file, and creates the signature file itself. The input is the dataset URI, N and the encryption/signature key; plus (optionally) the list of file names (with or without full path).

 

On the reader side, the tool parses and verifies the signature file, and provides the framework with the verified dataset name, number of files that must be accounted for, and the AAD prefix for each file;  plus (optionally) the list of file names (with or without full path). The input is the expected dataset URI and the encryption/signature key.

 

 

 

**Reporter**: [Gidon Gershinsky](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=gershinsky) / @ggershinsky
**Assignee**: [Gidon Gershinsky](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=gershinsky) / @ggershinsky
#### Related issues:
- [Parquet modular encryption](https://github.com/apache/parquet-java/issues/2110) (depends upon)

**Note**: *This issue was originally created as [PARQUET-1457](https://issues.apache.org/jira/browse/PARQUET-1457). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Es wird keine Quelldatei, kein Test und kein Einstiegspunkt für die Implementierung genannt. Beginne damit, das zugehörige Parquet-Problem zur modularen Verschlüsselung zu prüfen, und bestimme, wo ein Tool auf Datensatzebene hingehört. Der Abschluss sollte ein definiertes Writer- und Reader-Design umfassen, das die Dataset-URI, die Dateianzahl oder -liste, die Authentifizierung und AAD-Präfixe pro Datei abdeckt.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
cpp
Bereich
cryptography, data-engineering
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
20/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.