duckdb / duckdb/duckdb-java

Partitioned Export with COPY Generates Excessively Many Small Files within Each Partition

Open
#152 1 comment 1 reaction 0 assignees View on GitHub
Dominant language
C++
Stars
127
Forks
80
Avg merge
13h 49m
Merged PRs (30d)
48

Description

**Duckdb Version**: 1.1.3 and 1.2.0

When using the COPY command to export data into Parquet format with a partition key (e.g., batch_id), DuckDB produces multiple very small files (approximately 50–100KB each). In our use case—exporting around 100 million line items (1 crore) into batches of 5,000 line items each—this results in roughly 20,000 separate partition directories and small file fragments within each.
**Current output**
```
\batch_id=1000\
records_0.parquet (30kb)
records_1.parquet(70kb)
```

**Expected output**
```
\batch_id=1000\
records_0.parquet (100kb)
```

**Command used to export :**
```sql
COPY my_schema.my_table
TO '/path/to/export'
(FORMAT 'parquet', PARTITION_BY (batch_id), OVERWRITE_OR_IGNORE, FILENAME_PATTERN 'records_{i}');
```

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.