apache / apache/iceberg-rust

avro schema writer does not sanitize field names that violate avro naming rules

Open
#2,535 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Rust
Stars
1.4k
Forks
567
Avg merge
2d 2h
Merged PRs (30d)
93

Description

## problem

when writing manifest files, iceberg-rust copies iceberg field names verbatim into avro record field names. the avro spec requires names to match `[A-Za-z_][A-Za-z0-9_]*`. partition field names that start with digits (e.g. hash-derived names like `815d3b5701b94c78884835c1bea174bb_day`) produce invalid avro schemas.

java handles this correctly in `TypeToSchema.java`:
```java
String origFieldName = structField.name();
boolean isValidFieldName = AvroSchemaUtil.validAvroName(origFieldName);
String fieldName = isValidFieldName ? origFieldName : AvroSchemaUtil.sanitize(origFieldName);
Schema.Field field = new Schema.Field(fieldName, ...);
if (!isValidFieldName) {
field.addProp(AvroSchemaUtil.ICEBERG_FIELD_NAME_PROP, origFieldName);
}
```

sanitization rules (from `AvroSchemaUtil.sanitize()`):
- leading digit → prefix with `_` (e.g. `9col` → `_9col`)
- special chars → `_x` (e.g. `a.b` → `a_x2Eb`)

the original name is preserved in the `iceberg-field-name` avro field property.

## relevant code

`crates/iceberg/src/avro/schema.rs` — `schema_to_avro_schema` uses `field.name.clone()` directly as the avro field name without validation or sanitization.

## impact

- manifests written by iceberg-rust with digit-leading partition field names are invalid avro
- other engines (spark, trino, flink) using strict avro readers will reject these manifests
- any table using hash-based partition transforms can produce such names

## expected behavior

implement the same sanitize-on-write protocol as java:
1. validate name against `[A-Za-z_][A-Za-z0-9_]*`
2. if invalid, sanitize and store original in `iceberg-field-name` property
3. always store `field-id` property (already done)

Contributor guide

Open the contributing guide

Research direction

Start in crates/iceberg/src/avro/schema.rs at schema_to_avro_schema, where field.name is copied into the Avro schema. Compare the required validation and sanitization rules with the Java behavior described in the issue. Done means invalid names are sanitized, original names are preserved in iceberg-field-name, field-id remains present, and generated schemas are accepted by strict Avro readers.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
data-engineering
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
75/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.