apache / apache/beam

[Task]: Python schema generated types uses schema registry and coder registry

Open
#37,893 1 comment 0 reactions 0 assignees View on GitHub
awaiting triage P2 stale task
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
1d 20h
Merged PRs (30d)
196

Description

### What needs to happen?

Currently a Python Row (encoded with Row Coder) go through serialization/deserialization becomes a schema generated types named tuple. There are many caveats for this behavior

- original type get lost

- #22714

with cloudpickle becomes default and schema registry coder registry saved on pipeline submission, we should be able to use the schema id registered in the schema registry to obtain the user type, then use coder registry for the user type to get registered (row) coder, that makes user_type->GBK still produces user_type

### Issue Priority

Priority: 2 (default / most normal work should be filed as P2)

### Issue Components

- [ ] Component: Python SDK
- [ ] Component: Java SDK
- [ ] Component: Go SDK
- [ ] Component: Typescript SDK
- [ ] Component: IO connector
- [ ] Component: Beam YAML
- [ ] Component: Beam examples
- [ ] Component: Beam playground
- [ ] Component: Beam katas
- [ ] Component: Website
- [ ] Component: Infrastructure
- [ ] Component: Spark Runner
- [ ] Component: Flink Runner
- [ ] Component: Samza Runner
- [ ] Component: Twister2 Runner
- [ ] Component: Hazelcast Jet Runner
- [ ] Component: Google Cloud Dataflow Runner

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.