apache / apache/arrow-java

[Java][arrow-jdbc] Ability to customize JdbcConsumer construction

Đang mở
#224 1 bình luận 0 reaction 0 người được giao Xem trên GitHub
Type: enhancement
Ngôn ngữ chính
Java
Star
94
Fork
152
Merge trung bình
3 ngày 16 giờ
Pull request đã merge (30 ngày)
11

Mô tả

### Describe the enhancement requested

I'm working on a project that heavily uses the arrow-jdbc library to convert results from a number of different JDBC drivers to Arrow. When the integration was initially written a bit over 2 years ago it forked `ArrowVectorIterator` to add a number of customization points in the JDBC->Arrow conversion which were missing at the time. I'm updating the project and happily in the intervening time it looks like a lot of this customization is now possible natively via customization points in `JdbcToArrowConfig` - in particular the ability to override `JdbcFieldInfo` per column and most importantly the ability to override a `Function`.

I'd love to fully remove our forked `ArrowVectorIterator` but there's one last customization point we're depending on that I don't see a way to handle currently: The ability to customize the `JdbcConsumer` instances constructed for each column. We depend on overriding these to paper over idiosyncrasies with various JDBC drivers. As an example on the top of my mind this morning: The Snowflake JDBC driver [does not support](https://docs.snowflake.com/en/user-guide/jdbc-api.html#id14) `getBinaryStream` and so is not compatible with the default `BinaryConsumer` and must be overriden with a consumer that uses `getBytes` instead. We also add a layer of wrapping to all of our `JdbcConsumers` to wrap error messages to aid debugging (by adding info on the specific column an error originated from).

Would it be possible to add an extension point to the `JdbcToArrowConfig` to customize this, similar to `jdbcToArrowTypeConverter`? I wish we didn't need it but in practice we've found enough variability across JDBC drivers to necessitate it.

### Component(s)

Java

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Hướng nghiên cứu

Bắt đầu bằng cách đọc ArrowVectorIterator và JdbcToArrowConfig, sau đó theo dõi cách các instance của JdbcConsumer được tạo cho từng cột. Xác định một extension point tương tự như jdbcToArrowTypeConverter, cho phép các consumer dành riêng cho driver và wrapping, đồng thời xác minh rằng workaround getBytes của Snowflake có thể được cung cấp thông qua đó.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
java
Lĩnh vực
databases
Loại issue
Tính năng
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
38/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.