snowflakedb / snowflakedb/snowflake-connector-python

SNOW-701383: Upload Arrow dataset to Snowflake

Open
#1,358 8 comments 20 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

feature status-triage_done
Dominant language
Python
Stars
730
Forks
574
Avg merge
5h 45m
Merged PRs (30d)
16

Description

What is the current behavior?

I do not think there is a way to upload an arrow dataset to Snowflake via the snowflake-connector.

What is the desired behavior?

Given an arrow dataframe, I would like to upload it to Snowflake without going via Pandas (which would result in the valuecolumn being cast to float, since numpy integer types don't support null values).

Here is an example of a frame I would like to upload to Snowflake, by creating or appending to a table.

from datetime import date
import pyarrow as pa
id = pa.array(["A", "B", "E"])
date_day = pa.array([date(2022, 1, 1), date(2022, 1, 2), date(2022, 1, 3)])
value = pa.array([2, 4, None])
columns = ["id", "date_day", "value"]

df = pa.Table.from_arrays([id, date_day, value], names=columns)
df
----
pyarrow.Table
id: string
date_day: date32[day]
value: int64
----
id: [["A","B","E"]]
date_day: [[2022-01-01,2022-01-02,2022-01-03]]
value: [[2,4,null]]

Losing datatype information

print(df.to_pandas())
--- # Note that the `value` column now is floating point
  id    date_day  value
0  A  2022-01-01    2.0
1  B  2022-01-02    4.0
2  E  2022-01-03    NaN

How would this improve snowflake-connector-python?

Arrow is the next-generation data format for dataframe types. Libraries like polars already take great advantage of it, and it would be good to support this in order to allow for more correct data interactions with Snowflake.

Currently, Snowflake already allows for fetching data in the arrow format, but as far as I can see, not write data.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating the connector's existing Arrow-fetch support and the upload or table-creation entry points. Use the provided pyarrow.Table example to trace how create and append operations could accept Arrow data without converting through Pandas. Done means nullable integer and other Arrow datatype information is preserved when uploading to Snowflake.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
databases
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.