duckdb / duckdb/duckdb-python

DuckDB Spark API is incompatible with the PySpark API's `spark.createDataFrame(list of dict)`

Open
#183 0 comments 0 reactions 0 assignees View on GitHub
needs triage
Dominant language
Python
Stars
187
Forks
112
Avg merge
13h 29m
Merged PRs (30d)
17

Description

### What happens?

The DuckDB Spark API is incompatible with the PySpark API's `spark.createDataFrame(list of dict)` method.

### To Reproduce

```python
from duckdb.experimental.spark.sql import SparkSession as DuckdbSparkSession
from pyspark.sql import SparkSession

sql_text = "SELECT * FROM t0"
data = [
{"c0": "1969-12-21"}
]
spark = SparkSession.builder.getOrCreate()
df = spark.createDataFrame(data)
df.createOrReplaceTempView("t0")

print("PySpark SQL result:")
pyspark_result = spark.sql(sql_text)
pyspark_result.show()

duckdb_spark = DuckdbSparkSession.builder.getOrCreate()
df = duckdb_spark.createDataFrame(data)
df.createOrReplaceTempView("t0")

print("Duckdb Spark SQL result: ")
duckdb_spark_result = duckdb_spark.sql(sql_text)
duckdb_spark_result.show()
```
```bash
PySpark SQL result:
+----------+
| c0|
+----------+
|1969-12-21|
+----------+

Duckdb Spark SQL result:
┌─────────┐
│ col0 │
│ varchar │
├─────────┤
│ c0 │
└─────────┘
```

### OS:

x86_64 Ubuntu 24.04 Linux-6.14.0-35-generic-x86_64-with-glibc2.39

### DuckDB Version:

1.4.2

### DuckDB Client:

Python

### Hardware:

_No response_

### Full Name:

asddfl

### Affiliation:

xxx

### Did you include all relevant configuration (e.g., CPU architecture, Linux distribution) to reproduce the issue?

- [x] Yes, I have

### Did you include all code required to reproduce the issue?

- [x] Yes, I have

### Did you include all relevant data sets for reproducing the issue?

Yes

Contributor guide

Open the contributing guide

Research direction

Start at duckdb.experimental.spark.sql.SparkSession.createDataFrame and run the provided Python reproduction against the DuckDB Spark API and PySpark. Compare the resulting temporary-view output, and consider the issue complete when DuckDB produces the same column name and row value as PySpark for a list of dictionaries.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.