apache / apache/datafusion

Error when joining dataframes with duplicate column names if dataframes generated from file

Open
#14,147 4 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
Rust
Stars
9.3k
Forks
2.4k
Avg merge
3d 7h
Merged PRs (30d)
344

Description

### Describe the bug

Encountered an issue joining dataframes with duplicate column names if they generated from file read (I tried csv and parquet).
Dataframes produced from python dict do join without problem.

I did my testing with latest version of Datafusion on Windows.

### To Reproduce

Fine with dataframes from dict
```
from datafusion import SessionContext
ctx = SessionContext()
x1 = ctx.from_pydict({'id1': [1, 2, 4, 5, 6], 'col2': [3, 4, 3, 5, 2], 'col3': [3, 4, 1, 2, 3]})
x2 = ctx.from_pydict({'id1': [1, 2, 4, 5, 6], 'col2': [3, 4, 3, 5, 2], 'col3': [5, 6, 7, 8, 9]})
x1.join(x2, on="id1")
Out[16]:
DataFrame()
+-----+------+------+-----+------+------+
| id1 | col2 | col3 | id1 | col2 | col3 |
+-----+------+------+-----+------+------+
| 1 | 3 | 3 | 1 | 3 | 5 |
| 2 | 4 | 4 | 2 | 4 | 6 |
| 4 | 3 | 1 | 4 | 3 | 7 |
| 5 | 5 | 2 | 5 | 5 | 8 |
| 6 | 2 | 3 | 6 | 2 | 9 |
+-----+------+------+-----+------+------+
```

Continue to file read
```
x1.write_csv("df1.csv")
x2.write_csv("df2.csv")

x1_f = ctx.read_csv("df1.csv")
x2_f = ctx.read_csv("df2.csv")

x1_f.join(x2_f, on="id1")
---------------------------------------------------------------------------
Exception Traceback (most recent call last)
Cell In[21], line 1
----> 1 x1_f.join(x2_f, on="id1")

File ~\prj\datafusion_test\venv\Lib\site-packages\datafusion\dataframe.py:468, in DataFrame.join(self, right, on, how, left_on, right_on, join_keys)
465 if isinstance(right_on, str):
466 right_on = [right_on]
--> 468 return DataFrame(self.df.join(right.df, how, left_on, right_on))

Exception: Schema error: No field named id1. Valid fields are "?table?"."1", "?table?"."3".
```

### Expected behavior

_No response_

### Additional context

_No response_

Contributor guide

Open the contributing guide

Research direction

Start with the Python binding shown in datafusion/dataframe.py at DataFrame.join, then reproduce the failure using ctx.read_csv with the two generated files. Trace how file-backed schemas are exposed to the join and compare them with from_pydict; done means the file-read dataframes join on id1 without a schema error.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, rust
Domain
data-engineering
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.