JuliaPy / JuliaPy/PythonCall.jl
pandas.Categorical not preserved during DataFrame conversion and jl.convert fails
- 主要語言
- Julia
- 星號
- 1.1k
- 分支
- 86
- 平均合併
- 1 天 22 小時
- 30 天內合併 PR
- 3
描述
Hi and thanks for the great package!
While working with `juliacall` + `PythonCall.jl`, I ran into two issues related to pandas.Categorical handling.
(In all examples below, `jl`refers to Main from `juliacall`, i.e., f`rom juliacall import Main as jl`.)
**DataFrame conversion ignores Categorical columns**
When passing a `pandas.DataFrame` with categorical columns (i.e., `dtype='category'`), those columns are silently converted to `Int64` vectors in Julia (presumably the .codes). This results in CategoricalArray semantics being lost — so interactions in Julia formulas like `id & η1` are treated as numeric rather than generating dummy variables.
**`jl.convert()` can’t convert pandas.Categorical to any Julia type**
I tried using `jl.convert(CategoricalArray, col)` directly on a `pandas.Series` with categorical dtype, but got a `MethodError`. It appears `PythonCall` doesn’t yet support converting `pandas.Categorical` to any Julia-native type.
To work around this, I convert the column to `str` in Python (so it arrives as a `Vector{String}`), then manually wrap it in `categorical(...)` on the Julia side. This works, but it's not ideal for type fidelity or automatic translation.
Let me know if there's a cleaner workaround — or if you'd be open to a PR to improve automatic `CategoricalArray` support.
Thanks again!
貢獻指南
這個儲存庫沒有索引到貢獻指南
評估
這個 Issue 還沒有評估資料。