apache / apache/arrow

Expose built_in columns using ds.field() in datasets

Open
#38,518 1 comment 0 reactions 0 assignees View on GitHub
Component: Python Type: enhancement
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 13h
Merged PRs (30d)
88

Description

### Describe the enhancement requested

For all dataset/scanner functions like to_table(), head(), etc.. which support @columns..

This works..
`to_table(columns=["my_column", "__filename"])`

But built-in columns are not supported using expressions:
`to_table(columns = {"my_column": ds.field("my_column") , "filename": ds.field("__filename")})`

I'm trying to create an expressions to parse out YYYYMMDD from the following filename:
VAR_pnl.20230829.csv.gz
But I need to be able to reference the built-in column "__filename" in the expression..

My only other option is to extract everything into memory first using @columns= ["a","b", "__filename"], but this isn't ideal for memory management.

### Component(s)

Python

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.