apache / apache/datafusion

[Question] Optimize multiple reads on same DataFrame

Open
#2,845 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Rust
Stars
9.3k
Forks
2.4k
Avg merge
3d 7h
Merged PRs (30d)
344

Description

Hey,

I have a scenario where I have to run the same filter expression but with different values on the same RecordBatch

For example

```
let c2: Vec = ....
let provider = datafusion::datasource::MemTable::try_new(c2[0].schema(), vec![c2])
.map_err(|e| {
log::error!("Error MemTable {}", e);
e
})
.unwrap();

let ctx = SessionContext::new();

ctx.register_table("t", provider ).unwrap();
let df = ctx.table("t").unwrap();

let expr: Expr = get_expression(id, from_time, to_time)

let df = df.filter(expr).unwrap();

let res = df.collect().await.unwrap();
ctx.deregister_table("t").unwrap();
```

It is pretty fast, a few ms on a 80MiB in-memory array with filtering on 2 columns.
I might run 1000 queries on the same MemTable and was wondering if there is anything that could be optimized:

- pre computing an execution plan on the MemTable if it's cost effective
- Is SessionContext thread safe and shareable between multiple threads and be optimized across executions?
- Somehow create an index (not sure if an index is created by one of the calls or supported at all) if it's cost effective

Thanks!

Contributor guide

Open the contributing guide

Research direction

Start by reading the MemTable, SessionContext, DataFrame::filter, and DataFrame::collect entry points mentioned in the question. Determine whether repeated filtered reads can reuse planning or execution work, whether SessionContext supports the requested concurrent use, and whether indexing is supported; done would be a documented answer or a scoped enhancement proposal.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
databases
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.