[Question] Optimize multiple reads on same DataFrame
- Dominant language
- Rust
- Stars
- 9.3k
- Forks
- 2.4k
- Avg merge
- 3d 7h
- Merged PRs (30d)
- 344
Description
Hey,
I have a scenario where I have to run the same filter expression but with different values on the same RecordBatch
For example
```
let c2: Vec = ....
let provider = datafusion::datasource::MemTable::try_new(c2[0].schema(), vec![c2])
.map_err(|e| {
log::error!("Error MemTable {}", e);
e
})
.unwrap();
let ctx = SessionContext::new();
ctx.register_table("t", provider ).unwrap();
let df = ctx.table("t").unwrap();
let expr: Expr = get_expression(id, from_time, to_time)
let df = df.filter(expr).unwrap();
let res = df.collect().await.unwrap();
ctx.deregister_table("t").unwrap();
```
It is pretty fast, a few ms on a 80MiB in-memory array with filtering on 2 columns.
I might run 1000 queries on the same MemTable and was wondering if there is anything that could be optimized:
- pre computing an execution plan on the MemTable if it's cost effective
- Is SessionContext thread safe and shareable between multiple threads and be optimized across executions?
- Somehow create an index (not sure if an index is created by one of the calls or supported at all) if it's cost effective
Thanks!
Contributor guide
Research direction
Start by reading the MemTable, SessionContext, DataFrame::filter, and DataFrame::collect entry points mentioned in the question. Determine whether repeated filtered reads can reuse planning or execution work, whether SessionContext supports the requested concurrent use, and whether indexing is supported; done would be a documented answer or a scoped enhancement proposal.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust
- Domain
- databases
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100