Improve memory consumption of AggregateNumericRangeEquality

Open
#16 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
35/100
Issue type
Refactor
Clarity
Mostly clear
Activity status
Stale
Tech stack
python
Domain
databases

Research direction

Locate AggregateNumericRangeEquality and the .fetchall() call first, then inspect how the database results are consumed. The change is done when checks process rows without retaining the full result set and the existing equality behavior remains covered by the relevant tests.

Written by the indexing model from the issue text.

Description

AggregateNumericRangeEquality requires ~ 20 GiB of memory (~ 50 M rows).

.fetchall() returns a list. Could we change this to perform the checks in a streaming fashion that doesn't require all of the data in memory at once?

Dominant language
Python
Stars
46
Forks
3
PR merge metrics
No merged PRs in 30d

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Quantco/datajudge

All issues in Quantco/datajudge

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.