cmu-delphi / cmu-delphi/forecast-eval
Prediction cards: Memory and runtime optimization
- Dominant language
- R
- Stars
- 6
- Forks
- 3
- PR merge metrics
- No merged PRs in 30d
Description
I ran a mem+cpu profile of `create_prediction_cards` using profvis to investigate high memory use and long runtimes when building the prediction cards.
[profiler_run.zip](https://github.com/cmu-delphi/forecast-eval/files/6173083/profiler_run.zip)
You should be able to open the profile_run in RStudio, or just open it in a browser (it's html) after unzipping. Beware because of the long run-time, it's huge.
# Findings
There are three main sections:

1. `get_covidhub_predictions`: Which is downloading the predictions from github and loading/processing the CSVs into memory
2. Filtering the combined file, most of which is spent in:
```
# Only accept forecasts made Monday or earlier
predictions_cards = predictions_cards %>%
filter(target_end_date - (forecast_date + 7 * ahead) >= -2)
```
3. Saving the RDS file
In `get_covidhub_predictions` there are many blocks of processing for each forecaster/epiweek, which are split into two parts
1. An HTTP call to grab the data
2. CSV processing

# Recommendation
1. Create a pipline that processes data on a per-forecaster basis, instead of batching operations for all forecasters.
2. Parallelize processing multiple forecasters at once
3. Merge them as a last step into one predictions_cards file and save.
4. Evaluate filter operations to see which filters can't be applied sooner, perhaps when loading the CSVs.
(1) should present a more constant memory usage (instead of linear WRT # of forecasters currently). GC should be good at reclaiming memory after a particular forecaster is complete. Because of the amount of I/O in (1), the task should be able to take advantage of multi-processing. The only consideration I can think of is how to break the `rbind` of old and new cards on a per-forecaster basis.
Merging and saving the final result will still take linear time with number of forecasters, but GC should've pruned intermediary objects from step (1) and there should be room to do this in-memory.
Finally, evaluate the filter operations to see which can't be moved up sooner to make sure we're not processing a lot of data that is later pruned.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by opening profiler_run from profiler_run.zip in RStudio or a browser, then inspect create_prediction_cards and get_covidhub_predictions. Trace the download, CSV processing, filtering, merging, and RDS-saving steps identified in the profile. Done means reducing peak memory and runtime through per-forecaster processing, appropriate parallelization, and earlier applicable filtering while preserving the final predictions_cards output.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- r
- Domain
- data, performance
- Issue type
- Refactor
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100