cmu-delphi / cmu-delphi/epipredict

Eliminate/explicate differences in training windowing between flatline and arx forecasters

Open
#321 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
R
Stars
18
Forks
13
Avg merge
21d 58m
Merged PRs (30d)
1

Description

#290 highlighted that training window sizes similar to the ahead value can trip up the flatline forecaster. But this also indicates that the flatline forecaster is not using anywhere near `n_training` instances per epikey if `ahead` is within an order of magnitude of `n_training`. This is not the case for `arx_forecaster`:

``` r
library(epipredict)
#> Loading required package: epiprocess
#>
#> Attaching package: 'epiprocess'
#> The following object is masked from 'package:stats':
#>
#> filter
#> Loading required package: parsnip
trace(slather, quote({
if (inherits(object, "layer_residual_quantiles")) {
trace(dplyr::summarize, quote({
cat("Number of non-NA residuals:\n")
print(.data %>% tidyr::drop_na(.resid) %>% nrow())
}))
}
}), quote(untrace(dplyr::summarize)))
#> Tracing function "slather" in package "epipredict"
#> [1] "slather"
case_death_rate_subset %>% flatline_forecaster("case_rate", flatline_args_list(ahead = 28L, n_training = 29L))
#> [...]
#> Number of non-NA residuals:
#> [1] 56
#> [...]
case_death_rate_subset %>% arx_forecaster("case_rate", "case_rate", args_list = arx_args_list(ahead = 28L, n_training = 29L))
#> [...]
#> Number of non-NA residuals:
#> [1] 1624
#> [...]
```

Created on 2024-04-19 with [reprex v2.0.2](https://reprex.tidyverse.org)

However, `?flatline_args_list` doesn't explicate this
```
n_training: Integer. An upper limit for the number of rows per key that
are used for training (in the time unit of the 'epi_df').
```
and the message from `slather.layer_residual_quantiles` when output residuals are NA is something specific to flatline forecaster (and off by one for `flatline_forecaster`):
```
! Residual quantiles could not be calculated due to missing residuals.
ℹ This may be due to `n_train` < `ahead` in your .
```

Approach 1: eliminate these differences. Make `n_training` make sense for `flatline_forecaster` by using the same NA omission pre training window approach as `arx_forecaster`. Remove the mention of the inequality above in the `layer_residual_quantiles` error message since it won't be an issue anymore.

Approach 2: explain the difference in `?flatline_args_list`, and mention `n_train` --> **`<=`** <-- ` ahead` is an issue --> **for `flatline_forecaster`** <-- in the residual quantiles error message.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.