tidymodels / tidymodels/probably

Problem with names of columns like `probability_*`

Open
#121 2 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

bug reprex
Dominant language
R
Stars
123
Forks
16
PR merge metrics
No merged PRs in 30d

Description

From this Stack Overflow question, this will not work:

library(tidyverse)
library(probably)
#> 
#> Attaching package: 'probably'
#> The following objects are masked from 'package:base':
#> 
#>     as.factor, as.ordered

set.seed(100)
test_df <- tibble(
  probability_x = runif(100),
  Label = as.factor(case_when(probability_x > 0.5 ~ "x", TRUE ~ "y"))
)

cal_plot_breaks(test_df, Label, probability_x)
#> Error in `purrr::map()`:
#> ℹ In index: 2.
#> Caused by error in `estimate_str[[.x]]`:
#> ! subscript out of bounds
#> Backtrace:
#>      ▆
#>   1. ├─probably::cal_plot_breaks(test_df, Label, probability_x)
#>   2. ├─probably:::cal_plot_breaks.data.frame(test_df, Label, probability_x)
#>   3. │ └─probably:::cal_plot_breaks_impl(...)
#>   4. │   ├─probably::.cal_table_breaks(...)
#>   5. │   └─probably:::.cal_table_breaks.data.frame(...)
#>   6. │     └─probably:::.cal_table_breaks_impl(...)
#>   7. │       └─probably:::truth_estimate_map(...)
#>   8. │         └─purrr::map(seq_along(truth_levels), ~sym(estimate_str[[.x]]))
#>   9. │           └─purrr:::map_("list", .x, .f, ..., .progress = .progress)
#>  10. │             ├─purrr:::with_indexed_errors(...)
#>  11. │             │ └─base::withCallingHandlers(...)
#>  12. │             ├─purrr:::call_with_cleanup(...)
#>  13. │             └─probably (local) .f(.x[[i]], ...)
#>  14. │               └─rlang::sym(estimate_str[[.x]])
#>  15. │                 └─rlang::is_symbol(x)
#>  16. └─purrr (local) `<fn>`(`<sbscOOBE>`)
#>  17.   └─cli::cli_abort(...)
#>  18.     └─rlang::abort(...)

Created on 2023-07-06 with reprex v2.0.2

But the same code works if we change the column name to .pred_x:

library(tidyverse)
library(probably)
#> 
#> Attaching package: 'probably'
#> The following objects are masked from 'package:base':
#> 
#>     as.factor, as.ordered

set.seed(100)
test_df <- tibble(
  .pred_x = runif(100),
  Label = as.factor(case_when(.pred_x > 0.5 ~ "x", TRUE ~ "y"))
)

cal_plot_breaks(test_df, Label, .pred_x)

Created on 2023-07-06 with reprex v2.0.2

I see that the docs say:

A vector of column identifiers, or one of dplyr selector functions to choose which variables contains the class probabilities. It defaults to the prefix used by tidymodels (.pred_). The order of the identifiers will be considered the same as the order of the levels of the truth variable.

But it doesn't seem clear that they have to be .pred_x and similar.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the supplied reprex with cal_plot_breaks(), comparing probability_x with .pred_x. Read the cal_plot_breaks entry point and the related documentation to determine whether arbitrary probability-column names should work. Done means the behavior is fixed and covered by a regression check, or the naming requirement is made explicit.

Written by the indexing model from the issue text.

Assessment

Tech stack
r
Domain
data-visualization
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
62/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.