tidymodels / tidymodels/embed

`..new` is calculated wrong in lencode steps

Open
#243 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug target encoding tidy-dev-day :nerd_face:
Dominant language
R
Stars
146
Forks
23
PR merge metrics
No merged PRs in 30d

Description

The unseen levels are calculated based on the mean of the coeficients rather than the mean of global data. This should be fixed to better reflect the literature.

Make sure that the documentation is changed accordingly.

This change will be easily backward compatible as it changes how new values will change only.

data <- data.frame(
  outcome = rnorm(1000) + c(rep(10, 900), rep(0, 100)),
  predictor = c(rep("Big", 900), rep(letters[1:10], each = 10))
)

library(tidyverse)

data |>
  count(predictor)
#>    predictor   n
#> 1        Big 900
#> 2          a  10
#> 3          b  10
#> 4          c  10
#> 5          d  10
#> 6          e  10
#> 7          f  10
#> 8          g  10
#> 9          h  10
#> 10         i  10
#> 11         j  10

data |>
  summarize(
    mean = mean(outcome),
    .by = predictor
  )
#>    predictor        mean
#> 1        Big  9.92621834
#> 2          a -0.12884918
#> 3          b  0.24802560
#> 4          c  0.12339453
#> 5          d  0.33307724
#> 6          e  0.08705590
#> 7          f  0.86433875
#> 8          g  0.42452332
#> 9          h  0.42548890
#> 10         i -0.07257279
#> 11         j -0.67403943

embed:::glm_coefs(y = select(data, outcome), x = pull(data, predictor))
#> # A tibble: 12 × 2
#>    ..level ..value
#>    <chr>     <dbl>
#>  1 a       -0.129
#>  2 b        0.248
#>  3 Big      9.93
#>  4 c        0.123
#>  5 d        0.333
#>  6 e        0.0871
#>  7 f        0.864
#>  8 g        0.425
#>  9 h        0.425
#> 10 i       -0.0726
#> 11 j       -0.674
#> 12 ..new    0.256

mean(data$outcome, trim = 0.1)
#> [1] 9.717217

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start at the embed:::glm_coefs entry point shown in the example and trace how the ..new level is calculated. Use the supplied unbalanced-data example to verify that ..new reflects the global trimmed mean rather than the mean of coefficients, then update the related documentation accordingly.

Written by the indexing model from the issue text.

Assessment

Tech stack
r
Domain
documentation, machine-learning
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.