autogluon / autogluon/autogluon

Custom model that reshapes features in _fit(): confusing KeyError on the bagged predict path — what's the supported pattern?

Open
#5,898 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
10.7k
Forks
1.2k
Avg merge
21h 29m
Merged PRs (30d)
57

Description

**Question / rough-edge report** (not sure this is a bug vs. me holding the API wrong — happy to be told it's the latter).

## Summary

A custom `AbstractModel` that applies a representation-changing feature transform
(raw spectrum → a smaller set of derived components) works fine un-bagged, and its
`fit()` + out-of-fold scoring succeed when bagged, but `predictor.predict()` on a
**bagged** model then fails with:

```
KeyError: "None of [Index(['feature_0', ..., 'feature_N'], dtype='object')] are in the [columns]"
```

I think the underlying cause is on my side (see "What I think is going on"), but the
failure is opaque and only shows up on the bagged predict path, so I wanted to (a) confirm
the intended pattern and (b) ask whether a clearer error or a supported hook would make
sense.

## What the model does

The transform maps the ~1000 raw input columns to `K` derived features (`K` fixed at fit;
deterministic, same width on train and test). I implemented it by:

- overriding `_fit`: transform `X`, then re-run `self._preprocess_set_features(X_transformed)`
so `self.features` matches the transformed columns (`['feature_0', ..., 'feature_{K-1}']`),
then fit the underlying estimator;
- overriding `_predict_proba`: apply the same transform, then predict.

The `_preprocess_set_features` re-sync was added because some underlying models (e.g.
`KNNModel`) call `self.preprocess(X)` inside their own `_fit`, and by then `X` is already
transformed while `self.features` still holds the original column names → `KeyError` at fit
time. Re-syncing fixes that.

## What I think is going on

`AbstractModel.fit()` snapshots `self.features` against the **original** columns before
`_fit()` runs. My `_fit()` then overwrites `self.features` to the transformed names. That's
what the bagged predict path trips over:

```
BaggedEnsembleModel._predict_proba_internal
-> self.preprocess(X_raw, model=child)
-> child.preprocess(X_raw, preprocess_stateful=False)
-> AbstractModel._preprocess_nonadaptive(X_raw)
-> if list(X_raw.columns) != self.features:
X_raw = X_raw[self.features] # self.features == ['feature_0', ...]
# X_raw still has the ORIGINAL columns -> KeyError
```

The bagged ensemble (correctly) passes the raw input to the child's *non-adaptive*
`preprocess` and never routes through my `_predict_proba`, so `self.features` — which I
mutated to describe the post-transform space — no longer matches the raw frame it's being
applied to. Un-bagged predict and the bagged OOF path both go through my `_predict_proba`,
so they're fine.

Running the repro below with the re-sync disabled (`--no-resync`) makes bagged predict
succeed, at the cost of re-breaking the fit-time `KNNModel` case.

## Questions

1. Is the supported way to do a representation-changing transform in a custom model to put
it in `_preprocess` (so it runs per-child at fit and predict and AutoGluon tracks the
resulting feature space), and to **never** touch `self.features`? If so it'd be great to
have that stated in the custom-model docs — the natural-looking place (`_fit` /
`_predict`) is a trap here.
2. Is there a supported way for a model to declare "my inference-time input columns differ
from what I fit the underlying estimator on", so `_preprocess_nonadaptive`'s
`X[self.features]` projection is skipped for it?
3. Would a clearer error at the `X = X[self.features]` line (e.g. "model X's `self.features`
are not a subset of the input columns — did a custom model mutate `self.features`?") be
worth adding? The current `KeyError` from pandas took a while to trace to a mutated
`self.features`.

## Reproducer

Self-contained, scikit-learn only.

```python
"""Bagged custom model that reshapes its input in _fit and re-syncs self.features
-> KeyError at predict time. `--no-resync` predicts fine (but breaks fit-time
models that call self.preprocess() in their own _fit, e.g. KNNModel)."""
from __future__ import annotations
import argparse, shutil, sys, tempfile
import numpy as np, pandas as pd
from sklearn.linear_model import Ridge
from autogluon.core.models import AbstractModel
from autogluon.tabular import TabularPredictor

N_ROWS, N_COLS = 400, 60
RESYNC = True # flipped by --no-resync

class WidthChangingModel(AbstractModel):
"""Preprocessing halves the feature count (stand-in for an NMF / concat / projection step)."""
ag_key = "WIDTHCHG"
ag_name = "WidthChanging"

@staticmethod
def _transform(X: pd.DataFrame) -> pd.DataFrame:
arr = X.to_numpy(dtype=float)[:, ::2] # width halves; deterministic, same on train/test
return pd.DataFrame(arr, columns=[f"feature_{i}" for i in range(arr.shape[1])], index=X.index)

def _resync_features(self, X: pd.DataFrame) -> None:
self.features = list(X.columns)
self.feature_metadata = None
self._preprocess_set_features(X)

def _fit(self, X, y, **kwargs):
X_t = self._transform(X)
if RESYNC and list(X_t.columns) != list(X.columns):
self._resync_features(X_t)
self.model = Ridge().fit(X_t.to_numpy(), np.asarray(y, dtype=float))

def _predict_proba(self, X, **kwargs):
return self.model.predict(self._transform(X).to_numpy())

@classmethod
def supported_problem_types(cls):
return ["regression"]

def make_data(seed: int = 0) -> pd.DataFrame:
rng = np.random.default_rng(seed)
X = rng.normal(size=(N_ROWS, N_COLS))
y = X[:, ::2].sum(axis=1) + rng.normal(scale=0.1, size=N_ROWS)
df = pd.DataFrame(X, columns=[f"ch_{i:03d}" for i in range(N_COLS)])
df["target"] = y
return df

def main() -> int:
global RESYNC
ap = argparse.ArgumentParser()
ap.add_argument("--no-resync", action="store_true")
RESYNC = not ap.parse_args().no_resync

df = make_data()
train, test = df.iloc[:320], df.iloc[320:].drop(columns=["target"])
tmp = tempfile.mkdtemp(prefix="ag_repro_")
try:
predictor = TabularPredictor(label="target", problem_type="regression", path=tmp, verbosity=1)
predictor.fit(
train_data=train,
hyperparameters={WidthChangingModel: {
# sequential_local only avoids pickling this __main__ class across ray workers
"ag_args_ensemble": {"fold_fitting_strategy": "sequential_local"},
}},
num_bag_folds=2, # bagging is required to trigger it
num_bag_sets=1,
fit_weighted_ensemble=False,
)
print(f"\nfit + OOF scoring OK (resync={'on' if RESYNC else 'off'})")
preds = predictor.predict(test)
print(f"predict OK - not reproduced\n{preds.head(3).to_string()}")
return 1 if RESYNC else 0
except KeyError as e:
print(f"\n>>> REPRODUCED - KeyError: {str(e)[:120]}...")
return 0
finally:
shutil.rmtree(tmp, ignore_errors=True)

if __name__ == "__main__":
sys.exit(main())
```

```
python repro.py # -> >>> REPRODUCED - KeyError: "None of [Index(['feature_0', ...
python repro.py --no-resync # -> predict OK
```

## Traceback

```
File ".../autogluon/tabular/trainer/abstract_trainer.py", line 1364, in get_model_pred_proba_dict
model_pred_proba_dict[model_name] = model.predict_proba(X, **preprocess_kwargs)
File ".../autogluon/core/models/abstract/abstract_model.py", line 1577, in predict_proba
y_pred_proba = self._predict_proba_internal(X=X, normalize=normalize, **kwargs)
File ".../autogluon/core/models/ensemble/bagged_ensemble_model.py", line 629, in _predict_proba_internal
X = self.preprocess(X, model=model, **kwargs)
File ".../autogluon/core/models/ensemble/stacker_ensemble_model.py", line 274, in preprocess
X = super().preprocess(X, **kwargs)
File ".../autogluon/core/models/ensemble/bagged_ensemble_model.py", line 227, in preprocess
return model.preprocess(X, preprocess_stateful=False)
File ".../autogluon/core/models/abstract/abstract_model.py", line 598, in preprocess
X = self._preprocess_nonadaptive(X, **kwargs)
File ".../autogluon/core/models/abstract/abstract_model.py", line 674, in _preprocess_nonadaptive
X = X[self.features]
File ".../pandas/core/frame.py", line 4119, in __getitem__
...
KeyError: "None of [Index(['feature_0', 'feature_1', ..., 'feature_29'], dtype='object')] are in the [columns]"
```

(Line numbers from `1.5.1b20260731`; same frames on `1.6.2b20260809`.)

## Installed versions

```
autogluon.core / .tabular : 1.5.1b20260731 (also repro'd on 1.6.2b20260809)
python : 3.12.13
platform : macOS-26.6.2-arm64 / also Linux
pandas 2.3.3 | numpy 2.3.5 | scikit-learn 1.7.2
```

Contributor guide

Open the contributing guide

Research direction

Start by running the self-contained repro.py with and without --no-resync to confirm the bagged prediction difference. Read autogluon/core/models/abstract/abstract_model.py around preprocess, _preprocess_nonadaptive, and _fit, then follow the child preprocessing call in autogluon/core/models/ensemble/bagged_ensemble_model.py. Done means establishing the supported custom-model pattern and adding the documented behavior, hook, or clearer failure described by the maintainers.

Written by the indexing model from the issue text.

Assessment

Tech stack
numpy, pandas, python, scikit-learn
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.