feast-dev / feast-dev/feast

bug: FeatureStore.materialize() OOM-crashes on large time ranges with no built-in workaround

Open
#6,307 0 comments 0 reactions 0 assignees View on GitHub
kind/feature
Dominant language
Python
Stars
7.3k
Forks
1.4k
Avg merge
1d 21h
Merged PRs (30d)
15

Description

## Summary

`FeatureStore.materialize()` and `FeatureStore.materialize_incremental()` load the entire requested time range into memory in a single pass. On production deployments this causes out-of-memory (OOM) crashes that are silent, difficult to recover from, and currently require external orchestration to work around.

## Steps to reproduce

```python
from datetime import datetime, timedelta

fs.materialize(
start_date=datetime(2026, 1, 1),
end_date=datetime(2026, 3, 1), # ~60-day window
)
```

With a high-frequency feature view (e.g. 10-minute ETL batches → ~8 640 rows/day per entity × many entities), a worker with ≤ 8 GB RAM will exhaust memory and crash with no informative error from Feast.

## Root cause

Both `materialize()` and `materialize_incremental()` delegate to `materialize_single_feature_view` exactly once per feature view, passing the full `[start, end]` window. The underlying data source query materialises the entire range in one shot — there is no pagination, chunking, or streaming at the Feast SDK layer.

## Impact

| Scenario | Effect |
|---|---|
| Multi-day / multi-week backfill | Worker OOM crash |
| Sub-minute sensor data at scale | Worker OOM crash |
| Large number of entities × long window | Worker OOM crash |
| Crash recovery | `most_recent_end_time` is not committed until the entire range succeeds, so a crash forces a full re-run |

## Current workaround

Users must implement their own loop outside Feast:

```python
chunk = timedelta(hours=6)
cursor = start
while cursor < end:
next_cursor = min(cursor + chunk, end)
fs.materialize(cursor, next_cursor)
cursor = next_cursor
```

This is error-prone, not integrated with `materialize_incremental`'s watermark, and must be reimplemented by every affected user.

## Expected behaviour

Feast should provide a native `chunk_size` option so users can cap peak memory usage without external orchestration. Chunking should:

- Be **opt-in and backward-compatible** (default: no chunking).
- Support a **project-level default** in `feature_store.yaml` under `materialization.chunk_size`.
- Commit `most_recent_end_time` **per chunk** so a crash mid-run allows `materialize_incremental` to resume from the last committed chunk.
- Be available via both the **Python SDK** and the **CLI**.

## Related

A PR implementing the above is available for review: https://github.com/feast-dev/feast/pull/6277

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.