bevyengine / bevyengine/bevy

App hang on AppExit when scenario adds per-frame system or overwrites init_resource (Bevy 0.18.1, macOS Metal)

Open
#24,035 4 comments 1 reaction 0 assignees View on GitHub
A-App C-Bug C-Testing S-Needs-Reproduction
Dominant language
Rust
Stars
48.2k
Forks
4.8k
Avg merge
3d 16h
Merged PRs (30d)
171

Description

## Bevy version

`0.18.1` (workspace-pinned in our project's `Cargo.toml`)

## Platform

- macOS 26.3.1 (Apple Silicon, M1 Max)
- Metal backend (wgpu)
- `cargo 1.97.0-nightly`

## What I'm doing

Writing a small "scenario harness" that boots the full Bevy app to capture screenshots of game state for visual QA. The harness queues several `Screenshot::primary_window()` entities at known frame counts (4 captures across a ~360-frame run), then sends `AppExit::Success` via `MessageWriter` once everything has been captured.

## What I expect

The app exits cleanly within a frame or two of `AppExit::Success` being sent, the same way a single-screenshot scenario does.

## What actually happens

The `AppExit::Success` write succeeds (I can see the log line from the writer's surrounding system), but the app **never terminates** — it idles at 0% CPU indefinitely. Has to be SIGKILLed.

I've isolated this to two specific user actions that each independently trigger the hang. Either alone causes it; both together also trigger it.

### Trigger 1 — Adding ANY new per-frame system to a scenario after multiple `Screenshot` entities have queued

A scenario does:

```rust
super::schedule_screenshot(app, 60, output_dir.join("a.png"));
super::schedule_screenshot(app, 120, output_dir.join("b.png"));
super::schedule_screenshot(app, 180, output_dir.join("c.png"));
super::schedule_screenshot(app, 240, output_dir.join("d.png"));
super::schedule_exit(app, 360);

app.add_systems(Update, force_phase_per_frame); // <-- this line triggers the hang
```

Where `force_phase_per_frame` does NOTHING relevant to the screenshots — it just writes a `Resource` value:

```rust
fn force_phase_per_frame(mut solar: ResMut, frame: Res) {
solar.phase = match frame.0 {
n if n < 120 => 0.0,
n if n < 180 => FRAC_PI_2,
n if n < 240 => PI,
_ => 3.0 * FRAC_PI_2,
};
}
```

I also tried registering it in `PreUpdate` instead of `Update` — same hang. I also tried writing to a different (otherwise unused) `Resource` from the system body — same hang. The hang is triggered by the **act of registering an additional per-frame system in a scenario that has multiple queued screenshot entities**, regardless of what the system body does.

Removing this `add_systems` call (and dropping the per-frame override mechanism) makes the test pass cleanly in ~6 seconds.

### Trigger 2 — `app.insert_resource(X)` overwriting a `Resource` already `init_resource`'d by a `Plugin::build`

A different scenario configures the camera target like this:

```rust
app.insert_resource(CameraTarget {
position: Vec3::new(0.0, 0.0, 25.0),
view_height: MAX_VIEW_HEIGHT,
});
```

Where `CameraTarget` was already `init_resource::()`'d in our `RenderPlugin::build`. Same hang signature: screenshots all save, `AppExit::Success` log line appears, app then idles forever.

Replacing `insert_resource` with a `Startup` system that mutates `CameraTarget` via `ResMut` (so the same instance is updated rather than replaced) avoids the hang.

## Workarounds we're using

1. For **Trigger 1**: instead of registering a new per-frame system, we extended the existing scheduled-actions processor (which is the system that queues the screenshots and sends `AppExit`) to also write the override `Resource` itself. Keeping the system count constant sidesteps the hang.
2. For **Trigger 2**: scenarios mutate init-resourced types via a `Startup` system + `ResMut` instead of `insert_resource`.

Both workarounds are stable. The visual-QA scenarios run end-to-end in 4–6 seconds. We have ~3 scenarios passing in our integration test suite.

## What I haven't done yet

I haven't built a minimal standalone repro outside our codebase. The hang reproduces 100% of the time in our test suite when either trigger is present, and removing the trigger fixes it 100% of the time. Happy to build a minimal repro if it would help — wanted to file the report first since the symptom + isolation is very clean.

## Reference

- Repository: https://github.com/bryancostanich/lunie (private at time of filing — can share access on request)
- The two affected scenarios are at `source/lunie-game/src/scenarios/sun_phases.rs` and `source/lunie-game/src/scenarios/terrain_overview.rs`.
- The scenario harness lives at `source/lunie-game/src/scenarios/mod.rs`.
- The `process_scheduled_actions` system showing the workaround pattern is in the same file.

Happy to provide more context, build a minimal repro, or test patches on this hardware.

Contributor guide

Open the contributing guide

Research direction

Start with source/lunie-game/src/scenarios/mod.rs and the affected scenarios in sun_phases.rs and terrain_overview.rs; inspect process_scheduled_actions and reproduce the scenario suite with each trigger isolated. Compare the added per-frame system and resource replacement with the documented workarounds, then reduce the behavior to a minimal reproduction if possible. Done means AppExit::Success terminates cleanly without relying on either workaround.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
game-dev
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.