Scheduled recipe runs that fail mid-stream are recorded and reported as successful
- Ngôn ngữ chính
- Rust
- Star
- 54.2k
- Fork
- 6.2k
- Merge trung bình
- 3 ngày 4 giờ
- Pull request đã merge (30 ngày)
- 240
Mô tả
**Describe the bug**
When a scheduled recipe run fails partway through — provider error, rate limit, network drop, anything that surfaces as an `Err` on the agent event stream — the scheduler records and reports it as a successful completion. There is no failure signal anywhere the user can see: the CLI prints success, the job list shows a normal completed run, and the `scheduler_job_failed` telemetry never fires.
`execute_job` breaks out of the stream loop on error and sets a local `stream_error` flag, but that flag is only ever used to pick a label for a tracing counter. The function then returns `Ok(session.id)` on both paths.
`crates/goose/src/scheduler.rs:1123-1141` — the error is captured:
```rust
let mut stream_error = false;
while let Some(message_result) = stream.next().await {
tokio::task::yield_now().await;
match message_result {
Ok(AgentEvent::Message(msg)) => conversation.push(msg),
Ok(AgentEvent::HistoryReplaced(updated)) => conversation = updated,
Ok(_) => {}
Err(e) => {
tracing::error!("Error in agent stream: {}", e);
stream_error = true;
break;
}
}
}
```
`crates/goose/src/scheduler.rs:1154` — its only consumer is a metrics label:
```rust
let exit_type = if stream_error { "error" } else { "normal" };
```
`crates/goose/src/scheduler.rs:1214` — and the function returns success regardless:
```rust
Ok(session.id)
```
Two things follow from that unconditional `Ok`:
1. The cron completion handler at `scheduler.rs:392-396` takes the `Ok(_)` arm, logs `"Job '{id}' completed"`, and skips the `scheduler_job_failed` telemetry emission in the `Err` arm.
2. The `#[cfg(feature = "telemetry")]` block at `scheduler.rs:1191-1210` unconditionally emits `schedule_job_completed` with `"status": "completed"` before returning, so the failed run is affirmatively reported as a success rather than merely going unreported.
`run_now` propagates the same `Ok`, so `goose schedule run-now` prints its success message (`crates/goose-cli/src/commands/schedule.rs:252`) for a run that produced nothing.
---
**To Reproduce**
1. Create a schedule for any recipe.
2. Make the provider fail mid-turn — easiest is to point the configured provider at an unreachable base URL, or use a provider/key combination that returns a rate-limit or auth error after the stream has opened.
3. Trigger the job: `goose schedule run-now --id `, or wait for the cron activation.
4. Observe the CLI reports the run was triggered successfully and returns a session id.
5. `goose schedule sessions --id ` lists the run alongside genuine successes with nothing marking it as failed.
6. The only trace of the failure is an `ERROR Error in agent stream: ...` line in the logs, which nobody is watching for a background job.
---
**Expected behavior**
A run whose event stream terminated on an error should be reported as a failed run:
- `execute_job` should return `Err` (or otherwise propagate a failed outcome) when the stream ended on an error, so the cron handler's existing `Err` arm — error log plus `scheduler_job_failed` telemetry — actually runs.
- The completion telemetry should report a failed status for that run rather than `"completed"`.
- `run_now` should surface the failure to the caller so the CLI does not print a success message.
- The failure should be discoverable after the fact from the schedule's session listing, not only from log scraping.
The partial session should still be persisted — the value of a failed run is being able to inspect how far it got — so this is about the reported outcome, not about discarding work.
---
**Screenshots**
N/A.
---
**Please provide the following information**
- **OS & Arch:** macOS 15 arm64 (code path is platform-independent)
- **Interface:** CLI (and any surface that reads schedule run state)
- **Version:** `main` @ dafdbb736
- **Extensions enabled:** N/A — reproduces with any recipe
- **Provider & Model:** any; the trigger is a mid-stream provider error
---
**Additional context**
This is the same shape as #11047, where a failed ACP turn is reported as a successful `session/prompt` with `stopReason: end_turn`. Scheduled jobs make it worse in practice because there is no human in the loop: a nightly job that has been failing on every activation for a week looks healthy in every surface goose exposes.
While reading this code I also noticed that the cron activation path has no overlap guard — the closure at `scheduler.rs:336` checks only `job.paused`, whereas `run_now` at `scheduler.rs:788` refuses to start with `"Job '{id}' is already running"` when `currently_running` is set. A run that outlasts its own interval therefore runs concurrently with itself, and because both runs write the same key in `running_tasks` (`scheduler.rs:366`, `scheduler.rs:379`) the second run's cancellation token is dropped, which makes `kill_running_job` report success without cancelling anything. That is a distinct bug with a distinct fix; I have kept it out of scope here and am happy to file it separately if it is wanted.
I am not starting implementation — filing this for triage per the contribution workflow.
Hướng dẫn đóng góp
Hướng nghiên cứu
Lỗi nằm trong crates/goose/src/scheduler.rs, ở hàm execute_job. Trước tiên, hãy kiểm tra vòng lặp stream và việc trả về Ok vô điều kiện. Xem handler hoàn tất cron và khối phát telemetry. Bản sửa lỗi bao gồm thay đổi kiểu trả về để truyền tiếp lỗi, cập nhật nhãn trạng thái telemetry và đảm bảo lệnh CLI run_now hiển thị lỗi. Chạy các bài kiểm thử hiện có cho tương tác giữa scheduler và agent để xác minh thay đổi.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- rust
- Lĩnh vực
- backend
- Loại issue
- Lỗi
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức độ hoạt động
- Sôi nổi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức phù hợp với người mới
- 65/100