agronholm / agronholm/apscheduler
AsyncScheduler crashes (KeyError) when a job's lease expires and the job is re-acquired/run twice (MemoryDataStore)
- Langage dominant
- Python
- Étoiles
- 7.6k
- Forks
- 783
- Merge moyen
- 4 j 8 h
- PR mergées (30 j)
- 5
Description
### Summary
With the in-memory data store, if a running job's lease expires before it
finishes, the scheduler re-acquires the same job and runs it a **second time**.
The duplicate run then crashes the scheduler on two unguarded operations:
- `MemoryDataStore.release_job`: `self._jobs_by_id.pop(result.job_id)` → `KeyError`
- `AsyncScheduler._run_job`: `self._running_jobs.remove(job)` → `KeyError`
The exception escapes the scheduler's TaskGroup, so `run_until_stopped()`
raises and the scheduler stops. In a long-lived service this is fatal.
Lease expiry of a still-running job is reachable in normal operation: the
lease-renewal coroutine (`extend_job_leases`, every `lease_duration / 2`) is
just another task on the event loop, so if the loop is saturated — e.g. a
burst of jobs that do blocking work inline at startup — the renewal is starved
and a concurrently-suspended job loses its lease.
### Reproduction (deterministic, ~5s)
```python
import asyncio, time
from datetime import datetime, timezone
from apscheduler import AsyncScheduler
from apscheduler.triggers.date import DateTrigger
async def suspended_job(): # awaits across the lease window
await asyncio.sleep(6)
async def loop_blocker(): # starves the lease-renewal coroutine
time.sleep(6)
async def main():
# short lease for a fast demo; in the wild it's a normal lease whose
# renewal is starved by a saturated event loop
async with AsyncScheduler(lease_duration=2, max_concurrent_jobs=5) as sched:
now = datetime.now(timezone.utc)
await sched.add_schedule(suspended_job, DateTrigger(run_time=now), id="job")
await sched.add_schedule(loop_blocker, DateTrigger(run_time=now), id="blk")
await sched.run_until_stopped()
asyncio.run(main())
```
### Expected
A job whose lease expired while it is still running should not be re-acquired
and run a second time; at minimum `release_job` / `_running_jobs.remove` should
tolerate a job that is already gone, instead of raising `KeyError` out of the
scheduler.
### Actual
```
Scheduler crashed
...
File ".../apscheduler/datastores/memory.py", line 300, in release_job
job = self._jobs_by_id.pop(result.job_id)
KeyError: UUID('...')
...
File ".../apscheduler/_schedulers/async_.py", line 1238, in _run_job
self._running_jobs.remove(job)
KeyError: Job(id=UUID('...'), task_id='__main__:suspended_job', ...)
```
### Notes
- The SQL and MongoDB stores release via an idempotent `DELETE`, so they don't
crash on a redundant release; only `MemoryDataStore`'s `dict.pop` does. The
`_running_jobs.remove` in `_run_job` is store-independent, though.
- Whether the *real* bug is "a still-running job should never be re-acquired"
(the lease/renewal robustness) or "release must be idempotent" (defensive),
the two unguarded `KeyError`s turn it into a hard scheduler crash either way.
### Environment
- apscheduler 4.0.0a6
- Python 3.14
- `MemoryDataStore` (the default for `AsyncScheduler()`)
Guide de contribution
Ouvrir le guide de contribution
Évaluation
Cette issue n'a pas encore été évaluée.