99designs / 99designs/cmdstalk

Jobs that timeout will never be able to run again

Open
#2 6 comments 0 reactions 0 assignees View on GitHub
Dominant language
Go
Stars
76
Forks
14
PR merge metrics
No merged PRs in 30d

Description

When a job overruns it's TTR, beanstalkd will increment the job's timeout stat and put it back on the work queue for another worker to reserve.

In an effort to prevent pathological jobs from dog-piling all available workers, `cmdstalk` [will bury a task it reserves that has timeouts greater than 1](https://github.com/99designs/cmdstalk/blob/master/broker/broker.go#L112). This means that once a task is buried because of a timeout, it will _always_ re-bury instantly each time it is kicked: the job becomes un-runnable.

Using just the `buried`, `kicked` and `timeout` counters, there does not appear to be a way to differentiate between "kicks due buries due to timeouts" in the way that would allow `cmdstalk` to bury a job the next time it is reserved after a timeout.

The `beanstalkd` protocol docs make mention of [a one second grace period](https://github.com/kr/beanstalkd/blob/master/doc/protocol.txt#L224) at the end of a reserve time - would it be possible to use this grace period to bury a timed out job in the "same run" as the timeout occurred?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.