bytechefhq / bytechefhq/bytechef

[bug] Quartz JDBC job store is never configured as clustered (isClustered defaults to false)

Open
#5,650 0 comments 0 reactions 1 assignee Claimed by @ivicac View on GitHub
backend bug
Dominant language
Java
Stars
1k
Forks
170
Avg merge
11h 25m
Merged PRs (30d)
115

Description

> *This was generated by AI during triage.*

## Summary

Every app that uses the Quartz JDBC job store configures persistence but never enables
clustering. `org.quartz.jobStore.isClustered` and `org.quartz.scheduler.instanceId` are not
set anywhere in the repo, so Quartz falls back to its defaults (`isClustered=false`,
`instanceId=NON_CLUSTERED`).

A JDBC job store on its own gives **durability**, not **coordination**. With `isClustered=false`
each node assumes it is the sole owner of the job store: it does not take the `QRTZ_LOCKS` row
locks, does not check in to `QRTZ_SCHEDULER_STATE`, and does not honour other nodes' entries in
`QRTZ_FIRED_TRIGGERS`. Two instances pointed at the same database will both fire every trigger.

This is latent rather than active today, because the Helm chart ships a single replica
(`kubernetes/helm/bytechef/values.yaml:5`, `replicaCount: 1`). It becomes a live duplicate-execution
bug the moment anyone scales past one instance.

## Affected configuration

| File | Line |
| --- | --- |
| `server/apps/server-app/src/main/resources/config/application.yml` | 272-278 |
| `server/ee/apps/config-server-app/src/main/resources/config/apps/scheduler-app.yml` | 18-23 |
| `server/ee/apps/runtime-job-app/src/main/resources/config/application.yml` | 46-51 |

All three set `job-store-type: jdbc` and then set only `driverDelegateClass` under
`properties` — no clustering keys.

## Impact

Jobs registered through `QuartzTriggerScheduler` / `QuartzConnectionRefreshScheduler`
would each run once per node:

- **`ConnectionOAuth2TokenRefreshJob`** — the sharpest case. It calls
`connectionFacade.executeConnectionRefresh(connectionId)` and then reschedules itself via
`scheduler.rescheduleJob(...)`. Two nodes firing together means two concurrent OAuth2
refresh-token grants for the same connection; against providers that rotate refresh tokens
this can invalidate the credential and force the user to re-authorise. Both nodes then race
to reschedule the same trigger key.
- **`PollingTriggerJob`** / **`ScheduleTriggerJob`** — duplicate workflow executions, i.e.
duplicated user-visible side effects.
- **`DynamicWebhookTriggerRefreshJob`**, **`OneTimeSchedulerJob`** — duplicate refresh and
duplicate one-shot execution.

## Why the fix should be small

The Liquibase schema at
`server/libs/platform/platform-scheduler/platform-scheduler-impl/src/main/resources/config/liquibase/changelog/quartz/quartz_postgres_init.sql`
**already creates the cluster-critical tables** — `QRTZ_LOCKS`, `QRTZ_SCHEDULER_STATE` and
`QRTZ_FIRED_TRIGGERS` are all present. No migration is needed; this looks like a
configuration-only change plus a decision about which deployments should enable it.

## Suggested direction (needs a maintainer decision)

Roughly:

```yaml
spring:
quartz:
properties:
org.quartz.jobStore.isClustered: true
org.quartz.jobStore.clusterCheckinInterval: 20000
org.quartz.scheduler.instanceId: AUTO
```

`instanceId: AUTO` matters — clustered nodes must have distinct instance ids, so it cannot be
a shared fixed value.

Open questions for a maintainer:

1. Should CE's `server-app` enable clustering, or should the monolith instead be **documented
as single-instance** and the chart constrained to `replicaCount: 1`? Enabling clustering has
a cost: Quartz's cluster check-in is a DB polling loop and is sensitive to clock skew
between nodes.
2. Does EE's dedicated `scheduler-app` scale beyond one replica in practice? If yes it needs
the same change; if it is deliberately a singleton that should be written down.
3. Should there be a startup guard that fails fast when `job-store-type: jdbc` is combined with
`isClustered=false` on a multi-replica deployment, so this cannot silently regress?

Question 1 is a genuine design trade-off rather than a mechanical fix, which is why this is
filed for triage rather than as a ready-to-implement task.

## Reproducing

Not reproduced against a live multi-node deployment — this is a static configuration finding.
Verified at commit `880f8aa30a9`:

```bash
# no clustering configuration in any production config
grep -rn "isClustered\|scheduler.instanceId" --include="*.yml" --include="*.java" server/ \
| grep -v "/build/" | grep -v "/src/test/"

# but the cluster tables do exist
grep -oiE "qrtz_(locks|scheduler_state|fired_triggers)" \
server/libs/platform/platform-scheduler/platform-scheduler-impl/src/main/resources/config/liquibase/changelog/quartz/quartz_postgres_init.sql \
| sort -u
```

The first command returns nothing; the second returns all three table names.

The only occurrence of `isClustered` anywhere in the tree is in test scaffolding —
`QuartzIntTest` sets it to `false` explicitly for its single-node `SchedulerFactoryBean`
(`.../src/test/java/com/bytechef/platform/scheduler/QuartzIntTest.java:265`). That is correct
for that test and is not itself a problem; noted here only so the grep result isn't
mistaken for production configuration.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.