apache / apache/gravitino

[Subtask] Close the multi-node consistency gaps around the entity cache

Open
#12,377 0 comments 0 reactions 0 assignees View on GitHub
subtask
Dominant language
Java
Stars
3.2k
Forks
935
Avg merge
1d 16h
Merged PRs (30d)
298

Description

### Describe the subtask

Follow-up to #12376, which covers the `batchGet` write-back race and the reads that a fresher read can fix. This subtask covers what remains: the invalidation channel's own failure modes, the reads whose correctness a fresher read cannot restore, and the design document's contract.

**1. A missed invalidation never self-heals on hot keys.**

`CaffeineEntityCache` builds the cache with `expireAfterAccess`, so a continuously read key has its TTL reset on every read and never expires. This is item 3 of #11736, which was closed while PR #11739 addressed only listener retry. The consequence is that any lost invalidation becomes permanent staleness precisely on the hottest keys, and `Configs.CACHE_EXPIRATION_TIME` documents itself as "after it is written", which does not match the implementation.

**2. The poller can still lose a row (accepted limitation, worth re-evaluating).**

The cursor is `WHERE id > #{lastConsumedId} ORDER BY id` over an auto-increment id, so a transaction holding a lower id that commits after a higher id has its row skipped forever. This is item 1 of #11736 and was consciously accepted at the time. It should be re-evaluated now that the same feed backs the entity store cache for all cacheable types rather than only the catalog and jcasbin id caches: combined with (1), the result is unbounded staleness rather than a bounded window.

**3. Two check-then-act paths are racy across nodes even with the cache disabled.**

`PolicyManager.changePolicyEnabledState()` short-circuits when the current `enabled` state already matches, and `JobManager.cancelJob()` skips `jobExecutor.cancelJob()` on a terminal status. `TreeLockUtils` locks are in-process only, so nothing serializes these across nodes. A fresher read narrows the window but cannot close it, and the failure mode is silent: the operation is skipped and the API still reports success.

**4. `exists()` trusts the cache.**

`RelationalEntityStore.exists()` returns `true` on a cache hit without consulting the backend. A miss falls through to the store, so the error is one-directional: an entity dropped on another node is still reported as existing, surfacing as a spurious `AlreadyExistsException` on create paths. The short-circuit saves a single indexed query.

**5. The design document overstates what is safe.**

`design-docs/gravitino-entity-cache-multinode-design.md` says the cache holds only id/audit for tables and that a stale read is "at worst cosmetic". Neither holds for entity-store-backed (managed) catalogs: `ManagedTableOperations.loadTable()` returns the cached entity's columns, properties, partitioning and indexes as the final result, and `ManagedSchemaOperations` does the same for comment and properties.

**Proposed work**

- [ ] Add `expireAfterWrite` (alone or alongside `expireAfterAccess`) so a lost invalidation degrades to bounded staleness, and fix the config documentation.
- [ ] Re-evaluate the poller cursor so that a committed row is consumed exactly once by every node.
- [ ] Replace the check-then-act in `PolicyManager` and `JobManager` with conditional updates that report affected rows; keep the cached read as a fast path only.
- [ ] Drop the cache short-circuit in `RelationalEntityStore.exists()`.
- [ ] Rewrite the per-entity table in the design document along the managed / connector-backed split, state the staleness bound as one poll interval with the TTL as the fallback, and drop the "cosmetic" wording.
- [ ] Roadmap: read-your-writes via a change-log watermark returned by writes, so a lagging node catches up before serving a request that carries it.

Related: #12376, #11736, #11739, #12151.

### Parent issue

#11737

Contributor guide

Open the contributing guide

Research direction

Start by tracing CaffeineEntityCache and Configs.CACHE_EXPIRATION_TIME, then inspect the poller cursor and the PolicyManager, JobManager, and RelationalEntityStore.exists() entry points. Read the managed catalog and schema operations plus design-docs/gravitino-entity-cache-multinode-design.md. Done means the listed consistency gaps are addressed and the document states the actual staleness guarantees.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
backend, databases, distributed-systems
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.