linkedin / linkedin/venice

[BUG] Prevent chunking config mismatch between parent and child data centers

Open
#659 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Java
Stars
611
Forks
124
Avg merge
3d 1h
Merged PRs (30d)
26

Description

### Venice version

0.4.139

### System information

- **OS Platform and Distribution (e.g., Linux Ubuntu 20.0)**: Mariner 5.15.111.1-1.cm2
- **JDK version**: 17

### Describe the problem

Issue encountered by user in prod: # A store has write compute enabled.

- The customer can't read the data in in DC1, but can read the same keys in other fabrics.

Root cause: # DC1 has chunking flag enabled, but not in parent and other child fabrics.

1. In the read path, Venice Server is looking at StoreVersionState to check whether chunking is enabled or not and the chunking flag in StoreVersionState is decided by the StartOfPush control message generated by VPJ.
2. Even DC1 has chunking enabled in Version metadata, but StoreVersionState doesn't have it, so the read path won't append the chunking suffix, so the lookup always fail.

Potential mitigation: # Always looking at StoreVersionState in the ingestion path.

- Prevent the update to Child Controllers (which aligns with the decision of ParentController SPoF project).

1 seems more robust, and it seems good to have 2 regardless. But open to other approaches.

Theres another very similar issue regarding partition count mismatch between the colos. So it's be good to fix that here as well.

### Tracking information

_No response_

### Code to reproduce bug

_No response_

### What component(s) does this bug affect?

- [X] `Controller`: This is the control-plane for Venice. Used to create/update/query stores and their metadata.
- [ ] `Router`: This is the stateless query-routing layer for serving read requests.
- [ ] `Server`: This is the component that persists all the store data.
- [ ] `VenicePushJob`: This is the component that pushes derived data from Hadoop to Venice backend.
- [ ] `VenicePulsarSink`: This is a Sink connector for Apache Pulsar that pushes data from Pulsar into Venice.
- [ ] `Thin Client`: This is a stateless client users use to query Venice Router for reading store data.
- [ ] `Fast Client`: This is a stateful client users use to query Venice Server for reading store data.
- [ ] `Da Vinci Client`: This is an embedded, stateful client that materializes store data locally.
- [ ] `Alpini`: This is the framework that fast-client and routers use to route requests to the storage nodes that have the data.
- [ ] `Samza`: This is the library users use to make nearline updates to store data.
- [ ] `Admin Tool`: This is the stand-alone client used for ad-hoc operations on Venice.
- [ ] `Scripts`: These are the various ops scripts in the repo.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by tracing StoreVersionState and the StartOfPush control message through the Venice Server read path and VPJ ingestion path, then inspect how parent and Child Controllers update version metadata. Define and verify a consistent source of truth for chunking and partition-count settings across parent and child data centers; the issue provides no named files or tests.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
database, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
28/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.