Disagg EstablishDisaggTask keeps MEET_LOCK for ~2s after large txn already committed
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 1k
- Forks
- 423
- Avg merge
- 1d 15h
- Merged PRs (30d)
- 24
Description
Bug Report
Please answer these questions before submitting your issue. Thanks!
1. Minimal reproduce step (Required)
Environment:
- TiFlash next-gen disaggregated build:
ENABLE_NEXT_GEN=ON,ENABLE_NEXT_GEN_COLUMNAR=OFF - Compute node (CN) reads via
StorageDisaggregated::readThroughTiFlashWrite - Write node (WN) handles
EstablishDisaggTaskand performs learner read (wait index+ resolve locks) - Storage engine is next-gen / cloud-storage-engine style TiKV (has
txn_info!resolve_lock logs)
Observed production-like case (TiFlash logs in investigate.md, plus tidb.log / tikv.log):
- A large TiDB write enters 2PC (
[BIG_TXN],txnStartTS=468120244537524231,keys=43320,size≈5.6MB). - An MPP query starts on CN with
start_ts=468120244616167462. - CN sends
EstablishDisaggTaskto WN for table scan regions. - WN learner read hits a local Put lock on region
91537:lock_version=468120244537524231(= write txn start_ts)min_commit_ts=468120244603060470(min_commit_ts < start_ts)txn_size=2166,lock_ttl≈20126
- WN returns
error_locked; CN callslock_resolver->resolveLocks(...)and retries EstablishDisaggTask. - TiKV begins
resolve_lock committing transactionfor the samestart_tswith
commit_ts=468120244616167447(commit_ts < query start_ts, same physical ms). - TiFlash keeps retrying
MEET_LOCKfor ~2s (wait_cost=0ms, identical lock fields) until EstablishDisaggTask finally succeeds.
Relevant CN / WN code paths:
// StorageDisaggregatedRemote.cpp — CN resolve then always retry Establish
auto before_expired = cluster->lock_resolver->resolveLocks(
bo,
sender_target_mpp_task_id.gather_id.query_id.start_ts,
locks,
pushed);
// TODO: Use `pushed` to bypass large txn.
throw Exception(error_msg, ErrorCodes::DISAGG_ESTABLISH_RETRYABLE_ERROR);
// LearnerReadWorker.cpp — WN disagg always throws LockException to CN
if (is_wn_disagg_read && !region_locks.empty())
throw LockException(std::move(region_locks));
// DAGStorageInterpreter.cpp — no local bypass retry for disagg task
catch (const LockException & e)
{
if (context.getDAGContext()->is_disaggregated_task)
throw; // let CN retry
...
}
2. What did you expect to see? (Required)
Once CN has resolved the lock status and TiKV has determined the write txn is already committed with commit_ts < query start_ts (i.e. the committed data is visible to this snapshot), the disagg read path should not keep failing EstablishDisaggTask on the same stale local LockCF entry for seconds.
Expected behavior options:
- After CN
resolveLocksconfirms the txn is committed / cleanup is in progress, the next EstablishDisaggTask (or WN local retry) should wait for a fresh read index that covers the resolve/commit raft logs, so WN no longer sees the old LockCF entry. - Or, once the lock's final status is known, propagate enough information (
bypass_lock_ts/ resolved txn id, or equivalent) so WN can skip the already-handled lock while secondary resolve is still applying. - Avoid a long CN↔WN EstablishDisaggTask retry loop with
wait_cost=0msthat re-observes an unchanged local lock.
In short: after “txn already committed + resolve_lock committing secondary”, disagg read should converge on learner apply / correct bypass, not spin on the same MEET_LOCK.
3. What did you see instead (Required)
Timeline (UTC, 2026-08-03)
| time | component | event |
|---|---|---|
| 06:03:37.095 | TiDB / client-go | [BIG_TXN] start 2PC: txnStartTS=468120244537524231, keys=43320, puts=43320, size=5678162 |
| ~06:03:37.121 | Query | MPP start_ts=468120244616167462 |
| 06:03:37.153 | TiFlash WN | Learner read MEET_LOCK on region 91537; same lock fields as below |
| 06:03:37.154+ | TiFlash CN | error_locked → resolveLocks → retry EstablishDisaggTask |
| 06:03:37.199 ~ 06:03:37.220 | TiKV (CSE) | 512× resolve_lock committing transaction for start_ts=468120244537524231, commit_ts=468120244616167447, request_source=external_Select (and empty) |
| 06:03:37.221 ~ 06:03:38.736 | TiFlash WN/CN | Same lock repeatedly returned; every learner read shows wait_cost=0ms and identical lock info |
| 06:03:39.307 | TiFlash CN | EstablishDisaggTask finally succeeds (error=false) |
Evidence from tidb.log / tikv.log
tidb.log (client-go txnkv/transaction/2pc.go, threshold keys>10000 or size>4MB):
[BIG_TXN] keys=43320 puts=43320 size=5678162 txnStartTS=468120244537524231
→ Confirms the blocking lock belongs to a very large 2PC write (tens of thousands of secondary locks).
tikv.log (next-gen / cloud-storage-engine resolve_lock.rs via txn_info!):
resolve_lock committing transaction
start_ts=468120244537524231
commit_ts=468120244616167447
min_commit_ts=468120244603060470
use_async_commit=false
lock_type=Put ttl=20126 for_update_ts=468120244537524231
request_source=external_Select # Select-path lock resolver (matches CN resolveLocks)
Important ts relation (same physical ms):
- write
commit_tslogical = 23 - query
start_tslogical = 38 - ⇒
commit_ts < start_ts⇒ committed data is visible to this query
So this is not primarily “alive large txn waiting for TTL / pushed bypass”.
It is: txn already committed; secondary locks are being committed by resolve_lock; TiFlash learner still observes the old LockCF and keeps retrying.
Why MVCC still reports the lock on WN (correct locally)
On WN LockCF check (DecodedLockCFValue::getLockInfoPtr):
lock_version < start_tsmin_commit_ts < start_ts- lock type is Put (not skippable PessimisticLock/SharedLock)
Until the resolve/commit raft log is applied on the TiFlash learner (or the lock is explicitly bypassed), WN must treat it as blocking.
Why the retry is slow / inefficient (problem)
- CN
resolveLockstriggers TiKV secondary commit (request_source=external_Select), but EstablishDisaggTask retries do not effectively wait for learner to catch the new resolve/commit index (wait_cost=0msrepeatedly). - Disagg WN path always rethrows
LockExceptionto CN; no local bypass / short-wait retry like non-disagg MPP/batch-cop (#10992 does not apply tois_disaggregated_task). - Large txn (
43320keys) means secondary resolve/apply can take noticeable time; during that window CN↔WN spins on the same local lock for ~2s.
Notes on related PRs:
- #10986: invalidate stale read-index cache after local lock is found. Relevant to “resolve already committing on TiKV, but learner reused old read index and did not wait for resolve logs”. Worth verifying whether this branch includes it; it addresses part of this failure mode.
- #10992: local bypass-lock retry for MPP/batch-cop. Explicitly rethrows on
is_disaggregated_task, so this CN↔WN EstablishDisaggTask loop is unchanged. - CN
TODO: Use pushed to bypass large txnhelps the “txn still alive / not expired” case; for this incident (already committed,commit_ts < start_ts), the missing piece is mainly wait for resolve apply / avoid stale read-index reuse, not pushed TTL wait.
4. What is your TiFlash version? (Required)
tidb: v8.5.4-nextgen.202510.24
tikv: v8.5.4-nextgen.202510.31
tiflash: v8.5.4-nextgen.202510.13
ENABLE_NEXT_GEN=ONENABLE_NEXT_GEN_COLUMNAR=OFF(reads usereadThroughTiFlashWrite, not columnar path)
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Trace StorageDisaggregatedRemote.cpp, LearnerReadWorker.cpp, and DAGStorageInterpreter.cpp through the EstablishDisaggTask retry path. Compare the behavior with #10986 and #10992, focusing on read-index invalidation and the disaggregated-task exception path. Done means a committed transaction visible to the query no longer causes repeated retries against the same stale LockCF entry.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- databases, distributed-systems
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100