graphprotocol / graphprotocol/graph-node

[Bug] v0.40.1 - "graphman reassign" operations fails ~ 50% of the time without any meaningful info in subgraph's logs

Open
#6,227 3 comments 3 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug Stale
Dominant language
Rust
Stars
3.2k
Forks
1.1k
Avg merge
4d 1h
Merged PRs (30d)
1

Description

Bug report

Hi Team,
I'm using grapnode v0.40.1. Roughly 50% of my "graphman reassign" operations look like below:

  1. Subgraph is running perfectly fine (healthy, no lag, not paused) on indexer node X
  2. I invoke "graphman reassign SUBGRAPH_HASH indexer_node_Y" command and it finalises without problems
  3. I check the status of the subgraph with "graphman info --status" command after few minutes and the subgraph is stuck, it isn't processing any new blocks
  4. Target index node doesn't produce logs for this subgraph
  5. Prometheus metrics (for example "deployment_head") are available on the previous index node and on the new/target indexer node, on the previous index node they have misleading values as processing doesn't really take place on it anymore
  6. I have to perform at least 1 more reassign operation to fix the subgraphs, sometimes up to 5 reassignment operations have to be done
    I checked existing issues and it may be the case that https://github.com/graphprotocol/graph-node/issues/5253 is related so I've added a comment to it.

Can someone please take a look?

Relevant log output

IPFS hash

No response

Subgraph name or link to explorer

No response

Some information to help us out
  • Tick this box if this bug is caused by a regression found in the latest release.
  • Tick this box if this bug is specific to the hosted service.
  • I have searched the issue tracker to make sure this issue is not a duplicate.
OS information

None

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the intermittent graphman reassign SUBGRAPH_HASH indexer_node_Y workflow and inspect graphman info --status afterward, comparing logs and deployment_head on both indexer nodes. Review related issue #5253 for context. Done means reassignment consistently starts processing on the target node, with accurate metrics and useful logs when it fails.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
cli, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.