[Question][GitHub] PR records remain duplicated and stale after a repository is moved between GitHub connections
- Dominant language
- Go
- Stars
- 3.1k
- Forks
- 808
- Avg merge
- 1d 8h
- Merged PRs (30d)
- 49
Description
## Question
We are investigating how Apache DevLake handles repositories that are moved from one GitHub connection to another.
Our DevLake instance contains the same GitHub repository and pull request under multiple connection-scoped identities. For example, GitHub PR `#137` from an example repo
is present in the `pull_requests` table as follows:
```text
github:GithubPullRequest:1:3375195798 OPEN
github:GithubPullRequest:35:3375195798 MERGED
github:GithubPullRequest:36:3375195798 MERGED
github:GithubPullRequest:37:3375195798 MERGED
```
GitHub confirms that PR `#137` was merged:
```text
created_at: 2026-03-09T20:00:14Z
updated_at: 2026-03-10T03:38:33Z
closed_at: 2026-03-10T03:38:31Z
merged_at: 2026-03-10T03:38:31Z
```
The most recent pipeline for the project collected the repository through connection `35`:
```json
{
"plugin": "github_graphql",
"options": {
"connectionId": 35,
"fullName": "",
"githubId":
}
}
```
That pipeline correctly created or updated the connection-`35` record to `MERGED`. However, the older connection-`1` record remains `OPEN` and is not updated by subsequent pipeline runs.
Our reporting queries join pull requests to repositories by repository name and therefore include both the active connection-`35` record and the stale connection-`1` record. This results in incorrect open-PR counts unless we explicitly filter by the currently active connection-specific repository ID.
Is this behavior expected by design, or should DevLake reconcile records across GitHub connections when the same GitHub repository and pull request are collected through a different connection?
In particular:
1. Is the connection ID intentionally part of the identity of a GitHub repository and pull request?
2. When a repository is removed from connection `1` and added to connection `35`, should historical records under connection `1` remain permanently queryable?
3. Should a future full-sync or backfill pipeline update the old connection-`1` record, even though connection `1` is no longer present in the project blueprint?
4. Is there a supported procedure for migrating repository data between connections without creating duplicate tool-layer and domain-layer entities?
5. Is there a recommended way to mark or remove obsolete connection-scoped records?
6. What is the recommended query or dashboard strategy for counting PRs and identifying stale open PRs when the same repository exists under multiple connection IDs?
7. Is there a supported cleanup or deduplication task that preserves related PR comments, reviews, commits, deployments, and DORA metrics?
**Additional context**
The project uses the GitHub GraphQL plugin and has `skipOnFail` enabled. The affected pipeline also experienced a separate network timeout while preparing a GitHub GraphQL task:
```text
unable to get github API client instance
Failed to connect
dial tcp 20.26.156.210:443: i/o timeout
```
However, for PR `#137`, the active connection-`35` record was successfully updated to `MERGED`. The issue appears specifically related to the older connection-`1` record remaining stale and being included in cross-connection reporting.
The project’s sync policy is currently:
```text
Data time range: 2026-05-07 01:00 to Now
Sync frequency: Custom
Skip failed tasks: Enabled
```
We would appreciate clarification on whether this is:
- expected connection-scoped identity behavior,
- a limitation of connection migration,
- a data cleanup/configuration issue,
- or a defect in the GitHub GraphQL collector, extractor, converter, or project mapping logic.
**Expected outcome**
We would like to understand the supported DevLake behavior and obtain guidance for:
- migrating repositories between GitHub connections,
- preventing duplicate records,
- excluding obsolete connection data from project metrics,
- and safely cleaning up stale records without corrupting related DevLake data.
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue identifies the pull_requests table, the github_graphql plugin, project blueprints, and reporting queries. Start by tracing how connection-scoped repository and pull request identities are collected and mapped. Done means documenting whether reconciliation and cleanup are supported, along with a safe migration or reporting procedure.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- github, graphql
- Domain
- analytics, data-engineering, databases
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100