apache / apache/devlake

[Question][GitHub] PR records remain duplicated and stale after a repository is moved between GitHub connections

Offen
#9,040 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
type/question
Vorherrschende Sprache
Go
Sterne
3.1k
Forks
812
Ø Merge
1 T. 18 Std.
Gemergte PRs (30 T.)
50

Beschreibung

## Question
We are investigating how Apache DevLake handles repositories that are moved from one GitHub connection to another.

Our DevLake instance contains the same GitHub repository and pull request under multiple connection-scoped identities. For example, GitHub PR `#137` from an example repo

is present in the `pull_requests` table as follows:

```text
github:GithubPullRequest:1:3375195798 OPEN
github:GithubPullRequest:35:3375195798 MERGED
github:GithubPullRequest:36:3375195798 MERGED
github:GithubPullRequest:37:3375195798 MERGED
```

GitHub confirms that PR `#137` was merged:

```text
created_at: 2026-03-09T20:00:14Z
updated_at: 2026-03-10T03:38:33Z
closed_at: 2026-03-10T03:38:31Z
merged_at: 2026-03-10T03:38:31Z
```

The most recent pipeline for the project collected the repository through connection `35`:

```json
{
"plugin": "github_graphql",
"options": {
"connectionId": 35,
"fullName": "",
"githubId":
}
}
```

That pipeline correctly created or updated the connection-`35` record to `MERGED`. However, the older connection-`1` record remains `OPEN` and is not updated by subsequent pipeline runs.

Our reporting queries join pull requests to repositories by repository name and therefore include both the active connection-`35` record and the stale connection-`1` record. This results in incorrect open-PR counts unless we explicitly filter by the currently active connection-specific repository ID.

Is this behavior expected by design, or should DevLake reconcile records across GitHub connections when the same GitHub repository and pull request are collected through a different connection?

In particular:

1. Is the connection ID intentionally part of the identity of a GitHub repository and pull request?
2. When a repository is removed from connection `1` and added to connection `35`, should historical records under connection `1` remain permanently queryable?
3. Should a future full-sync or backfill pipeline update the old connection-`1` record, even though connection `1` is no longer present in the project blueprint?
4. Is there a supported procedure for migrating repository data between connections without creating duplicate tool-layer and domain-layer entities?
5. Is there a recommended way to mark or remove obsolete connection-scoped records?
6. What is the recommended query or dashboard strategy for counting PRs and identifying stale open PRs when the same repository exists under multiple connection IDs?
7. Is there a supported cleanup or deduplication task that preserves related PR comments, reviews, commits, deployments, and DORA metrics?

**Additional context**

The project uses the GitHub GraphQL plugin and has `skipOnFail` enabled. The affected pipeline also experienced a separate network timeout while preparing a GitHub GraphQL task:

```text
unable to get github API client instance
Failed to connect
dial tcp 20.26.156.210:443: i/o timeout
```

However, for PR `#137`, the active connection-`35` record was successfully updated to `MERGED`. The issue appears specifically related to the older connection-`1` record remaining stale and being included in cross-connection reporting.

The project’s sync policy is currently:

```text
Data time range: 2026-05-07 01:00 to Now
Sync frequency: Custom
Skip failed tasks: Enabled
```

We would appreciate clarification on whether this is:

- expected connection-scoped identity behavior,
- a limitation of connection migration,
- a data cleanup/configuration issue,
- or a defect in the GitHub GraphQL collector, extractor, converter, or project mapping logic.

**Expected outcome**

We would like to understand the supported DevLake behavior and obtain guidance for:

- migrating repositories between GitHub connections,
- preventing duplicate records,
- excluding obsolete connection data from project metrics,
- and safely cleaning up stale records without corrupting related DevLake data.

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Das Issue identifiziert die Tabelle pull_requests, das Plugin github_graphql, project blueprints und reporting queries. Beginne damit nachzuverfolgen, wie verbindungsspezifische Repository- und Pull-Request-Identitäten erfasst und zugeordnet werden. Als erledigt gilt, wenn dokumentiert ist, ob Reconciliation und Cleanup unterstützt werden, einschließlich eines sicheren Migrations- oder Reporting-Verfahrens.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
github, graphql
Bereich
analytics, data-engineering, databases
Issue-Typ
Bug
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Ruhig
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.