injectHeartbeat writing failed on teardown
- Vorherrschende Sprache
- Go
- Sterne
- 13.6k
- Forks
- 1.4k
- Ø Merge
- 2 Std. 31 Min.
- Gemergte PRs (30 T.)
- 4
Beschreibung
Hi all!
Environment:
Azure Mysql Flexible Sever, 8.0.44-azure
gh-ost 1.1.8
We have an instance with a number of partitioned tables and a large schema metadata. With this setup, the dry run check fails with the following errors (from the first to the last):
```code
Closed streamer connection. err=
Dropping table `_ghc`
Table dropped
Error 1146 (42S02): Table '_ghc' doesn't exist
...
Error 1146 (42S02): Table '_ghc' doesn't exist
...
injectHeartbeat writing failed 61 times, last error: Error 1146 (42S02): Table '_ghc' doesn't exist
```
It appears that on slow disks/thousands of tables, [finalCleanup](https://github.com/github/gh-ost/blob/v1.1.8/go/logic/migrator.go#L1723) is executed too early and drops the table before [teardown](https://github.com/github/gh-ost/blob/v1.1.8/go/logic/migrator.go#L373) (where injectHeartbeat actually stops). And deleting a '_ghc' table takes longer than `default-retries` (60) * `heartbeat-interval-millis` (100 ms).
Currently, this can be fixed by increasing `default-retries` or `heartbeat-interval-millis` (or both), but this seems like a workaround.
Is it possible to fix the order of execution?
Thanks in advance!
Beitragsleitfaden
Rechercherichtung
Beginne in go/logic/migrator.go bei finalCleanup und teardown und verfolge dann, wo injectHeartbeat während des dry-run-Ablaufs gestoppt wird. Reproduziere dies mit einer langsamen Bereinigung oder einem großen Schema und verifiziere, dass die Heartbeat-Schreibvorgänge beendet werden, bevor die Tabelle _ghc gelöscht wird, ohne die gemeldeten Teardown-Fehler.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- go, mysql
- Bereich
- database
- Issue-Typ
- Bug
- Schwierigkeit
- 3/5
- Geschätzter Aufwand
- 1-2 Tage
- Aktivitätsstatus
- Ruhig
- Klarheit
- Klar beschrieben
- Anfängerfreundlichkeit
- 58/100