canonical / canonical/postgresql-k8s-operator

Cannot restore PITR if the new cluster has a different app name from the one the backup was performed from

Open
#604 1 comment 0 reactions 0 assignees View on GitHub
bug
Dominant language
Python
Stars
15
Forks
33
Avg merge
1d 14h
Merged PRs (30d)
37

Description

## Steps to reproduce

1. Deploy the PostgreSQL K8s charm with the application name equal to `postgresql-k8s`.
2. Deploy the S3 integrator charm, configure it and relate it to the `postgresql-k8s` application.
3. Create a backup, then write some data and later create another backup.
5. Remove the `postgresql-k8s` application.
6. Deploy a new PostgreSQL charm with the application name equal to `db`.
7. Relate it to the S3 integrator application and trigger a PITR with `juju run db/leader restore restore-to-time="latest" --wait=1000s`.

## Expected behavior
Restore is successfully completed, and the unit has a status equal to `Move restored cluster to another S3 bucket`.

## Actual behavior

Restore fails, and the unit has a status equal to `cannot restore PITR, juju debug-log for details`.

## Versions

Operating system: Ubuntu 24.04 LTS

Juju CLI: 3.4.5-genericlinux-amd64

Juju agent: 3.4.5

Charm revision: 337

microk8s: v1.29.5 revision 6884

## Log output

Juju debug log:

```sh
unit-db-0: 23:34:35 ERROR unit.db/0.juju-log Restore failed: database service failed to reach point-in-time-recovery target. You can launch another restore with different parameters
unit-db-0: 23:34:35 ERROR unit.db/0.juju-log Can't tell last completed transaction time
```

## Additional context

The issue happens because the stanza name contains the name `postgresql` (from the previous PostgreSQL charm deployment), and the new application is using its application name (`db`) to search for a stanza from which it can perform the PITR.

Copied from VM issue: https://github.com/canonical/postgresql-operator/issues/562.

Contributor guide

Open the contributing guide

Research direction

Reproduce the failure with `juju run db/leader restore restore-to-time="latest" --wait=1000s` after changing the application name, then inspect the restore path and the reported debug-log errors about the stanza and last completed transaction time. Done means PITR succeeds for the new application name and the unit reaches `Move restored cluster to another S3 bucket`.

Written by the indexing model from the issue text.

Assessment

Tech stack
kubernetes, postgresql, python
Domain
databases, infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.