apache / apache/cloudstack

Enhance Disaster Recovery Scenario thru re-copy template/snapshot between zone

Abierto
#8,672 1 comentario 0 reacciones 0 asignados Ver en GitHub
component:storage type:enhancement
Lenguaje dominante
Java
Estrellas
3.1k
Forks
1.4k
Merge medio
6 d 19 h
PR fusionados (30 d)
32

Descripción

##### ISSUE TYPE

* Enhancement Request

##### COMPONENT NAME

~~~
Snapshot
Tempalte
~~~

##### CLOUDSTACK VERSION

~~~
4.19
~~~

##### CONFIGURATION

##### OS / ENVIRONMENT

##### SUMMARY

While ACS 4.19 brought a new feature that enable copy disk snapshot to another zone, an idea came up to extend this feature to become a disaster recovery approach.

Assume administrator ensure all template and disk snapshots are already made the copy to the partner zone,
when the victim zone's primary/secondary storage is unusable or corrupted, we will have an opportunity to just copy the template/snapshot from partner zone back to victim zone after the storage were rebuilt. However since the current ACS did not design to handle such scenario, so the VM originally host on victim zone has to deploy as a new instance on partner zone and start a new lifecycle.

My preliminary idea is

Scenario 1 - When victim zone primary storage is dead and unrecoverable.
1. Rebuilt a new primary storage
2. ACS found the victim zone instance volume are unavailable
3. We revert the volume from the snapshot image reside on secondary storage (Full Clone).

Scenario 2 - When victim zone secondary storage is dead and unrecoverable.
1. Rebuild a new secondary storage
2. Implement replace copy mechanism for template/snapshot from partner zone to victim zone
3. When we try to revert a snapshot, ACS found the victim zone snapshot disk is lost in victim zone, ACS then copy the snapshot from partner zone to the fresh secondary storage and do the disk recovery (Full Clone).

Scenario 3 - When victim zone both primary secondary storage is dead and unrecoverable.
1. Rebuild both new primary and secondary storage
2. Implement replace copy mechanism for template/snapshot from partner zone to victim zone
3. ACS found the victim zone instance volume are unavailable
5. We revert the volume from the snapshot image
6. ACS found the victim zone snapshot disk is lost in victim zone, ACS then copy the snapshot from partner zone to the fresh secondary storage and do the disk recovery (Full Clone).

With such implementation, when ACS setup multiple zone with scheduled disk snapshot, it will facilitate recovery scenario itself without engaging third party backup solution.

##### STEPS TO REPRODUCE

~~~

~~~

##### EXPECTED RESULTS

~~~

~~~

##### ACTUAL RESULTS

~~~

~~~

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

No se identifican archivos, pruebas ni puntos de entrada en el issue. Empieza por mapear los flujos de trabajo de snapshots, plantillas, almacenamiento y copia entre zonas de CloudStack; después, define el comportamiento de recuperación y las pruebas necesarias para los tres escenarios de disaster recovery.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
java
Área
cloud, infrastructure
Tipo de issue
Nueva funcionalidad
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Estancado
Claridad
Necesita aclaración
Aptitud para principiantes
25/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.