Enhance Disaster Recovery Scenario thru re-copy template/snapshot between zone
- Lenguaje dominante
- Java
- Estrellas
- 3.1k
- Forks
- 1.4k
- Merge medio
- 6 d 19 h
- PR fusionados (30 d)
- 32
Descripción
##### ISSUE TYPE
* Enhancement Request
##### COMPONENT NAME
~~~
Snapshot
Tempalte
~~~
##### CLOUDSTACK VERSION
~~~
4.19
~~~
##### CONFIGURATION
##### OS / ENVIRONMENT
##### SUMMARY
While ACS 4.19 brought a new feature that enable copy disk snapshot to another zone, an idea came up to extend this feature to become a disaster recovery approach.
Assume administrator ensure all template and disk snapshots are already made the copy to the partner zone,
when the victim zone's primary/secondary storage is unusable or corrupted, we will have an opportunity to just copy the template/snapshot from partner zone back to victim zone after the storage were rebuilt. However since the current ACS did not design to handle such scenario, so the VM originally host on victim zone has to deploy as a new instance on partner zone and start a new lifecycle.
My preliminary idea is
Scenario 1 - When victim zone primary storage is dead and unrecoverable.
1. Rebuilt a new primary storage
2. ACS found the victim zone instance volume are unavailable
3. We revert the volume from the snapshot image reside on secondary storage (Full Clone).
Scenario 2 - When victim zone secondary storage is dead and unrecoverable.
1. Rebuild a new secondary storage
2. Implement replace copy mechanism for template/snapshot from partner zone to victim zone
3. When we try to revert a snapshot, ACS found the victim zone snapshot disk is lost in victim zone, ACS then copy the snapshot from partner zone to the fresh secondary storage and do the disk recovery (Full Clone).
Scenario 3 - When victim zone both primary secondary storage is dead and unrecoverable.
1. Rebuild both new primary and secondary storage
2. Implement replace copy mechanism for template/snapshot from partner zone to victim zone
3. ACS found the victim zone instance volume are unavailable
5. We revert the volume from the snapshot image
6. ACS found the victim zone snapshot disk is lost in victim zone, ACS then copy the snapshot from partner zone to the fresh secondary storage and do the disk recovery (Full Clone).
With such implementation, when ACS setup multiple zone with scheduled disk snapshot, it will facilitate recovery scenario itself without engaging third party backup solution.
##### STEPS TO REPRODUCE
~~~
~~~
##### EXPECTED RESULTS
~~~
~~~
##### ACTUAL RESULTS
~~~
~~~
Guía de contribución
Línea de trabajo
No se identifican archivos, pruebas ni puntos de entrada en el issue. Empieza por mapear los flujos de trabajo de snapshots, plantillas, almacenamiento y copia entre zonas de CloudStack; después, define el comportamiento de recuperación y las pruebas necesarias para los tres escenarios de disaster recovery.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- java
- Área
- cloud, infrastructure
- Tipo de issue
- Nueva funcionalidad
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Estancado
- Claridad
- Necesita aclaración
- Aptitud para principiantes
- 25/100