apache / apache/cloudstack

Storage bandwidth is wasted during template uploads/imports

Abierto
#5,697 6 comentarios 0 reacciones 0 asignados Ver en GitHub
component:secondary-storage component:templates no-issue-activity status:stale type:improvement
Lenguaje dominante
Java
Estrellas
3.1k
Forks
1.4k
Merge medio
6 d 19 h
PR fusionados (30 d)
32

Descripción

At the moment when a template is being imported (via url or upload) there is a phase called "Installing template" or something similar. From what I noticed this phase calculates a hash and saves it to the database by reading the downloaded file over the storage network.
During this phase the template is not usable and I believe that should not be the case. I understand that the hash has a purpose and I am not saying it should be removed but I believe this should be an optional task that should be done in background.

In the following examples I do not consider disk speeds, just network speed/bandwidth.
For normal templates this is not necessarily noticeable. Consider this scenario (best case):
Template size: 4GB
Ingress bandwidth: 100 mb/s
Storage bandwidth: 1 gb/s
Download time required: 5.45 seconds
Installing template time required: ~0.6 seconds

Most templates (not ISOs) however are considerably larger, some could be even up to 500 GB (I do have a few templates that I have to import with very large sizes).
In such a case, installing template time required would be about 75 seconds (best case).

During this time (installing template time) the whole bandwidth available for the storage network (if it even is on a separate NIC) would be used up by this process resulting in bad performance for the cluster.

Ways to fix this would be:
1. either compute the hash as the transfer is happening.
2. make it optional (maybe even opt-in) and do it in background only (maybe even limit the bandwidth used for this)

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

Empieza rastreando la ruta de importación de la URL/carga del template alrededor de la fase «Installing template» y, después, identifica dónde se vuelve a leer desde el almacenamiento el archivo descargado y dónde se guarda su hash. Compara los pasos de transferencia y lectura desde el almacenamiento. Se considera terminado cuando la disponibilidad del template no está bloqueada por una lectura completa independiente del almacenamiento, el hash requerido sigue siendo correcto y el comportamiento propuesto en segundo plano o de ancho de banda está definido.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
java
Área
cloud, infrastructure, performance
Tipo de issue
Nueva funcionalidad
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Tranquilo
Claridad
Bastante claro
Aptitud para principiantes
45/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.