apache / apache/cloudstack

Storage bandwidth is wasted during template uploads/imports

Aperta
#5,697 6 commenti 0 reazioni 0 assegnatari Vedi su GitHub
component:secondary-storage component:templates no-issue-activity status:stale type:improvement
Lingua principale
Java
Stelle
3.1k
Fork
1.4k
Merge medio
6g 19h
PR unite (30g)
32

Descrizione

At the moment when a template is being imported (via url or upload) there is a phase called "Installing template" or something similar. From what I noticed this phase calculates a hash and saves it to the database by reading the downloaded file over the storage network.
During this phase the template is not usable and I believe that should not be the case. I understand that the hash has a purpose and I am not saying it should be removed but I believe this should be an optional task that should be done in background.

In the following examples I do not consider disk speeds, just network speed/bandwidth.
For normal templates this is not necessarily noticeable. Consider this scenario (best case):
Template size: 4GB
Ingress bandwidth: 100 mb/s
Storage bandwidth: 1 gb/s
Download time required: 5.45 seconds
Installing template time required: ~0.6 seconds

Most templates (not ISOs) however are considerably larger, some could be even up to 500 GB (I do have a few templates that I have to import with very large sizes).
In such a case, installing template time required would be about 75 seconds (best case).

During this time (installing template time) the whole bandwidth available for the storage network (if it even is on a separate NIC) would be used up by this process resulting in bad performance for the cluster.

Ways to fix this would be:
1. either compute the hash as the transfer is happening.
2. make it optional (maybe even opt-in) and do it in background only (maybe even limit the bandwidth used for this)

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Inizia tracciando il percorso di importazione dell’URL/upload del template attorno alla fase «Installing template», quindi individua dove il file scaricato viene riletto dallo storage e dove viene salvato il relativo hash. Confronta i passaggi di trasferimento e di lettura dallo storage. Il lavoro è completato quando la disponibilità del template non è bloccata da una lettura completa separata dallo storage, l’hash richiesto rimane corretto e il comportamento proposto per l’esecuzione in background o per la larghezza di banda è definito.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
java
Ambito
cloud, infrastructure, performance
Tipo di issue
Funzionalità
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Tranquilla
Chiarezza
Abbastanza chiara
Idoneità per principianti
45/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.