apache / apache/cloudstack

Storage bandwidth is wasted during template uploads/imports

Aberta
#5,697 6 comentários 0 reações 0 responsáveis Ver no GitHub
component:secondary-storage component:templates no-issue-activity status:stale type:improvement
Linguagem predominante
Java
Estrelas
3.1k
Forks
1.4k
Merge médio
6d 19h
PRs com merge (30d)
32

Descrição

At the moment when a template is being imported (via url or upload) there is a phase called "Installing template" or something similar. From what I noticed this phase calculates a hash and saves it to the database by reading the downloaded file over the storage network.
During this phase the template is not usable and I believe that should not be the case. I understand that the hash has a purpose and I am not saying it should be removed but I believe this should be an optional task that should be done in background.

In the following examples I do not consider disk speeds, just network speed/bandwidth.
For normal templates this is not necessarily noticeable. Consider this scenario (best case):
Template size: 4GB
Ingress bandwidth: 100 mb/s
Storage bandwidth: 1 gb/s
Download time required: 5.45 seconds
Installing template time required: ~0.6 seconds

Most templates (not ISOs) however are considerably larger, some could be even up to 500 GB (I do have a few templates that I have to import with very large sizes).
In such a case, installing template time required would be about 75 seconds (best case).

During this time (installing template time) the whole bandwidth available for the storage network (if it even is on a separate NIC) would be used up by this process resulting in bad performance for the cluster.

Ways to fix this would be:
1. either compute the hash as the transfer is happening.
2. make it optional (maybe even opt-in) and do it in background only (maybe even limit the bandwidth used for this)

Guia de contribuição

Abrir o guia de contribuição

Direção de pesquisa

Comece rastreando o caminho de importação de URL/upload de template em torno da fase “Installing template” e, em seguida, identifique onde o arquivo baixado é lido novamente a partir do armazenamento e onde seu hash é salvo. Compare as etapas de transferência e leitura do armazenamento. Considera-se concluído quando a disponibilidade do template não é bloqueada por uma leitura completa separada do armazenamento, o hash necessário permanece correto e o comportamento proposto em segundo plano ou de largura de banda está definido.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
java
Domínio
cloud, infrastructure, performance
Tipo de issue
Funcionalidade
Dificuldade
4/5
Tempo estimado
3-5 dias
Status de atividade
Pouca atividade
Clareza
Razoavelmente clara
Facilidade para iniciantes
45/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.