stackabletech / stackabletech/hdfs-operator

Research faster spin up times

Open
#261 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Rust
Stars
53
Forks
9
Avg merge
1d 13h
Merged PRs (30d)
10

Description

When creating a HDFS cluster the data-nodes will get created one after each other.
Compared to hbase where all region-servers are started in parallel this takes quite some time if the number of datanodes is large.
We should research if there is the need to start all datanodes sequentially or if we can improve the startup time somehow

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The issue names HDFS cluster creation and the sequential startup of DataNodes, but no files, tests, or entry points. Compare this behavior with the parallel HBase region-server startup described in the issue, determine whether sequential startup is required, and document a measurable startup-time improvement if it is not.

Written by the indexing model from the issue text.

Assessment

Tech stack
kubernetes, rust
Domain
distributed-systems, infrastructure, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.