awslabs / awslabs/data-on-eks

Dedupe the Karpenter Nodepools In the spark operator blueprint

Open
#867 0 comments 0 reactions 0 assignees View on GitHub
data-on-eks enhancement
Dominant language
Shell
Stars
857
Forks
303
Avg merge
11h 5m
Merged PRs (30d)
3

Description

### Community Note

* Please vote on this issue by adding a 👍 [reaction](https://blog.github.com/2016-03-10-add-reactions-to-pull-requests-issues-and-comments/) to the original issue to help the community and maintainers prioritize this request
* Please do not leave "+1" or other comments that do not add relevant new information or questions, they generate extra noise for issue followers and do not help prioritize the request
* If you are interested in working on this issue or have submitted a pull request, please leave a comment

#### What is the outcome that you are trying to reach?

Reduce the complexity of the karpenter nodepool configuration for the Spark examples.

#### Describe the solution you would like

The karpenter nodepools could be combined and node selectors added to the examples to tailor the instance selection down. Ideally we can combine the nodepools that are using the same instances/configurations (or close enough).

#### Describe alternatives you have considered

#### Additional context

Contributor guide

Open the contributing guide

Research direction

Start by locating the Spark examples' Karpenter nodepool configurations and comparing their instance and other shared settings. Combine compatible nodepools, add node selectors as described, and verify that the examples still provide appropriately tailored instance selection with less duplicated configuration.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, kubernetes, spark
Domain
data-engineering, infrastructure
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.