codeformuenster / codeformuenster/kubernetes-deployment
Cluster Infrastructure
Open
Nobody has claimed this yet.
epic
help wanted
- Dominant language
- No language data
- Stars
- 17
- Forks
- 4
- PR merge metrics
- No merged PRs in 30d
Description
FIXME Link here to doc about our current Kubernetes cluster and hosting setup.
- Monitoring
- Add customized dashboards using Grafonnet to kube-prometheus.
- cert-manager
- openebs
- Configure Alerts and send notifications to a matrix-channel.
- Maybe: matrix-alertmanager
- Add website analytics with Fathom.
- Create public status page with overview of current apps.
- Regularly check observatory.mozilla.org for all public sites.
- Add customized dashboards using Grafonnet to kube-prometheus.
- Authentication
- OpenID Connect via Keycloak for kube-apiserver and apps.
- Add gangway.
- Security
- Create restricted Pod Security Policy to only allow non-root.
- Default deny all ingress traffic
- RBAC
- Shared services
- Kinto
- Postgres
- Minio
- Elasticsearch
- Backup
- Push database snapshots and filestores regularly so some
s3storage.
- Push database snapshots and filestores regularly so some
- Stability
- Automatically replace the oldest node every twelve hours with a fresh one. Maybe with the help of kured.
- Make sure limits are set with every pod.
- Make every service be backed by at least two replicas. Label apps that can't deal with this.
- Set PodDisruptionBudget for all apps.
- Set recommended labels for all resources.
Random Ideas
- Try varnish with
trafficsandcrashes. - Add blackbox exporter for our public services.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the checklist in this issue and its linked Kubernetes, Grafonnet, kube-prometheus, and Matrix Alertmanager references; no repository file or test is named. Narrow the work to one unchecked infrastructure item before changing configuration. Done means the selected item is implemented and its operational outcome is verified.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- elasticsearch, grafana, kubernetes, postgres, prometheus
- Domain
- authentication, databases, devops, infrastructure, observability, security
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100