DataTalksClub / DataTalksClub/website
Implement AWS backup verification, alarms, and the operations dashboard
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 0
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
Parent: #66
Implementation repository: DataTalksClub/aws-infra
Product outcome
Implement the infrastructure-owned half of the approved observability and recovery contract: encrypted RDS backup/PITR/final-snapshot controls, scheduled backup verification, encrypted and retained safe evidence, exact query-backed alarms, notification routing, one operator dashboard, expiry monitoring, least-privilege runtime/delivery IAM, and bounded cost/scope controls for the website workload.
This issue is a cross-repository delivery record. Product authority and lifecycle evidence remain on this DataTalksClub/website issue; Terraform and infrastructure-runtime source belong only in DataTalksClub/aws-infra. It is not implementation-ready until the source catalog, backup protocol, root/state boundary, live-delivery authority, and HUMAN packet below are frozen.
Normative authority
_docs/PROCESS.md: role-separated engineer/tester/PM/on-call lifecycle, frozen evidence, and sensitive-artifact rules._docs/specs/07-security-privacy-operations.md: approved targets, release-critical signal families, encrypted backups/PITR, final snapshots, restore safety, expiry failures, named alert/runbook owners, and redacted evidence._docs/specs/08-aws-development-terraform.md:DataTalksClub/aws-infraownership, exact website workload boundary, encrypted RDS/S3/logging, observability/cost controls, separate OIDC plan/apply roles, protected approval, and production portability._docs/specs/09-migration-rollout-roadmap.md: production-shaped planning and the later restore/rollback/fault rehearsal gate._docs/specs/10-verification-strategy.md: Terraform policy/plan/readback, backup verification, alarms, recovery evidence, and no production-data/secret leakage.- Closed #26: approved target values. This issue does not reopen them.
- #264: authoritative provider-neutral
BackupVerificationRequirements,BackupSnapshotResult,BackupVerificationReceipt, and evaluators. - #265/#266: authoritative redacted source identifiers and the final zero-gap target/query/owner/alert/runbook/dashboard catalog.
- #78/#94: recurring protected Terraform delivery and the current legacy development root/state boundary.
Current audited baseline — not acceptance
Read-only inspection of DataTalksClub/aws-infra main at a83f6656d5ddd690d94079979b3f1a970551dac6 found:
- the current root remains
sandbox/websitewith state keysandbox/website/terraform.tfstateand legacywebsite-sandbox*physical identifiers; - RDS storage is encrypted and the example declares seven-day development backup retention, a backup window, retained automated backups, and a final snapshot, while development deletion protection remains disabled;
modules/django-website/observability.tfhas only basic web/worker running-task, ALB unhealthy/5xx, and RDS availability/free-storage alarms;- the example has no alarm actions, and no accepted scheduled backup verifier, encrypted verification-evidence contract, complete #266 query/alarm matrix, dashboard, expiry inventory, or functional alarm/notification evidence exists;
- the recurring Terraform workflow and policy tooling exist, but #78 still records live activation as HUMAN-blocked, and #94 has not migrated the current root/state identity.
This is a point-in-time source audit only. It neither pins the future implementation base nor proves live configuration, backup usability, alarm behavior, notification delivery, dashboard correctness, or operational readiness. Re-grooming must re-read the exact then-current external main, live authority, and dependency versions.
Dependency and dispatch gate
Keep needs grooming. Do not dispatch engineering in aws-infra until every item below exists and a PM rewrites this issue against the exact accepted versions.
- #264 is implemented, independently tested, PM-accepted, merged, and its exact schema/version/digest/evaluator interface is frozen.
- #265 is accepted and merged.
- #266 is in its final zero-gap state: every consumed row is
closure_ready, the canonical catalog digest is frozen, every query/source/window/exclusion/missing-data/action key is exact, and every owner/destination/escalation/runbook/panel key is approved. A source-onlyRefs #266slice is insufficient. - Every domain source/receipt referenced by the final #266 catalog is accepted and available in the intended environment. Generic event counts, health checks, deployment smoke, or an unresolved/null catalog row are not substitutes.
- #78's exact protected recurring plan/apply path, roles, permissions, evidence schema, and activation authority are accepted for this change.
- #94 either completes the root/state/OIDC migration or explicitly freezes the current compatibility root/state as this issue's approved target. The engineer may not choose between
sandbox/websiteanddevelopment/website. - The HUMAN packet below is answered and approved in a non-secret form, with protected values supplied only through the approved protected input/evidence path.
#267 is not a prerequisite for Terraform source or routine backup verification. It consumes the eventual accepted receipt and holds workload activation. #269 is downstream and owns isolated restore, RPO/RTO measurement, rollback/fault/expiry rehearsal, and clean return to service. Neither issue may be silently implemented here.
Source contract to freeze at re-grooming
The PM must pin all of the following to exact versions before engineering:
- External source identity: exact
aws-infrabase SHA, root, backend key, module path, provider lock, Terraform version, workflow SHA/path, development and production fixture paths, and the accepted #78 evidence schema. - Catalog identity: exact #266 catalog version/digest and the complete allowlist of row IDs, source contracts/versions/keys, CloudWatch query language/query, thresholds, periods/windows, missing-data behavior, exclusions/classifiers, alert actions, runbook keys, dashboard panel keys, owners, destinations, escalation, and review cadence.
- Backup identity: exact #264 requirements/result/receipt/evaluator versions, provider key/version, environment class, maximum recovery-point age, maximum receipt age, runtime identity, database-schema identity, safe snapshot identity projection, manifest definition, expected-resource arithmetic, error-code mapping, and canonical evidence digest.
- Execution bridge: exact owner and immutable artifact for the AWS snapshot adapter/verifier, its packaging and upgrade contract, schedule/event input, retry/concurrency/idempotency behavior, and how it consumes #264 without copying or weakening the website-owned schema. Terraform owns its AWS runtime, schedule, IAM, storage, and wiring; it must not invent a second receipt grammar.
- Evidence bridge: exact stable Terraform keys—not raw provider identifiers—for the encrypted evidence destination, KMS policy, prefix/object schema, retention/lifecycle, freshness marker, immutable artifact digest, and readback consumer. Raw AWS responses, ARNs/resource IDs where not already public authority, state, plans, URLs, credentials, database content, or provider payloads are not evidence.
Any source/catalog/protocol/root/workflow drift after grooming invalidates the handoff and returns this issue to PM.
Intended implementation scope after re-grooming
Only the final re-groomed contract may authorize these changes in DataTalksClub/aws-infra:
Backup and verification resources
- Environment-specific encrypted RDS automated-backup/PITR retention and backup windows, copy-tags behavior, retained automated backups, deletion protection, and final-snapshot policy that preserve the approved development/production differences.
- One bounded scheduled verifier runtime using the accepted immutable adapter/artifact and exact #264 requirements. It reads only the minimum safe provider metadata, produces only a canonical safe result/receipt, fails closed on absence/ambiguity/drift, and never reads database rows or performs a restore.
- Encrypted, public-blocked, versioned/immutable-as-approved verification evidence with exact retention/lifecycle, KMS policy, failed-write alarm, freshness alarm, and least-privilege producer/reader roles.
- Explicit schedule, timeout, memory/CPU, retry, dead-letter/failure destination, concurrency, idempotency, and maximum-run/cost bounds. A failed or stale verification remains non-green and loud.
Queries, alarms, notifications, and dashboard
- One Terraform-owned query/alarm mapping for each allowlisted final #266 row. Query text, source/version, units, periods, evaluation windows, comparisons, exact-boundary behavior, missing-data treatment, exclusions, and recovery behavior must remain byte/digest-traceable to the catalog.
- Separately attributable healthy, breach, recovery, missing, stale, version-drift, and query/panel-disagreement behavior. Expected validation denial, throttling, security denial, provider acceptance, delivery, ambiguity, and system failure may not be collapsed.
- Notification topics/subscriptions/routing only to the approved bounded destinations, with delivery-failure visibility, least privilege, owner/escalation metadata, and no secret/provider payload in subject or body.
- One Terraform-owned CloudWatch dashboard whose panel keys, queries, units, periods, labels, and alarm links exactly match #266. Missing/unknown is visible and non-green; dashboard agreement is verified by canonical API readback, not a mutable console screenshot.
- The approved expiry inventory and lead times for certificates, Relay client/callback-secret references, OIDC/trust material, database-credential synchronization, and other accepted provider dependencies. Monitoring must use safe metadata only and must not read or expose secret values.
IAM, security, portability, and cost
- Separate least-privilege Terraform delivery, verifier runtime, evidence writer/reader, query/alarm/dashboard, and notification permissions as required; explicit denials/boundaries prevent unrelated state, shared DNS, production from development, secret values, database contents, infrastructure outside the website workload, application release mutation, SES/Datamailer/provider send, and provider-event ingress.
- Encryption in transit/at rest, explicit retention, public-access block, bounded logs, deterministic redaction, safe labels, exact environment tags, and no sensitive/high-cardinality dimensions.
- Development and production fixtures use the same source shape with explicit environment inputs and separate account/backend/state/destinations. No development state/resource dependency is promoted to production.
- Explicit maximum resource count, schedule frequency/runtime, log/evidence retention, storage growth, notification volume, and monthly/one-time cost. Workload budget/cost-anomaly and approved 50/80/100% allowance/forecast alarms are included only from the accepted catalog and cost decision.
HUMAN decision packet — all fields required
The accountable owners must approve this packet before re-grooming. Do not infer an answer from labels, contributors, repository access, current resources, examples, incident history, or Terraform defaults.
- Environment and source: environment class; account and region; exact root/backend/state; current-root versus #94 migration disposition; exact source SHA; protected plan/apply workflow; operator/deployer role; maintenance/change window; go/no-go/abort/retry authority.
- Backup policy: exact RDS/resource scope; authoritative backup source; retention/PITR window; backup and maintenance windows; final-snapshot/deletion-protection/delete-automated-backup policy; copy/replication, if any; accepted recovery-point and receipt-age bounds; verification cadence and tolerated schedule delay.
- Verifier/evidence: adapter/runtime owner; immutable artifact and upgrade authority; expected-resource manifest/count; safe identity projection; evidence bucket/prefix/key reference; KMS/key owner; artifact retention/lifecycle; write/read roles; evidence audience; failure/DLQ handling; maximum attempts/runtime/concurrency.
- Every catalog row: accountable owner; notification destination; escalation owner/destination; exact query/window/exclusion/missing-data policy; breach/recovery/rollback/review action; runbook; dashboard panel; review cadence; exception owner and expiry.
- Expiry inventory: each credential/certificate/provider dependency; metadata source; warning and critical lead time; rotation/reconciliation owner; failure behavior; escalation; proof that no secret value is read or emitted.
- Cost and scope: maximum resources, schedule invocations/runtime, evidence/log storage and retention, notification volume, monthly recurring and one-time change budget, cost owner, 50/80/100% thresholds, anomaly action, and expansion/exception authority.
- Live evidence authority: permission to apply; permission to perform bounded safe verifier executions and synthetic alarm/no-data/recovery transitions; approved test dimensions and notification destinations; evidence retention/audience; stop conditions. This is not authority to restore, read protected data, send email, rotate a secret, mutate Relay/provider infrastructure, or run #269's drill.
Protected identifiers or destinations that are unsafe for a public issue must be approved through the protected workflow/input mechanism and represented here only by stable redacted keys plus an immutable approval/evidence digest.
Evidence ladder — classifications must not be collapsed
1. Source and policy evidence
- Frozen external base/head and diff digest;
terraform fmt; lockedinit -backend=false;validate; development and production mock tests; repository policy/IAM/workflow tests; secret/state/plan containment checks; and exact catalog/protocol fixture tests. - A deterministic plan-summary test proves expected Terraform addresses and rejects destroy/replacement, shared hosted-zone/backend mutation, public database/task ingress, unencrypted/unbounded storage/logs, wildcard IAM, secret values, unsupported rows, missing destinations, and cost/scope excess.
Passing this layer proves source consistency only. It does not prove a live plan, apply, backup, receipt, alarm, notification, dashboard, or cost.
2. Reviewed live plan evidence
- Exact current-main source/workflow identity, protected read-oriented role, refreshed locked state, exact root/backend/environment inputs, accepted catalog/protocol digests, redacted reviewable semantic plan, plan/input/artifact digests, no unapproved drift, and explicit HUMAN approval.
A plan proves proposed change only. It is not apply or readback evidence; no saved binary plan, unredacted rendering, state, credentials, sensitive values, or provider payload is attached to this issue.
3. Apply and convergence evidence
- The protected apply consumes the exact reviewed plan under the approved operator/window and produces the accepted redacted activation evidence.
- Post-apply inventory binds exact source/workflow/plan/apply identities and accepted Terraform keys/counts. A final refreshed locked plan reports terminal convergence/no change. Any failed/partial apply, replacement, unrelated drift, identity mismatch, or missing evidence is a mandatory hold, not success.
4. Functional readback evidence
- Canonical safe readback proves effective RDS backup/PITR/final-snapshot controls, encryption, retention, schedule, IAM, evidence lifecycle, alarms, dashboard queries, notification routing, expiry monitors, and cost controls match source and the exact catalog/protocol versions.
- One bounded safe verification produces an accepted #264 receipt at the healthy and exact freshness boundary; missing, partial, ambiguous, unencrypted, stale, future, wrong-runtime/schema/environment, digest/count drift, provider failure, and evidence-write failure stay blocked/non-green.
- Controlled synthetic or metric-math fixtures prove every alarm's healthy, exact boundary, breach, no-data/stale, recovery, and notification-failure behavior, plus query/alarm/dashboard agreement. No production record, secret, raw provider response, destructive action, or email send is used.
5. Downstream rehearsal evidence
Backup restoreability, measured RPO/RTO, isolated restore, tombstone replay, delivery reconciliation, active-content validation, workload release, rollback, expiry rotation, and clean return to service belong to #267/#269 and parent #66. They remain explicitly not_implemented_here; source, apply, readback, or a backup receipt cannot substitute for them.
Acceptance matrix after re-grooming
- The frozen source implements only the approved website workload resources at the exact root/state and consumes the exact #264/#266 identities without schema/query duplication or wildcard rows.
- RDS backup/PITR/final-snapshot/deletion controls match the approved environment policy and preserve encryption, non-public access, retention, and production portability.
- Scheduled verification is bounded, idempotent, least-privilege, version-pinned, fail-closed, and stores only canonical encrypted retained safe evidence; no restore, database-row read, or authority inference occurs.
- Every final catalog row maps exactly once to a query, alarm, owner, destination, escalation, runbook, panel, review cadence, missing/stale behavior, and tested healthy/boundary/breach/recovery disposition.
- Notification failure, dashboard/query disagreement, verifier failure/staleness, evidence-store failure, backup failure/freshness, and every approved expiry state remain visible and non-green.
- IAM and policy tests deny wildcard/unrelated infrastructure, state outside the exact key, shared DNS ownership, public database/task ingress, secret-value access, direct sender/provider mutation, and cross-environment access.
- Encryption, retention, redaction, bounded labels/cardinality, immutable identity, evidence containment, and safe error behavior pass canary tests without leaking identifiers, URLs/ARNs, credentials, state, plans, provider payloads, or production data.
- Development/production fixtures plan from the same source shape with separate account/backend/state/destination inputs and no development dependency in production.
- Scope and cost stay within the approved resource/invocation/runtime/storage/notification budgets, with the accepted budget/anomaly/allowance alarms and owner actions.
- Independent tester evidence classifies source, plan, apply, convergence, readback, functional alarm/notification/dashboard, and downstream rehearsal separately; no required item is skipped or inferred.
- [HUMAN] The named service, infrastructure, backup/recovery, security, cost, and escalation owners approve the exact packet, reviewed plan, live-operation window, bounded functional tests, and redacted evidence custody.
Lifecycle and closure
After re-grooming, an infrastructure engineer works in one isolated aws-infra worktree and leaves the candidate uncommitted. A separate tester recomputes the frozen source/evidence plan and verifies every acceptance criterion. PM source acceptance authorizes only a focused Refs #268 commit and local merge/push under the applicable repository process; on-call observes source CI. It does not authorize or claim a live apply.
The protected plan/apply/readback/functional-evidence phase then runs only under the separately approved HUMAN authority. #268 closes only after independent technical verification and PM acceptance cover the exact applied source plus all layer-4 evidence. If live authority remains pending, keep human/decision and the issue open.
No product/browser page changes are in scope. Playwright and screenshots are not_applicable; an AWS console screenshot is neither required nor sufficient and must not be used where deterministic redacted API/readback evidence exists.
Explicit non-goals
- No website repository code, Django model/service/job/route/API/Studio/template/UI, #264/#265/#266 domain implementation, or source-catalog invention.
- No production deployment, protected-data backup or read, database dump/row inspection, restore/PITR execution, replica promotion, tombstone replay, delivery/outbox reconciliation, content activation/repair, workload startup/release, application rollback, fault drill, or RPO/RTO claim.
- No Relay/provider infrastructure, sender identity, SES/Datamailer path, provider-event ingress, callback transition, secret creation/value mutation/readback, credential rotation, or email send.
- No raw provider/AWS response, state, binary/rendered plan, backend credentials, real tfvars, secret/token/cookie, database identifiers/content, user/subject data, URL/ARN, IP/header/query/body, exception text, or high-cardinality identifier in Git, issues, logs, dashboard labels, notifications, screenshots, or retained artifacts.
- No name-only hosted-zone lookup/creation, shared backend mutation, unrelated drift repair/import, root/state migration under #94, production-from-development state copy, public origin/database/task ingress, unencrypted storage/logging, wildcard IAM, mutable image tag, console-only resource/subscription, or unreviewed runtime registration.
- No missing-data-as-success, partial/ambiguous evidence accepted as green, query/dashboard disagreement ignored, alarm disabled to pass, stale receipt reuse,
--force/skip/ignore override, or eventual convergence treated as proof. - No Premium/advanced edge product, real-time log, targeted/advanced bot/fraud/challenge feature, cross-region backup copy, or expanded retention/schedule/resource/cost unless separately approved in the HUMAN packet and re-groomed scope.
Re-grooming exit
When every dependency and HUMAN field above is accepted, a PM must re-read the exact website and aws-infra sources and rewrite this issue with the actual source SHA, root/state/workflow, protocol/catalog digests, Terraform addresses, adapter artifact, queries, alarms, destinations, dashboard panels, IAM actions/resources, evidence keys, costs, and commands. Only then may needs grooming be removed and external engineering begin.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the dependency and dispatch gate, then read _docs/PROCESS.md and specs 07–10 before inspecting the pinned aws-infra source. Review sandbox/website and modules/django-website/observability.tf, but do not choose a root or implementation contract until the required issues and HUMAN packet are accepted. Done requires the frozen backup, catalog, delivery, evidence, IAM, alarm, dashboard, and cost controls to pass the specified verification gates.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, terraform
- Domain
- cloud, devops, infrastructure, observability-sre, security
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100