DataTalksClub / DataTalksClub/website

Write and rehearse restore, rollback, fault, and expiry operations

Open
#269 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

decision documentation human infra needs grooming operations P0 security testing
Dominant language
Python
Stars
0
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Parent: #66

Product outcome

Give an authorized operator one deterministic, fail-closed runbook for a synthetic production-like recovery rehearsal. The rehearsal must select and verify an authoritative backup, restore it only into the approved isolated target, measure the approved recovery objectives, keep every unsafe workload held, consume the accepted privacy/delivery/content receipts, either activate the exact intended release or abort safely, exercise immutable application rollback through #102, inject the approved fault/expiry cases, and return the environment to one known-good exact state.

This is child 6 and the final delivery lane under #66. It is intentionally blocked and retains needs grooming: the producer interfaces, applied infrastructure, operational owners, protected inputs, and live-operation authority do not yet exist in accepted form. This issue is not authority to write a runbook against guessed interfaces or to execute any AWS, restore, provider, credential, workload, or production action.

Normative authority

  • _docs/PROCESS.md: role-separated lifecycle, versioned verification, HUMAN criteria, operational evidence, and sensitive-artifact rules.
  • _docs/specs/01-platform-architecture.md: one Django deployment, after-commit durable work, immutable runtime identity, and business-service ownership.
  • _docs/specs/07-security-privacy-operations.md: approved service/recovery targets; encrypted backup/PITR; quarterly restore drills; restore holds; privacy tombstone replay; no historical resend; active-content validation; failure and expiry behavior.
  • _docs/specs/09-migration-rollout-roadmap.md: rehearsal, migration, workload hold, exact rollback, reconciliation, and cutover gates.
  • _docs/specs/10-verification-strategy.md: fault injection, restore/rollback evidence, production safety, no-resurrection/no-resend, and separate HUMAN deployed gates.
  • _docs/specs/open-decisions.md resolved decisions 10, 12, and 15: Relay is the sole sender, privacy periods are approved defaults, and production database RPO <=15 minutes/service RTO <=4 hours plus development RPO 24 hours/RTO one business day are approved.
  • #102: exact immutable web/worker restoration with fixed 240/420/720-second recovery bounds, receipt-bound terminal proof, singleton-worker invariant, and no success by eventual convergence.
  • #264/#266/#267/#268: accepted backup receipt/provider protocol, final target/query/owner catalog, live activation controller, and applied/read-back infrastructure are direct prerequisites.
  • #64/#258, #49, and the content-owning validator issue: authoritative tombstone replay, generation-bound delivery/outbox reconciliation, and side-effect-free active-content/pointer validation remain owned by those domains.

Approved recovery targets and clock contract

The objective values are already approved; #269 does not reopen them:

  • production database RPO: at most 15 minutes;
  • production service RTO: at most 4 hours;
  • development database RPO: 24 hours;
  • development service RTO: one business day.

Before re-grooming, the accountable recovery owner must bind each environment to exact named events supplied by accepted interfaces:

  1. incident_boundary_at: the authoritative UTC point at which protected service/data is declared lost or unavailable for this scenario;
  2. recovery_point_at: the canonical UTC point carried by the accepted #264 backup receipt for the selected restore point;
  3. rto_started_at: the monotonic-plus-UTC instant at the approved service-unavailable/maintenance transition, no later than the first recovery mutation;
  4. rto_completed_at: the first instant at which the exact restored database, exact application pair, accepted current receipt set, required public/auth/registration/enrollment checks, and all approved activation consumers are terminally ready; an intermediate healthy page, provider convergence, or partial consumer release is not completion;
  5. drill_completed_at: the later instant at which fault/rollback cases, evidence finalization, and clean return to the intended exact release finish. This is reported separately and cannot replace or shorten RTO.

rpo_age = incident_boundary_at - recovery_point_at; negative/future recovery points fail. rto_duration = monotonic(rto_completed_at) - monotonic(rto_started_at). Equality at the approved maximum passes; one unit beyond fails. UTC timestamps provide correlation, while monotonic time controls duration. Clocks never reset across retry, replay, reconciliation, activation, abort, #102 compensation, operator handoff, or provider wait. Any pre-approved exclusion remains visibly itemized and does not disappear from the raw elapsed duration. The owner must define the development business-day calendar, timezone, holidays, maintenance-window treatment, and every permitted exclusion before re-grooming. #102's 720-second phase cap is a nested bound and never extends the overall RTO.

Dependencies and blocked work

Engineering and rehearsal remain undispatched until all of the following are accepted, merged/applied where applicable, and pinned by immutable version/digest:

  • #264's provider-neutral requirements/result/receipt/evaluator and authoritative recovery-point semantics.
  • #265's redacted correlation/evidence schema as consumed by the final #266 catalog.
  • #266's zero-gap, closure_ready catalog, including every owner, destination, threshold/window/exclusion, alarm, dashboard panel, runbook key, escalation and review cadence.
  • #267's live restore/startup activation controller, exact receipt-set grammar, checkpoint/generation arithmetic, consumer hold/re-hold states, recovery-only seam, activation/abort terminal states, and evidence.
  • #268's exact applied infrastructure: backup/PITR/final-snapshot policy, scheduled verifier/evidence store, alarms/dashboard/notifications/expiry monitors, IAM, readback, convergence, and bounded cost evidence.
  • #64/#258's current authoritative subject-key/tombstone ledger, snapshot/high-watermark, replay/fence evaluator, legal-hold/exception behavior, retained proof, and privacy release authority.
  • #49's accepted logical-delivery/Relay contract and exact generation-bound outbox/reconciliation receipt, safe-state arithmetic, ambiguity resolution, provider reconciliation and delivery release authority.
  • A separately filed and accepted content-owned side-effect-free validator for the direct-sync model selected by #226, including active source/commit/pointer/projection identity, retained last-known-good behavior, and release authority. #269 must not revive the superseded site-wide ContentRelease model.
  • #102's exact-pair recovery interface and remaining HUMAN recovery drill, or a documented owner decision defining how the #269 rehearsal supplies that still-open evidence without weakening #102.
  • #78/#94 and any other infrastructure prerequisites selected by #268 have terminal accepted/applied identities; #269 consumes their result and does not repair Terraform state or OIDC.
  • Every HUMAN input and protected authority below is complete.

Any dependency/interface/catalog/infrastructure/runtime/runbook-source drift after final grooming invalidates the handoff. A PM must re-read the exact accepted sources and re-pin the runbook before engineering or rehearsal resumes.

Required runbook phases after re-grooming

The final runbook must be executable as an ordered state machine. Every phase records start/end, exact inputs/digests, owner, disposition, next permitted transition, safe abort, and redacted evidence. Missing, stale, partial, ambiguous, contradictory, expired, or mismatched evidence is a hold/failure, never a prompt to choose “latest,” force, skip, infer, or wait for convergence.

0. Authority, identity, and preflight freeze
  • Verify the approved environment/account/region using stable protected keys, operator/deployer/recovery roles, maintenance window, source and isolated target, synthetic-data manifest, budget/cost cap, evidence audience/destination/retention, communication/escalation tree, and separate go/no-go/abort/return-to-service authorities.
  • Pin exact website and infrastructure source SHAs, workflow/runbook versions, VERSION/source SHA/image digest, database schema identity, #264/#266/#267 receipt/evaluator/catalog versions and digests, #102 controller identity, Terraform root/state/workflow identity, and every domain receipt contract.
  • Prove the source environment, isolated target, evidence destination, and recipient/sender/provider boundaries cannot be confused. Reject name-only discovery, mutable tags, public/unprotected targets, production data outside the approved authority, or an unclean prior drill.
  • Start the evidence chain and RTO clock only at the approved events; perform no mutation during an incomplete preflight.
1. Establish holds and recovery isolation
  • Put all accepted activation consumers into their exact #267 held/recovery-only states before restore: public web behavior as approved, authentication/session issuance, web mutations, ordinary worker leases, scheduler/cron, delivery/outbox dispatch, exports, content sync/activation, cache/search/graph rebuilds, callbacks that could mutate state, and deployment finalization.
  • Preserve only the explicit recovery-only handlers and read paths authorized by #267. Each handler is deny-by-default, capability-scoped, private/no-store/noindex, auditable, and unable to send email or bypass a domain receipt.
  • Prove no old or new worker can lease held jobs, no second sender/Datamailer path is enabled, no callback can advance an unbound generation, and no consumer can self-release. A failed hold or re-hold handshake is an immediate abort before restore.
2. Select and verify the recovery source
  • Enumerate only the approved manifest through #268's bounded adapter. Select one authoritative candidate by explicit immutable identity and requested recovery point; never search for “latest,” accept a console label, or use a Git repository as a database backup.
  • Validate the exact #264 requirements/result/receipt, encryption, provider/environment/runtime/schema identity, expected/present/verified counts, canonical digest, creation/verification/expiry ordering, evidence retention, and recovery-point/RPO arithmetic.
  • Record rejected candidates only by allowlisted reason/count. Missing, stale, future, partial, unencrypted, ambiguous, corrupt, wrong-environment/schema/runtime/provider, or expired evidence aborts with no restore.
3. Restore into the approved isolated target
  • Restore only through the accepted #268 procedure and least-privilege role into the separately identified isolated target. No overwrite, promotion, source mutation, public ingress, sender/provider access, or production workload activation is permitted.
  • Bind every provider action to the selected receipt and one drill generation/checkpoint. Capture only canonical safe status and timing; raw provider payloads, database rows, dumps, identifiers, credentials, endpoints, SQL, and exception text are prohibited evidence.
  • Validate encryption, network isolation, engine/settings compatibility, runtime/database-schema identity, pending migration policy, and exact expected resource shape before any application-level read. Provider “available” is not restore success.
4. Replay privacy authority and reconcile retained state
  • Apply #258's accepted replay using the authoritative pre-restore and current tombstone ledger snapshots/high-watermarks, generation/checkpoint, subject-key grammar, retention horizon, and legal-hold/exception policy. Replay is idempotent and crash/retry safe.
  • Prove erased/quarantined/disabled subject data cannot be read, authenticated, projected, cached, indexed, exported, delivered, or used by jobs after restore. The accepted privacy receipt must account for every required tombstone and exception without exposing subject material.
  • Reconcile dynamic records required by the accepted recovery plan, including post-recovery-point registrations/enrollments, through their owning service/interface. No whole-database/image rollback may silently discard them, and no guessed merge/direct SQL is permitted.
5. Reconcile delivery/outbox and durable work
  • Consume #49's exact generation-bound reconciliation receipt and compare website intent/job state with Relay's redacted authoritative projection. Provider acceptance remains distinct from delivery.
  • Historical, restored, duplicate, stale-generation, suppressed, terminal, ambiguous, expired, or unmatched intents/jobs remain held according to the owning contract. Ambiguity never triggers automatic resend; no bulk “retry all,” provider fallback, SES direct path, or Datamailer re-enable is allowed.
  • Job lease expiry/recovery, scheduler state, deduplication and callbacks must remain fenced to the drill generation. Reconciliation can release only the explicit safe set under named delivery authority, and #269 itself never decides transport state or sends a message.
6. Validate content and product-domain integrity
  • Consume the accepted content-owned validator against the direct-sync model: exact active sources/commits/pointers/projections, expected retained last-known-good state, route/link/search inputs, and drift/error disposition. The runbook never edits content tables/pointers directly.
  • Run bounded, synthetic, row-safe domain reconciliation for account/auth, registrations, enrollments, course/event integrity, durable jobs, and migration/schema invariants. Validate counts/digests/arithmetic only through accepted safe interfaces; do not expose records or production identifiers.
  • Validate required public reads plus registration/enrollment writes against approved synthetic canaries while all email effects stay disabled/held. Cache/search/graph warming or rebuild occurs only through its accepted generation-bound procedure and cannot publish stale/deleted content.
7. Activation decision or safe abort
  • Re-evaluate one exact current receipt set atomically through #267: backup, privacy, delivery/outbox, content, runtime/schema, checkpoint/generation, consumer manifest, and expiry state. No mixed-generation, partial, old-but-valid, or “eventually consistent” set is eligible.
  • ACTIVATE requires named HUMAN go/no-go approval plus all exact receipts current and passing. Release consumers only in #267's fixed order with acknowledgement after each transition; an attempted early start, drift, timeout, or contradiction immediately re-holds the accepted consumer set.
  • ABORT leaves the source untouched, keeps unsafe consumers held, records the allowlisted reason, runs the accepted isolated-target cleanup/retention path, and escalates. Abort is a successful safety outcome but never an RPO/RTO, restore, or drill pass.
  • Stop the RTO clock only at the exact rto_completed_at definition. Partial availability, a green load balancer, one healthy workload, a provider receipt, or later convergence does not stop it.
8. Immutable application rollback and post-cutover writes
  • From the activated synthetic candidate, create approved post-cutover registration/enrollment/durable-work canaries, then invoke only #102's exact immutable rollback path to the recorded prior VERSION/source SHA/image digest and exact web/worker task definitions/counts/receipts.
  • Preserve #102's fixed web/worker/phase 240/420/720 bounds, fixed initiation/observation rules, exact terminal pair, prior-SHA health, singleton worker, evidence redaction, and red original-operation result. Never use a reverse migration, mutable tag, fabricated timestamp, alternate image, or ultimate convergence as proof.
  • Prove the post-cutover dynamic writes remain present exactly once, idempotency/deduplication remains intact, held email is not sent, no duplicate job is leased, and the accepted content/privacy/delivery receipt set still matches or forces a re-hold/reconciliation before service resumes.
9. Fault and expiry matrix

For each approved case, the runbook must pin injection mechanism/owner, synthetic boundary, precondition, expected signal/query/alarm, safe degradation/hold, customer/operator impact, escalation deadline, recovery/rotation/reconciliation steps, abort threshold, terminal proof, and cleanup. Cases run one at a time unless a separately approved combination is necessary; an alarm test or simulated metadata fixture is not described as a real outage.

Family Minimum states to exercise Required safe behavior
GitHub/content source unavailable, invalid/stale commit, sync failure retain accepted last-known-good/direct-sync status; no unsafe pointer publication
Search/graph/cache builder/query failure, stale projection, invalidation failure retain/degrade without losing canonical pages or exposing deleted/private data
Worker/scheduler/jobs worker down, lease expiry, duplicate/replayed job, singleton risk idempotent fenced recovery; no duplicate mutation/send; loud backlog/heartbeat alarm
Relay submission timeout, unavailable, accepted-versus-delivered distinction, ambiguous acknowledgement no direct provider fallback and no automatic resend
Relay callback/reconciliation bad signature, replay, loss, reordering, stale generation, reconciliation unavailable reject/fence safely; hold ambiguity; reconcile before release
OIDC/deployer unavailable/denied/expired trust fail before mutation; no static credential fallback or widened trust
Database/restore unavailable, wrong schema/runtime, partial restore, connection loss during phase holds remain; transaction/idempotency boundaries win; no partial readiness
Edge/origin origin/ALB failure, invalidation failure, WAF/cache safe simulation approved private/maintenance or safe degradation; no broadened cache/origin access
Backup/verifier/evidence missing, stale, corrupt, partial, unencrypted, provider error, evidence-write failure no restore/activation; exact alarm/escalation; old success cannot mask current failure
Expiry warning, exact warning boundary, critical boundary, expired certificate/OIDC trust/Relay client/callback-secret reference/database credential synchronization/other final #268 inventory row alert at approved lead time; fail closed; named rotation/reconciliation only; no secret value read/emitted

Expiry exercises use safe metadata/synthetic fixtures unless the protected owner separately authorizes a real rotation. A credential/certificate must not be deliberately expired, rotated, read, or displayed merely to satisfy this issue.

10. Clean return and evidence closure
  • Return to the exact owner-selected intended release, not implicitly the pre-drill or newest release. Revalidate the full exact receipt set, runtime/application pair, database schema, consumer states, active content, synthetic domain canaries, alarms/dashboard, and public/auth/registration/enrollment checks.
  • Remove or retain the isolated target and synthetic data only under the approved cleanup/evidence policy. Prove no drill role, ingress, sender, callback, worker, schedule, hold, temporary exception, alarm suppression, credential, snapshot copy, or mutable resource remains.
  • End with exactly one terminal disposition: passed, aborted_safe, or failed_unresolved. A passed drill requires every required scenario and clock; aborted_safe proves containment only; any missing terminal proof remains failed_unresolved and loud. File follow-up issues for every deviation with owner/expiry; do not edit evidence to make it green.

Activation, abort, rollback, and escalation authority

The final runbook must separate these capabilities; one role/label must not silently imply another:

  • drill sponsor approves scope/cost and may cancel before execution;
  • recovery operator performs only the frozen steps in the approved window;
  • privacy owner accepts tombstone source/replay and no-resurrection proof;
  • delivery owner accepts ambiguity resolution and the safe release set but does not authorize infrastructure activation;
  • content owner accepts active-source/pointer/projection validity;
  • infrastructure owner authorizes backup selection, isolated restore, infrastructure mutation and cleanup;
  • release/on-call owner authorizes #102 rollback and exact application return;
  • go/no-go authority alone authorizes activation;
  • abort authority may stop/re-hold at any phase without waiting for go/no-go;
  • evidence custodian controls the encrypted destination, audience, retention and deletion;
  • incident/escalation owners receive alarm/failure handoff and own unresolved recovery.

Separation-of-duty and break-glass details, if any, require explicit protected approval. There is no --force, skip, ignore, choose-latest, manual database edit, raw console workaround, or broad administrator shortcut in the accepted path.

HUMAN decision and protected-input packet

Every field is required before final re-grooming. Public issue text records stable redacted keys and approval/evidence digests only; protected values stay in the approved confidential channel.

  1. Environment and identity: environment class; account/region; exact source and isolated target keys; Terraform root/backend/state/workflow; website/infrastructure source SHAs; intended return release; runtime/schema identities; network boundary; synthetic-data manifest.
  2. Recovery source: authoritative backup class/provider; candidate-selection rule; retention/PITR/final-snapshot policy; accepted recovery-point/receipt-age bounds; expected resource manifest; immutable #264/#268 evidence identities.
  3. Clock interpretation: exact five clock events above; inclusive boundary/unit; development business-day calendar/timezone/holidays; maintenance and third-party wait treatment; permitted exclusions; clock/evidence source; who may declare start/completion/failure.
  4. Owners and authority: named sponsor, operator, infrastructure, database, privacy, delivery, content, auth, worker, edge, security, release/on-call, go/no-go, abort, escalation, evidence and cost owners; primary/backup contacts and destinations.
  5. Window and blast radius: maintenance/change window; maximum duration, AWS/provider operations, resources, concurrency, retries, storage, notifications and cost; protected/live versus synthetic boundaries; stop conditions and emergency containment.
  6. Activation consumers: exhaustive consumer manifest, held/recovery-only/released behaviors, fixed release order, acknowledgement/re-hold handshake, customer-visible maintenance/readiness/auth/callback behavior, and escalation per consumer.
  7. Domain authority: exact privacy ledger/high-watermark and exception policy; delivery safe-state/reconciliation generation and release authority; content validator identity; post-recovery-point dynamic-write reconciliation source and owner.
  8. Fault/expiry inventory: each approved injection, safe mechanism, lead time/boundary, expected signal/alarm, rotation/reconciliation owner, abort/recovery action, combination policy, and proof that secrets/provider payloads are never read or emitted.
  9. Evidence and communications: encrypted destination/prefix stable key, KMS/evidence owner, schema/version, audience, retention/legal hold, notification/escalation channels, redaction review, incident linkage, and permitted issue summary.
  10. Cleanup/return: intended exact release, isolated-target disposition, snapshot/evidence/synthetic-data retention, temporary-role/ingress/schedule/alarm-suppression removal, final validation owner, and unresolved-failure custody.
  11. Explicit permissions: separate authorization for isolated restore, protected infrastructure mutation, synthetic application writes, alarm/fault simulation, #102 rollback, credential-expiry simulation/real rotation, activation, abort, cleanup, and public evidence summary. Permission for one does not grant another.

No repository role, issue assignment, prior incident, existing AWS access, label, current resource, example value, or past deployment supplies these decisions.

Evidence contract

The final immutable drill report must bind:

  • runbook/schema/catalog/receipt/evaluator/controller versions and digests;
  • website/infrastructure source and workflow identity plus VERSION/source SHA/image digest/database schema identity;
  • environment/source/isolated-target/synthetic-manifest stable keys and one drill generation/checkpoint;
  • approved owner/authority packet digest, window, cost/scope envelope and evidence policy key;
  • UTC and monotonic phase timestamps, raw/included/excluded RPO/RTO arithmetic, exact boundary verdicts, #102 nested timings and total drill duration;
  • each hold/re-hold/release acknowledgement and exact consumer disposition;
  • backup, privacy, delivery, content, runtime/schema and activation receipt digests/statuses without protected payloads;
  • fault/expiry case identity, expected/observed query/alarm/escalation/recovery disposition, and cleanup;
  • pre/post synthetic canary arithmetic for tombstones, registrations/enrollments, jobs/deduplication and content, expressed only through approved safe counts/digests;
  • exact terminal application pair, intended release, alarm/dashboard state, isolated-target cleanup/retention, and one terminal drill disposition;
  • every exception/deviation/follow-up with owner, risk, expiry and approval.

Evidence is append-only/canonical as approved, encrypted, access-controlled, retained for the approved period, and independently recomputable. It must never contain production database contents/rows/dumps, subject/account/email/profile/submission data or reversible hashes, raw tombstones, message bodies, recipient/provider payloads, raw AWS responses/state/plans, URLs/ARNs where not explicitly accepted safe identity, host/database names, IP/header/query/body, cookies/tokens/credentials/secrets, exception text/stack traces, screenshots of protected consoles, or high-cardinality metric labels. Redaction failure is a drill failure and security escalation.

Repository/source tests, Terraform validation, a console screenshot, an available restored database, a green load balancer, provider acceptance, eventual convergence, or an operator statement cannot substitute for the exact evidence ladder.

Acceptance matrix after final re-grooming

  • The versioned runbook implements phases 0-10 as one deterministic state machine with exact prerequisites, owner/capability, inputs, transitions, timeout/expiry, evidence, abort, and cleanup for every phase.
  • Production and development clock fixtures prove the named RPO/RTO events, monotonic duration, UTC correlation, inclusive boundary, one-unit-over failure, no retry/reset, approved exclusions, business-day calendar, and nested #102 bounds.
  • Backup selection/restoration consumes the exact accepted #264/#268 identities, restores only the approved isolated target, and fails before restore/activation on missing, stale, future, partial, ambiguous, unencrypted, corrupt, mismatched, expired, or wrong-environment/runtime/schema evidence.
  • The exact #267 consumer set remains held through restore/replay/reconciliation/content/domain validation; early startup, mixed-generation receipts, drift, timeout, or partial release re-holds and aborts safely.
  • #258 replay proves every required tombstone/exception and prevents erased data from use, authentication, projection, cache/search, export, job, or delivery after restore.
  • #49 reconciliation proves no historical/duplicate/ambiguous intent is automatically resent, no Datamailer/direct-provider/dual-sender path exists, and only the explicitly authorized generation-bound safe set can release.
  • The accepted content validator retains the correct direct-sync active source/pointer/projection and last-known-good behavior without #269 directly mutating content-owned state.
  • Synthetic product-domain checks prove post-recovery-point and post-cutover registrations/enrollments persist exactly once, jobs remain idempotent/fenced, no duplicate leases/sends occur, and migrations/schema/content remain coherent.
  • Activation requires one exact current receipt set and named go/no-go; safe abort leaves source unchanged, consumers held, evidence loud, and isolated-target cleanup/retention controlled.
  • #102 rollback retains its exact identities, 240/420/720 bounds, receipts, singleton invariant, terminal pair and red failure semantics; no reverse migration, dynamic-write loss, or eventual-convergence inference occurs.
  • Every approved GitHub/content, search/graph/cache, worker/scheduler/job, Relay submission/callback/reconciliation, OIDC, database/restore, edge/origin, backup/verifier/evidence and credential-expiry state produces the exact safe degradation, query/alarm, escalation, recovery and cleanup.
  • Expiry warning/exact-boundary/critical/expired fixtures match the final #268 inventory and lead times, fail closed, use named rotation/reconciliation authority, and expose no secret value.
  • Independent tester evidence proves the canonical report, redaction/security canaries, phase/clock arithmetic, receipt identities, fault matrix, exact terminal state and clean cleanup; no required item is skipped or inferred.
  • Browser/screenshots are not_applicable unless the final #267 design exposes a human-facing maintenance/recovery surface. If it does, that owning issue must define and pass private/no-store/noindex, authorization/denial, accessibility and inspected desktop/mobile evidence before #269 runs.
  • [HUMAN] Every named owner approves the exact protected packet, source/target, maintenance window, synthetic boundary, scope/cost, RPO/RTO interpretation, fault/expiry mechanisms, evidence custody, activation/abort and return plan.
  • [HUMAN] The authorized rehearsal meets the environment's RPO/RTO, proves no resurrection, historical resend, lost/duplicated dynamic write/job, identity weakening, or silent convergence, exercises an actual safe abort and #102 rollback, and returns to the intended exact release.
  • [HUMAN] Privacy, delivery, content, infrastructure/database, security, release/on-call, evidence and final product owners separately accept their receipts, observed behavior and terminal evidence. Any unresolved criterion keeps the issue open and the affected environment held or in its approved safe state.

Repository and rehearsal verification scenarios

Before live authority, deterministic synthetic tests must cover every phase transition; absent/duplicate/out-of-order/mixed-generation/stale/expired/malformed receipts; crash/retry at every checkpoint; inclusive/over-bound clocks; hold/re-hold and early consumer start; privacy replay before/during/after deletion and legal-hold exception; safe/ambiguous/terminal/replayed delivery states; content drift/stale pointer; post-recovery-point writes; #102 rollback success/failure/deadline/singleton cases; every fault/expiry row; abort at every mutation phase; evidence write/redaction failure; and cleanup/return mismatch.

The independent rehearsal tester recomputes the exact frozen source/evidence plan and observes the authorized operation without becoming its operator or approver. A dry-run/simulation is classified separately from isolated restore, activation, fault injection, rollback, and live readback. Source CI and synthetic tests may land with Refs #269, but the issue cannot close until the authorized evidence and all independent HUMAN acceptances are terminal.

Explicit non-goals

  • No implementation or execution under #269 while needs grooming remains; no repository/code/test/runbook/drill/AWS/production/provider/credential/commit/push/merge/deploy action is authorized by this blocked contract.
  • No ownership of #264 backup grammar/provider adapter, #265 correlation schema, #266 catalog/query semantics, #267 activation controller, #268 Terraform/AWS resources, #258 tombstones/privacy decisions, #49 delivery/Relay states, content business rules, #102 recovery controller, or domain mutations.
  • No production/protected-data dump, row inspection, ad hoc copy, source overwrite, in-place restore, replica promotion, direct SQL, destructive reverse migration, whole-database/image rollback that loses dynamic writes, or Git-as-database-backup assumption.
  • No website sender, SES/Datamailer fallback, dual sender, provider-event ingestion workaround, automatic resend, retry-all, provider-accepted-as-delivered, or delivery ambiguity resolution by #269.
  • No direct content-row/pointer mutation, resurrection of the superseded site-wide staged ContentRelease design, silent cache/search rebuild, or publication before the content owner accepts the exact state.
  • No mutable image/tag, fabricated release identity/time, broadened IAM/OIDC/network/cache/origin access, static-credential fallback, secret readback, unencrypted/unbounded artifact, public recovery surface, or name-only resource discovery.
  • No --force/skip/ignore/choose-latest/manual-console-success, missing-data-as-green, partial receipt set, stale receipt reuse, clock reset, hidden exclusion, alarm suppression, automatic consumer release, or eventual convergence accepted as proof.
  • No credential/certificate value read, deliberate real expiry/rotation, external provider mutation, real recipient/message, or combined fault blast radius without separate explicit protected authority.
  • No product redesign. Documentation is the eventual repository change; infrastructure/application/domain changes belong to their owning issues. Protected-console screenshots are neither required nor sufficient evidence.

Re-grooming and closure

After every dependency and HUMAN field is accepted, a PM must re-read the exact accepted website and aws-infra sources and rewrite this issue with the real interface/type/state/error names, consumer manifest/order, receipt/catalog/controller digests, Terraform/runbook/workflow identities, protected stable keys, clock events, commands, fault mechanisms, alarms/destinations, costs, authorities and evidence schema. Only then may needs grooming be removed and a documentation engineer begin in an isolated worktree.

The engineer writes the versioned runbook and synthetic fixtures only, leaves the candidate uncommitted, and posts the versioned verification handoff. A separate tester verifies it; PM source acceptance permits only a focused Refs #269 commit because live HUMAN evidence remains. The orchestrator merges/pushes under PROCESS and on-call observes source CI. The separately authorized rehearsal then follows the frozen runbook with distinct operator/tester/domain/PM/HUMAN gates. Use Closes #269 only after the exact authorized rehearsal, clean return, independent evidence and every HUMAN acceptance pass; otherwise retain human/decision and keep the issue open.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with _docs/PROCESS.md and the listed architecture, security, roadmap, and verification specifications, then review prerequisites #264, #266, #267, #268, and #102. Do not implement or execute the rehearsal until those interfaces, infrastructure, owners, and HUMAN inputs are accepted and pinned. Done is an ordered runbook and rehearsal that records evidence, meets the approved recovery objectives, exercises the required cases, and returns the environment to the intended exact state.

Written by the indexing model from the issue text.

Assessment

Tech stack
django, python
Domain
devops, documentation, infrastructure, security
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.