DataTalksClub / DataTalksClub/website

Epic: Sync and serve GitHub content with canonical people identities

Open
#4 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

content epic integration P0
Dominant language
Python
Stars
0
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Normative authority:

PM disposition

GROOMED / OPEN / NOT ACCEPTED. This is a coordination epic, not a broad engineering lane. Its child issues own the independently testable source adapters, public projections, search/graph projection, direct-sync migration, and management contracts. Every child still requires the normal engineer → independent tester → PM → focused commit → local no-ff merge/push → on-call lifecycle.

Product outcome

Coordinate the migration of GitHub-owned editorial content to one Django deployment while preserving each accepted public route, link, fragment, asset, metadata, source relationship, and failure contract. The current specification names five editorial repositories; this epic does not infer that all five are enabled or already equivalent. The exact repository, branch, path, adapter, ownership, rollout, and public-authority choices are recorded by the direct-sync parent #38 and its accepted source/family manifests.

The website never writes editorial commits, branches, or pull requests to GitHub. The accepted lifecycle decision is the direct-sync model from #226: an allowlisted immutable checkout is parsed and validated, source-owned records are directly upserted, and source-scoped draft/soft-delete state is the public visibility boundary. Ordinary site-wide ContentRelease candidate/ready/activate/rollback is not product authority. Historical ContentRelease rows and their provenance remain read-only migration evidence under #219 and the final retirement contract #278.

Source and projection boundaries

The five editorial source families named by specification 03 are:

  • DataTalksClub/content: structured articles, podcast metadata and separate transcripts, books, and adopted media;
  • DataTalksClub/datatalksclub.github.io: the remaining legacy-main editorial collections and their migration provenance;
  • DataTalksClub/docs: Docs pages, navigation, and assets;
  • DataTalksClub/faq: FAQ courses, sections, questions, and JSON source; and
  • DataTalksClub/podwiki: Wiki pages, typed links/citations, graph, and search source.

This is the current source inventory, not a live-source approval. #38 owns the exhaustive source rollout/ownership manifest and the source-specific direct-sync contracts. Course-owned records and operational course data remain with the course epics and are not silently added to this editorial scope.

Checked/baked projections are compatibility evidence until their owning source/family and public-reader cutover gates pass. #253 owns the reproducible Main/Podcast/Book/public-projection envelope and is not yet accepted. The source-specific boundaries are:

  • #39 — remaining legacy-main adapter and later public tool/conference parity, with opaque Person keys until #40;
  • #40 — exact Person short identity, aliases, and relationship resolution, strictly separate from accounts/member profiles;
  • #41 — Docs adapter, checked/public compatibility, and later direct-sync/search integration through their owning contracts;
  • #42 — FAQ adapter, feeds, anchors, and later reader cutover through #276; and
  • #43 — normalized Podwiki projection onto the sole public /wiki family, with #44 downstream.

The checked projection is never a license to hand-edit generated bytes, follow a moving source branch, or claim database authority. Public-reader cutover is source/family-scoped and belongs to #38 / #276 after exact source/projection parity and the required owner evidence.

Search and graph boundary

#44 owns the unified Django search/graph projection: public-safe document construction, ranking/query/filter behavior, graph/link validation, build identities and digests, projection activation/fallback/rollback, parity, and the Lambda-retirement gate. Its candidate/active/rollback vocabulary describes the search/graph projection lifecycle, not the rejected source ContentRelease lifecycle. #44 consumes accepted public DTOs and source seeds; it does not ingest GitHub sources or decide public content authority. Direct-sync children call the accepted #44 interface where required.

The ownership sequence is intentionally acyclic:

#253 + #294 + #40 → #43
#292 → focused Docs adoption/public-projection slice under #41
#293 → focused FAQ adoption/public-projection slice under #42
accepted public inputs → #44 search/graph projection
#219 + source/ownership manifest + #253 → #38 direct-sync slices #273–#278
#38/#276 source-family reader cutover → #278 staged-path/spec/runbook retirement

#72, #76, and #77 own project-wide classification, evidence-producer, and release-report consumers. This epic supplies accepted source/projection evidence to those gates; it does not duplicate their report or grant rehearsal, provider, deployment, or production authority.

Child ledger

The ledger records delivery ownership, not acceptance. A checked box below is permitted only after that issue's complete lifecycle and post-push evidence pass.

  • #12 — closed source-authority/no-write-back decision; its old staged examples are superseded by #226.
  • #24 — closed PostgreSQL search/public-contract decision; implementation remains #44.
  • #37 — closed historical immutable-release/read-model foundation; it is not current direct-sync authority.
  • #38 — open needs grooming direct-sync parent; owns source authority, ingestion, reconciliation, family cutover, management parity, and #278 final contract removal through its bounded children.
  • #39 — open needs grooming remaining legacy-main adapter/public parity lane.
  • #40 — open needs grooming exact Person source/resolver/relationship lane.
  • #41 — open groomed Docs epic, blocked on its offline/adoption inputs and later direct-sync/public gates.
  • #42 — open groomed FAQ epic, with #293 first and direct-reader cutover downstream.
  • #43 — open groomed /wiki projection/presentation lane, blocked on #253, #294, and #40; #44 is downstream.
  • #44 — open groomed search/graph projection lane, blocked on accepted public inputs and its own parity/management gates.

The closed foundations #103 (network-free structured content adapter) and #156 (provider-neutral webhook authenticity/delivery fence) are reusable prerequisites, not substitutes for #38 or its children.

Epic completion gate

  • The five-source inventory, source ownership/rollout manifest, and source/family public-authority cutover manifest are explicitly accepted; no source, branch, adapter, path, relationship, or baked/direct boundary is inferred.
  • #39–#43 deliver accepted deterministic source adapters/public projections and compatibility evidence, with exact Person and cross-source relations, routes, links, fragments, assets, metadata, and approved exceptions.
  • #38 and its direct-sync children deliver source-owned direct upsert, partial-recovery, ingress/reconciliation, status/evidence, management parity, source/family cutover, and the safe expand/reconcile/contract path; #278 alone owns final staged-path/spec/runbook removal.
  • #44 delivers the accepted backend-portable search/graph projection and its separate build, parity, fallback, rollback, and management contracts without taking source-ingestion or public-content authority.
  • All required focused/full verification, independent tester reports and screenshots for render-impacting children, PM acceptance, focused commits, local no-ff integration, push CI, deployment/readiness, and on-call evidence pass for the applicable child gates. Release-level HUMAN/provider/production decisions remain with their owning issues.

Explicit non-goals

  • No broad #4 implementation, combined source sync, cross-source atomic snapshot, or speculative source enablement.
  • No source repository write, website-created commit/branch/pull request, moving-branch runtime fetch, direct provider fallback, or production/protected-data operation.
  • No revival, rename, or emulation of staged ContentRelease candidate/ready/activate/rollback as source authority; no automatic rollback or arbitrary older-SHA sync. See #38/#278 for the accepted direct-sync and historical-contract boundaries.
  • No duplication of #38's direct-sync consumer manifest, #44's search/graph projection contract, or #72/#76/#77 release traceability/report contracts.
  • No course/account/event/email/Studio behavior outside the explicitly owned child/domain issue, and no acceptance of a child from a checked projection or synthetic fixture alone.

Dependencies and next action

Closed decisions/foundations (#12, #24, #37, #103, #156, and #226) establish direction and reusable seams. Current blockers are the open/re-grooming source/projection baseline #253, historical migration classification #219, the #38 source/authority manifests and direct-sync child sequence, and the accepted inputs for #39–#44. The orchestrator should continue those bounded lanes independently; no #4 child may be accepted or merged from stale ContentRelease wording, old source counts, or the failed 9a491cd regression evidence.

This parent issue does not edit the normative specifications. Their staged-content reconciliation is explicitly owned by #278; until that contract is accepted, the current specs are read together with the closed #226 decision and the current #38/#278 issue bodies.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

There is no single implementation entry point: start with _docs/PROCESS.md and the linked architecture, GitHub content, API, migration, and verification specifications, then review child issues #38–#44. This epic is complete only when its source manifests, child adapters and projections, direct-sync lifecycle, search/graph contracts, and verification gates all pass; implementation work belongs in the child issues.

Written by the indexing model from the issue text.

Assessment

Tech stack
django, github, python
Domain
api, backend, database, search
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.