CodeForPhilly / CodeForPhilly/codeforphilly-ng

Post-cutover: migrate body-heavy entities to gitsheets v1.2 content-typed (markdown) records

Open
#44 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
TypeScript
Stars
1
Forks
1
Avg merge
5d 3h
Merged PRs (30d)
9

Description

[gitsheets v1.2.0](https://github.com/JarvusInnovations/gitsheets/releases/tag/v1.2.0) added content-typed records — sheets opt into `format.type = 'markdown'` to store records as `.md` files with TOML frontmatter and a designated body field. Plus lazy body loading via `query({ withBody: false })`.

This is the biggest one-time upgrade we'd take from the gitsheets 1.x line. Not urgent — defer to **after `cutover-prep` ships** so we're not refactoring entities mid-migration.

## Why migrate

- **Snapshot is actually-readable.** Contributors cloning `codeforphilly-data-snapshot` see real `.md` files in any markdown viewer — instead of parsing TOML records to find the prose.
- **Authoring via PR.** Staff / maintainers can edit a project overview in any markdown editor and PR it; currently they roundtrip through the API.
- **Listing performance.** `queryAll({ withBody: false })` on hot paths: projects-index, activity feed, FTS seeding, snapshot scrub.
- **Indexes stay fast.** Index builds use body-less reads natively in v1.2.

## What changes

Entities with substantial body content:
- `Project` — `overview` (markdown body), `summary` (short markdown) → migrate `overview` as the body field, keep `summary` in frontmatter
- `ProjectUpdate` — `body` (markdown) → migrate `body` as the body field
- `ProjectBuzz` — `summary` (markdown) → migrate as body
- `Person` — `bio` (markdown) → migrate as body
- `HelpWantedRole` — `description` (markdown) → migrate as body
- `Tag` — `description` (markdown, short) → optional; cheaper to leave as TOML field

The migration is bounded; entities without long bodies (`ProjectMembership`, `SlugHistory`, `Revocation`, `TagAssignment`, `HelpWantedInterestExpression`) stay as TOML records.

## Tasks

1. Schema reshape in `packages/shared/src/schemas/` — one designated body field per content-typed entity (rename or restructure the existing `overview` / `body` / `bio` / `description` / `summary` fields).
2. Update `.gitsheets/.toml` configs with `[gitsheet.format] type = 'markdown' body = ''`.
3. In-memory loader in `apps/api/src/store/memory/loader.ts` — use `{ withBody: false }` for index-building reads; lazy-load via `Sheet.loadBody(record)` when serving record detail responses.
4. Serializers in `apps/api/src/services/serializers/` — `*Html` / `*Excerpt` derived from the body field instead of the legacy string field.
5. FTS pipeline in `apps/api/src/store/fts.ts` — body included in the indexed text via lazy-load batch.
6. `apps/api/scripts/import-laddr.ts` — write the new markdown format for migrated entities.
7. `apps/api/scripts/scrub-data.ts` — the snapshot now contains real `.md` files; verify the scrub still strips PII correctly across the new file shape.
8. The data repo's existing TOML records need migration once — write a one-shot `apps/api/scripts/migrate-to-content-typed.ts` that reads existing records and rewrites as `.md` per the new format.
9. Update `specs/behaviors/markdown-rendering.md` and `specs/data-model.md` to reflect content-typed entities.

## Why defer

- `cutover-prep` is next and depends on every other plan; this would invalidate frozen plans (`storage-foundation`, `read-api`, `write-api`, `laddr-import`, `public-snapshot-scrub`).
- The benefit is real but landing is post-cutover work, not pre-cutover refactor.

## Out of scope

- `gitsheets check` pre-commit hooks belong in the data repo, not this code repo.
- The bundled Claude Code skill at `node_modules/gitsheets/skills/gitsheets/` is available once we bump the dep range; future plans touching gitsheets can load it.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.