internetarchive / internetarchive/openlibrary
2026 H2 Planning: Engineering Initiatives Recap
- Dominant language
- Python
- Stars
- 6.7k
- Forks
- 2k
- Avg merge
- 2d 14h
- Merged PRs (30d)
- 126
Description
## Summary
**This is a planning/tracking issue, not an implementation spec.** It recaps major 2026 H2 engineering initiatives so each one can get a clear path to its own Epic issue and a slot on a 2026 H2 project board. Nothing here is ready to be picked up as-is — several items still need internal discussion before they're scoped enough to become a real epic. **Close this issue once every item below has a real Epic issue and the project board exists** — it's meant to be short-lived.
Some items already map to existing epics; some need discussion before they're scoped enough to become one. Top-level items are numbered sequentially; **Browsing** (4) and **Participation** (5) are umbrella categories with lettered sub-items, since each pair is the same underlying activity at different granularity.
## Checklist
- [ ] **1. Imports** (@rattnak) — broadly 3 things: (a) a registry of partner feeds we can crawl over time, plus making it easier to register new identifier types; (b) a database of acquisition options per edition (prices, download, etc — whether this eventually goes into Solr is still TBD); (c) an insulated batch import system where anyone with an account can stage records for import (including covers) without overloading production DB/website, supporting patrons, partners, Amazon, etc. Needs further scoping/discussion before this is final.
- #12844 — Feed Registry & Acquisitions Import System (Epic) — the integrated end-state for (a) and (b): feed registry + acquisitions DB + BWB ingestion + query-time Solr surfacing. Reopened and updated 2026-07-23 with real status (acquisitions table merged but not wired; feed registry PR unmerged/conflicting; BWB script drafted, not integrated; Solr surfacing not started). Still needs scope reconciled against #5792 and #12655 below.
- #12655 — BookWorm: Modernize the Import Pipeline (candidate existing epic for (c), the general import queue/DB modernization — distinct from but adjacent to #12844)
- [ ] **2. Metrics [Core Vitals]** — needs to land this cycle, ASAP. **2026 priority is sequential, not parallel: Retention → Participation.** Content Value is measured this year (a diagnostic — identifying high-demand/low-usefulness records) but *actively growing* content value is a 2027 goal, not 2026's. Open question: does this need one umbrella epic tying the three sub-scores together, or do they stay independent? The existing issue(s) below need expanding.
- #11954 — Usefulness & Demand Scores
- #11955 — Participation Score
- #11956 — Retention Score
- [ ] **3. Performance** — Solr, ModSecurity, API isolation, import insulation. No single epic currently covers this combination — likely needs a new umbrella epic linking the pieces.
- Solr: general visibility into why/how Solr gets saturated — which queries are killing it, when, where, why. Very little observability into Solr internals today.
- ModSecurity: existing PRs need discussion + review, both for performance and to unblock false positives. Separately, jimchamp is building a daily pipeline — an anonymized ModSecurity error log uploaded to an archive.org item each day, enabling repeatable rounds of audits.
- API isolation: isolate API usage to specific ol-web heads (haproxy or similar) so heavy API traffic can't take down the production website.
- Import insulation: see item 1c.
- #12718 — ModSecurity: 403 Forbidden (+ 460 sub-thread)
- #7726 — Performance: Borrowing Bloat
- [ ] **4. Browsing** — carousels and tag/genre-driven browsing are the same underlying patron activity (finding books by facet) at different scope; grouped here so the two don't diverge on taxonomy/UX independently.
- [ ] **4a. Carousels++** (@lokesh, @BearSunny) — filter/facet carousels by tag, order carousel results by Dewey Decimal in Solr. Nuanced addition: **carousel controls** — a dropdown/option to add "topics" (tags, genres, content warnings, etc.) as a facet, plus an option to change sort order (e.g. Trending, DDC). Separately, **content-warning moderation via `content_warning:*` tags should be an option at the global/page level**, not just per-carousel — a site-wide patron preference, not a one-off carousel toggle. @lokesh is redesigning the carousel away from Slick; @BearSunny is assigned to #10512 and open to help prototyping.
- #9828 — Carousel Fixups (candidate existing epic)
- #10512 — facet/filter carousels by tag
- #13159 — order carousel results by Dewey Decimal in Solr
- [ ] **4b. Tags & Genre Explorer** (@Chisomnwa) — the broader Tags initiative Chisom is co-leading with Mek, plus the Genre Explorer feature (genre/subgenre bookshelf browsing backed by the Tag hierarchy) that sits directly on top of it.
- #13158 — Genre Explorer (prototype already underway)
- #11610 — RFC: genres field on Work records
- [ ] **5. Participation** — patron-driven contribution and community signal; Prompts, Lists, and Micro-edits are all ways patrons create value for other patrons, grouped here for the same reason as Browsing above.
- [ ] **5a. Micro-edits** (Ore James, @lokesh) — move away from the monolithic, confusing book-edit UI (separate work-level/edition-level metadata tabs) toward something closer to Wikimedia's [Micro-task Generator for Organizers](https://meta.wikimedia.org/wiki/Research:Micro-task_Generator_for_Organizers_on_Wikipedia) — e.g. click a title, edit it in place. Falls under librarian editing. Related to the Data Quality Table (#7661/#10764, identifies records needing metadata improvement based on search query) as a source of what to surface for micro-editing. Now that jimchamp's `transaction_details` table work (#12319, #12434) exists, we could synthesize commit messages on the fly instead of requiring editors to type one manually — directly relevant to #2386 (auto-generate comments for edits with missing comments).
- #2386 — Auto-generate comments for edits w/ missing comments
- #7661 / #10764 — Data Quality Table (work search & author pages)
- [ ] **5b. Prompts** (Lauren) — Weekly curated homepage prompt widget, staff-curated with patron nomination/voting. Already well-scoped (MVP Lit component, data model, concrete success criteria) — closest of any item here to being genuinely ready to build. Expanded scope: also **Lists** — boosting signal for good/quality patron-created Lists as a participation mechanism, alongside Prompts.
- #12858 — Weekly Prompt homepage widget
- [ ] **6. Solr improvements** (@cdrini) — real-time loan availability in Solr (work already started on a branch), acquisition/price data, and hopefully also cleaning up our legacy loans infrastructure.
- #7450 — Book Availability in Solr (@benbdeitch)
- #12138 — Solr Bulk Indexing BWB & Lenny OPDS (price)
- [ ] **7. Ada & AI automations** — finish up our AI automations and workflows to make project management reasonable. Open question whether the existing epic also covers the broader agent-based dev workflow, or whether that stays a separate effort.
- #13160 — Epic: AI Workflows — GitHub Actions automation
- [ ] **8. Lenny** (@ronibhakta1) — target: October. Stand up a lennyforlibraries.org instance (one already exists) loaded with all the EPUBs archive.org has purchased from BRIET, connecting each edition via **1a** (the partner feed registry). End goal: search Open Library for books available via the "Archive Labs Lenny Instance," with Lenny as a borrow option (for archive.org account holders) on equal footing with any other trusted book provider. Neither candidate epic below mentions Lenny by name yet.
- #5792 — Trusted Book Providers
- #10251 — Simplify Trusted Book Providers integration
- [ ] **9. Retention** (@Sadashii) — Activity Feed as a discovery/retention surface. Same umbrella question as item 2.
- #11956 — Retention Score
- #10242 — Add Social Activity Feed to My Books page
## Next steps
- For each checked-off item: confirm or create its Epic issue, then add it to the 2026 H2 project board.
- Item 5a (Micro-edits, and parts of item 1) need more internal discussion/scoping before they're ready for their own epic.
- **Fix Acquisition** (preserving registration intent / letting patrons engage before signing up) isn't yet its own item above — #10335 already has a concrete, community-proposed design for exactly this: pre-registration persona + subject/mood chip selection stored client-side, homepage personalizing immediately with no signup required, preferences auto-migrating to `/account/subscriptions` once the patron does register. Worth a look before deciding whether this becomes its own top-level item.
- Close this issue once the board exists and every item above has a linked Epic.
Contributor guide
Assessment
This issue has not been assessed yet.