filecoin-project / filecoin-project/devgrants
Open Grant Proposal: Decentralized Scholarly Archive on Filecoin (QDC)
- Dominant language
- No language data
- Stars
- 409
- Forks
- 311
- PR merge metrics
- No merged PRs in 30d
Description
# Open Grant Proposal: Decentralized Scholarly Archive on Filecoin — Provenance, Retrieval, and a Reference Architecture for AI-Assisted Research Corpora
**Project Name:** QNFO Decentralized Corpus (QDC)
**Proposal Category:** Research & protocols
**Individual or Entity Name:** Individual — Rowan Brad Quni-Gudzinas (ORCID 0009-0002-4317-5604)
**Proposer:** rwnq8
**Project Repo(s):**
- https://github.com/QNFO/qnfo-workers (platform, paper pipeline, gateway)
- https://github.com/QNFO/QWAV (strategy)
**(Optional) Filecoin ecosystem affiliations:** None. This is a new contribution to the ecosystem.
**(Optional) Technical Sponsor:** None yet — open to connecting with the Filecoin team.
**Do you agree to open source all work you do on behalf of this RFP under the MIT/Apache-2 dual-license?:** Yes
# Project Summary
AI-assisted research is producing open corpora at a scale that outpaces the trust and persistence infrastructure of the current web. The QNFO platform alone has published approximately 1,000 method papers (all DOI-registered, audited, and served from a living database), yet they live on centralized infrastructure with no content-addressed, verifiable, long-term persistence story. When a research corpus is decentralized and content-addressed, citation integrity becomes cryptographically checkable, provenance becomes portable, and the corpus cannot be silently altered or lost. This proposal builds that layer for a real, operating corpus — and publishes the reference architecture so any research organization can replicate it.
The deliverable is the QNFO Decentralized Corpus (QDC): a pipeline that (1) mirrors the full QNFO paper corpus (markdown, PDF, HTML, plus provenance manifest) onto Filecoin storage deals, (2) publishes content-addressed citation records so every paper's CID is derivable from its DOI, and (3) benchmarks retrieval (CID-to-content, deal-verification, retrieval speed) against the centralized baseline. This is not an IPFS-only project — it uses Filecoin storage deals, deal verification, and retrieval as the persistence and audit layer, with IPFS as the transport.
## Impact
Scholarly publishing has a persistence problem: content lives on corporate clouds, links rot, and provenance is unverifiable. Decentralized storage solves this only if someone actually demonstrates the full loop — corpus → content addressing → storage deals → verifiable retrieval → public reference architecture. This project does that with a live corpus and publishes the benchmarks, which gives the Filecoin ecosystem a credible, reproducible case study for the research-publishing vertical — a large, growing, and currently underserved market for long-term data persistence.
Risks of not doing this: the research-publishing vertical defaults to centralized archives (or closed systems), and the Filecoin network misses a natural, high-retention use case. Academic data is exactly the "humanity's information" category the Filecoin mission targets — data that must remain available and verifiable for decades. Success looks like: 1,000+ papers content-addressed and stored on Filecoin, retrieval benchmarks published, and the reference architecture adopted by at least one other research organization within 12 months.
## Outcomes
**Milestone 1 — Provenance & content addressing.** Build the corpus manifest generator: every paper (markdown, PDF, HTML) hashed to a CID; a provenance manifest mapping DOI → slug → CIDs → license → audit trail; verification tool that recomputes CIDs from the live corpus and reports drift. Deliverables: `qdc-manifest` tool, published provenance manifest, verification script.
**Milestone 2 — Filecoin storage deals.** Onboard the corpus via Filecoin storage: deal-making for the corpus blocks (target: the full corpus, ~1-5 GiB), verified deals with a public deal browser record, and an open dashboard showing deal state, term, and renewal schedule. Deliverables: deal-making scripts, public deal records, storage dashboard.
**Milestone 3 — Retrieval benchmarks & reference architecture.** Publish retrieval benchmarks (CID-to-content latency, deal verification, retrieval success rate over time, cost per GiB) against the centralized baseline, and write the reference architecture document (how any research org can replicate the pipeline). Deliverables: benchmark report, reference architecture doc, public demo.
## Data Onboarding
- Month #1: 250 papers (manifest + CIDs)
- Month #3: full corpus (~1,000 papers) content-addressed
- Month #6: full corpus on storage deals + retrieval verified
- Month #12: maintenance, renewal, benchmarks published
## Adoption, Reach, and Growth Strategies
Target audience: (1) open-science and AI-assisted research organizations publishing large corpora; (2) the Filecoin developer ecosystem seeking reference implementations; (3) academic libraries evaluating decentralized preservation. QNFO already operates a public corpus (papers.qnfo.org), a knowledge graph, and a public GitHub org — the first 10 "users" are the QNFO publication pipeline itself (every future paper automatically enters the decentralized archive); the first 100 are reached through the open reference architecture, the benchmark report, and presentation of the case study in decentralized-storage and open-science communities.
## Development Roadmap
**Milestone 1 — Provenance & content addressing (completed ~45 days from grant).** Corpus manifest generator, CID verification tool, published provenance manifest. Funding: $8,000. Single developer (proposer) with full-stack Cloudflare/Workers/D1/R2/Vectorize and IPFS deployment experience (prior IPFS deployment of the QNFO platform, 2026-07-18).
**Milestone 2 — Filecoin storage deals (~90 days).** Deal-making pipeline, public deal records, storage dashboard. Funding: $10,000. Same developer; engagement with Filecoin docs/community for deal mechanics.
**Milestone 3 — Retrieval benchmarks & reference architecture (~120 days).** Benchmark report, reference architecture doc, public demo, adoption outreach. Funding: $7,000. Total: $25,000.
## Total Budget Requested
| Milestone # | Description | Deliverables | Completion Date | Funding |
|===|===|===|===|===|
| 1 | Provenance & content addressing | Manifest generator, CID verification, published manifest | 2026-11-15 | $8,000 |
| 2 | Filecoin storage deals | Deal pipeline, public deal records, dashboard | 2027-01-15 | $10,000 |
| 3 | Retrieval benchmarks & reference architecture | Benchmark report, architecture doc, demo | 2027-02-15 | $7,000 |
## Maintenance and Upgrade Plans
The corpus is continuously growing (QNFO publishes new papers continuously); the pipeline is designed as a repeatable service — new papers automatically enter the decentralized archive. The dashboard and verification tool run as open-source services; deal renewal is monitored and documented. The reference architecture is maintained as a living document in the project repo.
# Team
## Team Members
- Rowan Brad Quni-Gudzinas (solo) — 20+ years software/systems engineering, quantum/AI research, platform architect; founder of QNFO (open AI-assisted research platform) and QWAV.
## Team Member LinkedIn Profiles
- https://www.linkedin.com/in/rowanquni (public profile)
## Team Website
- https://qnfo.org · https://qwav.tech
## Relevant Experience
The proposer has built and operated the entire QNFO platform: Cloudflare Workers/D1/R2/Vectorize stack, a knowledge graph with 8,000+ nodes, a publication pipeline that has deposited 900+ records to Zenodo with live DOIs, and a prior IPFS/decentralized deployment of the platform (2026-07-18). Published method papers on audit methodology (10.5281/zenodo.21901984, 10.5281/zenodo.21901983) and funding strategy (10.5281/zenodo.21922589) demonstrate the discipline for provenance and verification work. The corpus being archived is real, live, and already DOI-registered — this is not a greenfield demo.
## Team code repositories
- https://github.com/QNFO/qnfo-workers (publication pipeline, gateway, paper server)
- https://github.com/QNFO/QWAV (research strategy)
# Additional Information
- Learned about the Open Grants Program via the Filecoin Foundation grants page (fil.org/grants) while researching decentralized-storage funding.
- Best email for grant agreement and next steps: rowan.quni@qwav.tech
- The QNFO corpus is a live, operating asset: papers.qnfo.org serves ~1,000 DOI-registered papers; the knowledge graph is queryable; all evidence links are public. The proposer is an individual applicant — the agreement and payments would be completed as an individual (per the template note).
Contributor guide
No contributing guide indexed for this repository
Research direction
The proposal points to the qnfo-workers and QWAV repositories; start by reviewing their publication pipeline, gateway, and strategy materials, then examine the three milestone deliverables. Done would mean the proposed manifest and verification tool, Filecoin storage and dashboard work, and retrieval benchmarks and reference architecture are delivered as described.
Written by the indexing model from the issue text.
Assessment
- Domain
- data-engineering, distributed-systems, infrastructure
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100