forrtproject / forrtproject/flora-extractor
Stage 3 audit 2026-08-13: partial results coded successful, self-record links, quote-source provenance
- Dominant language
- Python
- Stars
- 2
- Forks
- 1
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 4
Description
A hand audit of 30 random `successful` and 15 random `failed` rows in the 8b3d01ec5ad7 render (seed 42) found no failure coded as success and clean evidence on the failed side, but three defects:
1. **Partial→successful drift (the substantive one).** 6 of 30 sampled successes are arguably `mixed`: the coder takes the authors' summary sentence at face value and never weighs a discordant clause later in the same abstract. Recurring shapes: a headline effect replicated while a secondary result ran opposite to the original (10.31234/osf.io/mhpu9), and partial coverage framed as success (11 of 17 proteins, 10.1016/j.bpsgos.2024.100332). Upper bound if all such rows re-coded: successful drops from 57% to ~46% of settled rows. Candidate fix: prompt the outcome coder to state whether any reported result contradicts the original before settling on `successful`.
2. **Self-record link.** 10.31234/osf.io/bqzj9_v1: `title_o` is the replication's own record (a version/duplicate of itself), not the original (Triplett 1898). A same-title, same-author-list candidate should be refused at link time.
3. **Quote-source provenance.** 10.17605/osf.io/5bms8: `out_quote_source = abstract` but the row's stored abstract is an OSF registration stub that contains none of the quoted text — the quote came from a document the row does not show.
Full audit table in the 2026-08-13 handover round.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the 2026-08-13 handover round and its full audit table, using the 8b3d01ec5ad7 render with seed 42 to reproduce the three findings. Trace the outcome-coder prompt, same-title/same-author link check, and quote-source field; done means mixed outcomes are not settled as successful, self-records are rejected, and every quote source is represented by the row.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend, data
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 48/100