forrtproject / forrtproject/flora-extractor

Stage 3 audit 2026-08-13: partial results coded successful, self-record links, quote-source provenance

Open
#198 0 comments 0 reactions 0 assignees View on GitHub
stage-3
Dominant language
Python
Stars
2
Forks
1
Avg merge
1d 6h
Merged PRs (30d)
4

Description

A hand audit of 30 random `successful` and 15 random `failed` rows in the 8b3d01ec5ad7 render (seed 42) found no failure coded as success and clean evidence on the failed side, but three defects:

1. **Partial→successful drift (the substantive one).** 6 of 30 sampled successes are arguably `mixed`: the coder takes the authors' summary sentence at face value and never weighs a discordant clause later in the same abstract. Recurring shapes: a headline effect replicated while a secondary result ran opposite to the original (10.31234/osf.io/mhpu9), and partial coverage framed as success (11 of 17 proteins, 10.1016/j.bpsgos.2024.100332). Upper bound if all such rows re-coded: successful drops from 57% to ~46% of settled rows. Candidate fix: prompt the outcome coder to state whether any reported result contradicts the original before settling on `successful`.

2. **Self-record link.** 10.31234/osf.io/bqzj9_v1: `title_o` is the replication's own record (a version/duplicate of itself), not the original (Triplett 1898). A same-title, same-author-list candidate should be refused at link time.

3. **Quote-source provenance.** 10.17605/osf.io/5bms8: `out_quote_source = abstract` but the row's stored abstract is an OSF registration stub that contains none of the quoted text — the quote came from a document the row does not show.

Full audit table in the 2026-08-13 handover round.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the 2026-08-13 handover round and its full audit table, using the 8b3d01ec5ad7 render with seed 42 to reproduce the three findings. Trace the outcome-coder prompt, same-title/same-author link check, and quote-source field; done means mixed outcomes are not settled as successful, self-records are rejected, and every quote source is represented by the row.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend, data
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.