Jordan-Hall / Jordan-Hall/browser

[P1][RES-01] Deep retrieval and exact-link discovery

Open
#75 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
No language data
Stars
0
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Programme: #1
Epic: #25

## Objective
Build a research retrieval pipeline that optimizes for useful exact sources and primary evidence rather than a traditional ranked page of links.

## Scope
- Query decomposition from GoalContract constraints.
- Source planning across installed connectors, public web, domain APIs and user-authorized sources.
- Canonical URL/object resolution and exact-link validation.
- Retrieval logs including query/source/time/result disposition.
- Primary-source preference and independent-source grouping.
- Record inaccessible/failed sources as explicit evidence gaps.
- Product/entity identifier extraction for downstream entity resolution.
- Cache only within source-specific access/retention rules.

## Research rules
- Search snippets are discovery hints, not sufficient evidence for material facts.
- Ten syndicated copies of one story count as one underlying information source where detected.
- Retrieval breadth does not override privacy/account permissions.

## Acceptance criteria
- [ ] Material results include accessible canonical source links/object IDs where available.
- [ ] Important facts are grounded in underlying evidence, not snippets alone.
- [ ] Primary sources are preferred/explained when available.
- [ ] Inaccessible sources remain explicit gaps rather than hallucinated contents.
- [ ] Duplicate/syndicated source groups are tracked.
- [ ] Research fixture scores exact-link validity, constraint coverage, source quality and missing-source honesty.

## Dependencies
- CONN-01
- DATA-01

**First phase:** P1
**Maturity target:** P2
**Owner:** connectors-domains

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading the CONN-01 and DATA-01 dependencies, then map the requested retrieval pipeline to query decomposition, source planning, URL/object resolution, logging, and evidence-gap handling. Use the research fixture to check exact-link validity, constraint coverage, source quality, missing-source honesty, and duplicate-source grouping. Done means the listed acceptance criteria are demonstrably satisfied.

Written by the indexing model from the issue text.

Assessment

Domain
data, search
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.