opensanctions / opensanctions/poliloom

Agentic entity mapping: give the LLM a search tool instead of dumping 100 candidates

Open
#149 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

loom
Dominant language
Python
Stars
22
Forks
2
PR merge metrics
No merged PRs in 30d

Description

Problem

The current two-stage entity mapping pipeline (positions, birthplaces, citizenships) searches Meilisearch for 100 candidates and dumps them all into the LLM context for selection. This has issues:

  • High noise: Semantic ratio was lowered from 0.8 to 0.5 because of irrelevant results, but this also reduces cross-lingual recall
  • Single-shot search: The query can't adapt when results are poor — no way to disambiguate (e.g. narrowing "Minister of Defence" with a country) or reformulate when the correct entity has a slightly different name
  • No retry path: If the correct entity isn't in the top 100, the mapping silently fails (returns None)
  • Expensive: 100 candidate descriptions (5000-8000 tokens) are fed into every mapping call

Proposal

Replace the "search then select" pattern with an agentic approach: give the mapping LLM a search tool and let it formulate its own queries.

Search tool interface:

  • query (string): The search text, formulated by the LLM
  • semantic_ratio (float): Balance between keyword and semantic search
  • Returns ~25 candidates per call

What this enables:

  • LLM can call the search tool multiple times with different queries — reformulating, narrowing, or broadening as needed
  • LLM adds jurisdiction context (e.g. appends country name) to disambiguate
  • LLM chooses semantic ratio based on the situation (high for cross-lingual, low for exact labels)
  • The number of search iterations is up to the LLM based on the difficulty of the mapping

Cost

With cached input tokens at 90% discount, multiple tool calls are cheap — each subsequent call only pays full price for the new search results while the rest of the context is cached. Even with 3-4 search rounds, this is likely comparable to or cheaper than the current single call with 100 candidates.

Scope

This applies to all three two-stage extraction types:

  • Position mapping (POSITIONS_CONFIG)
  • Birthplace mapping (BIRTHPLACES_CONFIG)
  • Citizenship mapping (CITIZENSHIPS_CONFIG)

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Locate the current two-stage entity-mapping pipeline and the POSITIONS_CONFIG, BIRTHPLACES_CONFIG, and CITIZENSHIPS_CONFIG definitions. Start by tracing how each config performs candidate search and passes results to the LLM. Done means all three mappings use the proposed search tool workflow, with reformulation and multiple searches available instead of dumping 100 candidates.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, search
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.