afx-team / afx-team/hebb-mind

[Feature] retrieval: cache IDF corpus stats to cut DB round-trips on repeated queries

Open
#54 1 comment 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
52
Forks
18
PR merge metrics
No merged PRs in 30d

Description

### Problem

Every `search()` call runs `_build_idf` (`src/hebb/retrieval/searcher.py:474`), which hits the DB twice with no caching:

1. **`corpus_size()`** — `SELECT count(*) FROM memory_fts` (`sqlite_store.py:404`, `pg_store.py:237`).

2. **`keyword_doc_freqs()`** — A per-term loop running one `SELECT count(*) ... MATCH ?` per query token (`sqlite_store.py:437`, `pg_store.py:249`).

Neither the retrieval nor storage layer caches anything today.

For a read-heavy workload, this re-computes statistics that rarely change between calls. Two in-tree examples:

- **Identical query re-searched** (UI retry, agent retry): Re-runs the exact same `corpus_size` + per-term Document Frequency (DF) from scratch.

- **Consolidation pass**: In `RecallAgent.recall`, the LLM's first 3 queries are executed in a single pass (`queries[:3]`, `src/hebb/agents/recall_agent.py:49`). The corpus does not change between those 3 calls, yet each re-runs the full DF query work.

### Proposed Solution

Implement a **short-TTL, instance-level cache** inside `MemorySearcher._build_idf` only — with zero storage-layer changes. Use two independent dicts, as the two statistics depend on different inputs:

- **`corpus_size` cache**: Keyed by `tuple(sorted(partition_ids))` $\rightarrow$ `(corpus_size, expires_at)`. Corpus size depends only on the partition set, not the query, so all queries on the same partition share one entry.

- **`keyword_doc_freqs` cache**: Keyed by `(token, tuple(sorted(partition_ids)))` $\rightarrow$ `(df, expires_at)`. DF depends on both the token and partition set, so the token must be in the key to prevent cross-query overwrites.

Both dicts use a short TTL (e.g., 60s). On a cache miss, query the store as before and populate the cache. Cache `None` results as well (empty corpus / no DF-eligible tokens) to prevent repeated DB hits.

#### Out of Scope (For This Issue)

- A global cross-query DF table that collapses all consolidation queries into a single DB hit (higher reward, but requires storage-layer changes and complex invalidation; leave for a follow-up).

- Event-driven invalidation on writes (a short TTL already bounds staleness sufficiently).

### Expected Outcomes & Impact

- Repeated identical `search()` calls within the TTL produce **zero** `corpus_size` / `keyword_doc_freqs` DB queries after the first call.

- `RecallAgent.recall`'s 3-query pass reuses cached DF for overlapping tokens and cached `corpus_size` for the same partition set.

- Partition-scoped DF stays correct (different partition sets produce independent entries; DF is keyed by token, avoiding cross-query overwrites).

- Unit tests will cover: *Cache Hit*, *Cache Miss*, *TTL Expiry*, and *Partition Isolation*.

- All static analysis and test suites must pass green:

Bash

```
ruff check src/
mypy src/hebb/ --strict
pytest tests/ -v
```

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.