agentscope-ai / agentscope-ai/ReMe
benchmark: evaluate personal knowledge and proactive memory
- Langage dominant
- Python
- Étoiles
- 3.4k
- Forks
- 298
- Merge moyen
- 19 h 52 min
- PR mergées (30 j)
- 55
Description
**Background**
ReMe supports personal, file-native knowledge and exposes proactive topics generated from accumulated memory, but conventional long-term memory benchmarks mainly test explicit recall after a user asks a question. We need to identify and adapt benchmarks that also measure user modeling, implicit memory retrieval, appropriate personalization, proactive timing, usefulness, and the cost of unnecessary or intrusive interventions.
**Changes**
- Survey public benchmarks and document their licenses, data availability, task definitions, metrics, reproducibility, and fit with ReMe's local-first model
- Define separate evaluation dimensions for personal knowledge capture, retrieval, freshness and updates, provenance, personalization, and proactive behavior
- Include negative cases where memory should not be retrieved or surfaced, measuring irrelevance, repetition, privacy risk, over-personalization, and unnecessary interruption
- Define proactive metrics for trigger detection, timing, evidence grounding, actionability, user value, and false-positive rate
- Select one primary public benchmark and a minimal ReMe-specific extension only where existing datasets leave material coverage gaps
- Design a reproducible adapter and result schema that separates memory-system quality from host-agent policy and generation quality
- Produce a recommendation, implementation plan, small validation subset, and documented baseline before adding a full benchmark harness
Guide de contribution
Aucun guide de contribution indexé pour ce dépôt
Évaluation
Cette issue n'a pas encore été évaluée.