mem0ai / mem0ai/memory-benchmarks
Pre-publication review: our re-scoring of Mem0's published LoCoMo answers under alternative judge prompts (corrections welcome)
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 112
- Forks
- 46
- PR merge metrics
- No merged PRs in 30d
Description
Dear Mem0 maintainers,
We are the authors of "The Judge Is the Benchmark: The Same Answers Score 91% or 35%, Depending on the Grading Prompt", a measurement paper we are posting to arXiv. Full disclosure first: the lead author is the founder of Mnemoverse.AI, a vendor of a competing agent-memory system. The paper discloses that conflict and is built so that its load-bearing result runs entirely on your published artifacts, where no system of ours can be advantaged.
The preprint PDF is attached to the repository release: https://github.com/mnemoverse/mnemoverse-benchmarks-paper/releases/tag/arxiv-v1 ; the arXiv identifier will be added to this issue once announced.
What the paper does with your artifacts (all from `mem0ai/memory-benchmarks@4b61c5d`, Apache-2.0):
- It re-scores the 1,539 scorable LoCoMo-10 answers you released, under your own distributed judge prompt and under a deliberately strict prompt of ours. Under your prompt the answers return 91.0%, within 1.5 points of the 92.5% you publish; under the strict prompt the identical answers return 35.0%. The paper reads the 91.0% as the field's realized number and the 56-point span as a constructed stress test, not as a claim about memory quality.
- It scores the same answers with the unmodified LongMemEval rubric on its pinned `gpt-4o-2024-08-06` judge: 81.7%, against 91.9% for your prompt on a `gpt-4o` backbone. The paper calls this a configuration contrast, because the two judges may resolve to different snapshots.
- It scores the same answers with a judge calibrated on blind human adjudication: 85.3% against your prompt's 91.0% on the 1,534 jointly scored answers, a provisional 5.7-point estimate with the stated caveats (calibration uncertainty excluded from the interval; the lead author is the primary annotator).
- The two prompts disagree on 863 answers, all in one direction (your prompt credits, the strict one rejects); the paper says explicitly that 855 of those were not hand-adjudicated and warns against reading un-adjudicated disagreement as real over-credit.
- The Acknowledgments credit the openness that makes the audit possible: you published the answers, the prompt and the harness, and this paper depends on that.
Three questions:
1. Is any reading above factually wrong, in particular our port of your judge prompt and our description of your harness? If so, what is the correct reading, with a pointer (file, line, commit or doc page)?
2. Which judge model, and which snapshot or alias, produced the published 92.5%, and on which date? The paper notes that the release does not pin the exact configuration behind the official number; your answer would settle it.
3. May we record your response verbatim in the paper's public repository (https://github.com/mnemoverse/mnemoverse-benchmarks-paper, release arxiv-v1, DOI https://doi.org/10.5281/zenodo.22348527) and acknowledge it in the next arXiv version?
Every number in the paper re-derives from committed artifacts in that repository (`REPRODUCING.md` is the claim-to-command registry). Corrections received will enter the next arXiv version (v2) and the repository as dated entries.
With thanks for your time,
Edward Izgorodin, Olga Timoshina, Andrey Ustyuzhanin
edward@mnemoverse.ai
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the artifacts at mem0ai/memory-benchmarks@4b61c5d and the cited REPRODUCING.md claim-to-command registry, then compare the paper's judge-prompt and harness descriptions with the repository. Determine the model and snapshot behind the published 92.5% result if the repository records them, and report factual corrections with file, line, commit, or documentation pointers.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100