mem0ai / mem0ai/memory-benchmarks

Pre-publication review: our re-scoring of Mem0's published LoCoMo answers under alternative judge prompts (corrections welcome)

Open
#29 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
112
Forks
46
PR merge metrics
No merged PRs in 30d

Description

Dear Mem0 maintainers,

We are the authors of "The Judge Is the Benchmark: The Same Answers Score 91% or 35%, Depending on the Grading Prompt", a measurement paper we are posting to arXiv. Full disclosure first: the lead author is the founder of Mnemoverse.AI, a vendor of a competing agent-memory system. The paper discloses that conflict and is built so that its load-bearing result runs entirely on your published artifacts, where no system of ours can be advantaged.

The preprint PDF is attached to the repository release: https://github.com/mnemoverse/mnemoverse-benchmarks-paper/releases/tag/arxiv-v1 ; the arXiv identifier will be added to this issue once announced.

What the paper does with your artifacts (all from `mem0ai/memory-benchmarks@4b61c5d`, Apache-2.0):

- It re-scores the 1,539 scorable LoCoMo-10 answers you released, under your own distributed judge prompt and under a deliberately strict prompt of ours. Under your prompt the answers return 91.0%, within 1.5 points of the 92.5% you publish; under the strict prompt the identical answers return 35.0%. The paper reads the 91.0% as the field's realized number and the 56-point span as a constructed stress test, not as a claim about memory quality.
- It scores the same answers with the unmodified LongMemEval rubric on its pinned `gpt-4o-2024-08-06` judge: 81.7%, against 91.9% for your prompt on a `gpt-4o` backbone. The paper calls this a configuration contrast, because the two judges may resolve to different snapshots.
- It scores the same answers with a judge calibrated on blind human adjudication: 85.3% against your prompt's 91.0% on the 1,534 jointly scored answers, a provisional 5.7-point estimate with the stated caveats (calibration uncertainty excluded from the interval; the lead author is the primary annotator).
- The two prompts disagree on 863 answers, all in one direction (your prompt credits, the strict one rejects); the paper says explicitly that 855 of those were not hand-adjudicated and warns against reading un-adjudicated disagreement as real over-credit.
- The Acknowledgments credit the openness that makes the audit possible: you published the answers, the prompt and the harness, and this paper depends on that.

Three questions:

1. Is any reading above factually wrong, in particular our port of your judge prompt and our description of your harness? If so, what is the correct reading, with a pointer (file, line, commit or doc page)?
2. Which judge model, and which snapshot or alias, produced the published 92.5%, and on which date? The paper notes that the release does not pin the exact configuration behind the official number; your answer would settle it.
3. May we record your response verbatim in the paper's public repository (https://github.com/mnemoverse/mnemoverse-benchmarks-paper, release arxiv-v1, DOI https://doi.org/10.5281/zenodo.22348527) and acknowledge it in the next arXiv version?

Every number in the paper re-derives from committed artifacts in that repository (`REPRODUCING.md` is the claim-to-command registry). Corrections received will enter the next arXiv version (v2) and the repository as dated entries.

With thanks for your time,

Edward Izgorodin, Olga Timoshina, Andrey Ustyuzhanin
edward@mnemoverse.ai

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the artifacts at mem0ai/memory-benchmarks@4b61c5d and the cited REPRODUCING.md claim-to-command registry, then compare the paper's judge-prompt and harness descriptions with the repository. Determine the model and snapshot behind the published 92.5% result if the repository records them, and report factual corrections with file, line, commit, or documentation pointers.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.