agentscope-ai / agentscope-ai/ReMe
benchmark: complete LongMemEval evaluation
- Ngôn ngữ chính
- Python
- Star
- 3.4k
- Fork
- 298
- Merge trung bình
- 19 giờ 52 phút
- Pull request đã merge (30 ngày)
- 55
Mô tả
**Background**
[LongMemEval](https://github.com/xiaowu0162/LongMemEval) evaluates information extraction, multi-session reasoning, knowledge updates, temporal reasoning, and abstention over long, timestamped chat histories. ReMe already contains initial LongMemEval steps and runner scripts, but the workflow should be completed and standardized as a reproducible benchmark.
**Changes**
- Consolidate the existing LongMemEval steps and scripts into a documented end-to-end workflow under `benchmark/longmemeval`
- Support the official cleaned LongMemEval-S, LongMemEval-M, and oracle datasets without committing downloaded data or generated workspaces
- Preserve chronological session ingestion and source identifiers for evidence-level evaluation
- Provide direct-context, retrieval, and agentic-answer baselines using consistent prompts and model settings
- Evaluate the five official memory abilities and report both aggregate and per-category results
- Separate retrieval recall from answer correctness and retain inspectable evidence and judge outputs
- Record dataset revision, configuration, model, embedding, latency, token usage, and cost metadata
- Add resumable runs, deterministic sample subsets, cleanup guidance, and focused tests with fixture data
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Đánh giá
Issue này chưa được đánh giá.