agentscope-ai / agentscope-ai/ReMe

benchmark: complete LongMemEval evaluation

Đang mở
#354 0 bình luận 0 reaction 1 người được giao Được @xyf2020 nhận Xem trên GitHub
enhancement
Ngôn ngữ chính
Python
Star
3.4k
Fork
298
Merge trung bình
19 giờ 52 phút
Pull request đã merge (30 ngày)
55

Mô tả

**Background**

[LongMemEval](https://github.com/xiaowu0162/LongMemEval) evaluates information extraction, multi-session reasoning, knowledge updates, temporal reasoning, and abstention over long, timestamped chat histories. ReMe already contains initial LongMemEval steps and runner scripts, but the workflow should be completed and standardized as a reproducible benchmark.

**Changes**
- Consolidate the existing LongMemEval steps and scripts into a documented end-to-end workflow under `benchmark/longmemeval`
- Support the official cleaned LongMemEval-S, LongMemEval-M, and oracle datasets without committing downloaded data or generated workspaces
- Preserve chronological session ingestion and source identifiers for evidence-level evaluation
- Provide direct-context, retrieval, and agentic-answer baselines using consistent prompts and model settings
- Evaluate the five official memory abilities and report both aggregate and per-category results
- Separate retrieval recall from answer correctness and retain inspectable evidence and judge outputs
- Record dataset revision, configuration, model, embedding, latency, token usage, and cost metadata
- Add resumable runs, deterministic sample subsets, cleanup guidance, and focused tests with fixture data

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.