Back to homepage
Memory Benchmarks

Fresh Equal-Conditions Benchmark Results for Metronix

The old benchmark story was flattering but structurally wrong. These results use the updated protocol from the current paper: same models, same question volume, and both retrieval plus end-to-end evaluation before a comparison is allowed on stage.

The Headline

Metronix leads on LoCoMo and MemoryAgentBench under the updated equal-conditions protocol

LoCoMo landed at 52.8% end-to-end accuracy and MemoryAgentBench at 63.6%. LongMemEval-S reached 59.0%, just behind Mem0 at 60.0%. BEAM 100K came in at 32.1%, which is the clearest signal of where the next round of product work should go.

Benchmark
What It Measures
Metronix
Best Peer
Notes
LoCoMo
Long-conversation memory questions
52.8%
Mem0 50.0%
Recall@10: 85.3%
LongMemEval-S
Cross-session assistant memory
59.0%
Mem0 60.0%
Recall@10: 95.4%
MemoryAgentBench
End-to-end agent memory accuracy
63.6%
Mem0 53.0%
Metronix leads
BEAM 100K
Large-ingest memory workload
32.1%
Mem0 42.0%
Directionally competitive
Plain English
  • Metronix leads the equal-conditions comparison on LoCoMo and MemoryAgentBench.
  • LoCoMo and LongMemEval-S show a persistent retrieval-greater-than-generation gap: the system usually finds the evidence before the answer model fully capitalizes on it.
  • The weak spots are clear instead of hidden: conflict resolution, preference following, and large-ingest BEAM workloads still need product work.
Why This Matters

Better memory changes whether agents preserve user preferences, recover the right prior context, and avoid repeating work. The important shift here is not just stronger numbers. It is that the numbers now survive scrutiny instead of depending on benchmark theater.

Methodology

Three Rules Before a Number Gets to Be a Headline

The updated paper formalizes an equal-conditions protocol specifically to avoid apples-to-oranges benchmark comparisons.

Same answer model and same blind judge across systems.
Same question volume and denominator for every system.
Both retrieval and end-to-end evaluation are required before a result is considered comparable.
Cross-System Comparison

Headline Layer B Accuracy Under the Same Harness

Metronix and Mem0 are the two leaders in this round. Metronix tops LoCoMo and MemoryAgentBench; Mem0 tops LongMemEval-S and BEAM 100K.

System
LoCoMo
LongMemEval-S
MAB
BEAM 100K
Metronix
52.8%
59.0%
63.6%
32.1%
Mem0
50.0%
60.0%
53.0%
42.0%
GBrain
18.7%
11.1%
N/A
N/A
MemClaw
14.5%
22.0%
20.0%
LoCoMo Retrieval Detail

Supporting Evidence, Not the Main Headline

The old site made retrieval-only LoCoMo numbers do too much work. They still matter, but now they sit where they belong: as supporting evidence.

Metric
Metronix
MemClaw
Layer B accuracy
52.8%
14.5%
Recall@10
85.3%
14.2%
Precision@1
39.4%
5.2%
nDCG@10
0.603
0.101
Caveats
  • All results are directional, N=1. Useful for strategy and product positioning, not a publication-grade final word.
  • Mem0 exposes headline Layer B numbers, but not the stable provenance needed for full retrieval-layer comparison.
  • BEAM 1M has not been run yet, and GBrain did not scale to the high-volume MAB and BEAM tiers in this test harness.
Current Take

The benchmark story is now stronger because it is more honest. Metronix looks like one of the top memory systems in the field, and the remaining gaps are specific enough to guide the roadmap instead of hiding behind shiny but incomparable numbers.