Fresh Equal-Conditions Benchmark Results for Metronix
The old benchmark story was flattering but structurally wrong. These results use the updated protocol from the current paper: same models, same question volume, and both retrieval plus end-to-end evaluation before a comparison is allowed on stage.
Metronix leads on LoCoMo and MemoryAgentBench under the updated equal-conditions protocol
LoCoMo landed at 52.8% end-to-end accuracy and MemoryAgentBench at 63.6%. LongMemEval-S reached 59.0%, just behind Mem0 at 60.0%. BEAM 100K came in at 32.1%, which is the clearest signal of where the next round of product work should go.
- ›Metronix leads the equal-conditions comparison on LoCoMo and MemoryAgentBench.
- ›LoCoMo and LongMemEval-S show a persistent retrieval-greater-than-generation gap: the system usually finds the evidence before the answer model fully capitalizes on it.
- ›The weak spots are clear instead of hidden: conflict resolution, preference following, and large-ingest BEAM workloads still need product work.
Better memory changes whether agents preserve user preferences, recover the right prior context, and avoid repeating work. The important shift here is not just stronger numbers. It is that the numbers now survive scrutiny instead of depending on benchmark theater.
- ›All results are directional, N=1. Useful for strategy and product positioning, not a publication-grade final word.
- ›Mem0 exposes headline Layer B numbers, but not the stable provenance needed for full retrieval-layer comparison.
- ›BEAM 1M has not been run yet, and GBrain did not scale to the high-volume MAB and BEAM tiers in this test harness.
The benchmark story is now stronger because it is more honest. Metronix looks like one of the top memory systems in the field, and the remaining gaps are specific enough to guide the roadmap instead of hiding behind shiny but incomparable numbers.