Home / Compare

Win the bake-off on the numbers.

Recall quality on LOCOMO judged by mem0's own harness, plus five regulated invariants scored across six products - all reproducible from the open repo.

Recall quality · LOCOMO · run 2026-07-09

Their benchmark. Their harness. Our number.

92.9%Lians · LLM-judged QA accuracy 92.5%Mem0 · current published top-200 score

Reported LoCoMo results, scored with Mem0's unmodified pipeline. Settings differ, so this is not a controlled head-to-head. Protocol and caveats →

Recall quality · LongMemEval-S · run 2026-07-12

LongMemEval: they lead, we publish it anyway.

89.4%Lians · LLM-judged, all 500 questions 94.8%mem0 · their published score

The same harness finds the answer sessions for 97% of questions. The remaining gap is answer assembly.

Regulated-memory eval · head-to-head

Five invariants. One winner.

Regulated invariantLiansZep / GraphitiLettaHindsightSupermemorymem0
Stale revision suppressedpasspartialpartialpartialpartialfail
Point-in-time (as-of) recallpasspartialabsentpartialabsentabsent
Provable erasure (crypto-shred + cert)passpartialpartialabsentpartialpartial
Lookahead / backtest guardpassabsentabsentabsentabsentabsent
Audit-state snapshot at Tpasspartialabsentabsentabsentabsent
Score (pass = 1, partial = ½)5.0 / 52.0 / 51.0 / 51.0 / 51.0 / 50.5 / 5

Lians, mem0 OSS, and Graphiti OSS executed live in their default documented configs; Letta, Hindsight, and Supermemory via real-SDK capability adapters. Last run 2026-07-04. Methodology →

Why audit matters

Storage is not the moat.

Bring the numbers to your compliance officer.

Get an API key →