Home / Compare

Compare the evidence.

Separate dated Lians measurements from current, documented product capabilities. This is an inspectable evaluation surface, not an independent vendor leaderboard.

Recall quality · LOCOMO · run 2026-07-09

Run the public test.

92.9%Lians · LLM-judged QA accuracy 10LOCOMO conversations · frozen run

Lians retrieval was scored end-to-end by the then-current, unmodified mem0ai/memory-benchmarks ↗ harness (gpt-5 answerer + judge), all 10 LOCOMO conversations. The result is frozen to July 9, 2026 and is not presented as a current cross-vendor ranking. Full report ↗

Recall quality · LongMemEval-S · run 2026-07-12

Publish the gap.

89.4%Lians · LLM-judged, all 500 questions 97%questions where retrieval surfaced an answer session

The dated run shows retrieval surfacing answer sessions more often than the final answerer converts them into correct answers. That measured answer-assembly gap remains visible in the open report.

Capability map · source-verified August 1, 2026

Compare the same evidence.

Temporal memory is no longer unique. The meaningful test is whether a system can reconstruct a named decision at both event-time and knowledge-time cutoffs, enumerate included and excluded source versions, detect future leakage, and verify a portable receipt.

SystemPublicly documented strengthQuestion for a live bake-offPrimary source
LiansDual-time reconstruction, supersession, leak checks, Evidence Packs, crypto-erasure certificates.Can the receipt prove exactly which source versions were knowable and excluded at both cutoffs?Open source ↗
Zep / GraphitiBitemporal knowledge graph and episode provenance.Can one portable artifact bind a decision, dual cutoffs, exclusions, erasure state, and integrity proof?Graphiti ↗
Mem0Temporal reasoning, as-of queries, memory history, and a broad integration ecosystem.Can the replay prove no post-cutoff source influenced the decision and verify that proof offline?Temporal reasoning ↗
HindsightQuery-time temporal recall, history and curation, audit controls, and OpenTelemetry.Can an exported receipt enumerate included and excluded source versions and survive crypto-erasure?Recall API ↗
SupermemoryContent versioning, temporal graph memory, connectors, and multimodal ingestion.Can a reviewer replay the exact decision boundary and verify its chain without trusting the UI?Graph memory ↗
LettaA full stateful-agent runtime with editable memory blocks and an agent development environment.Does the use case need a complete agent runtime, or a narrower evidence layer around an existing stack?Memory blocks ↗

Current Lians claim level: foundation verified. Production and competitive-leadership labels stay gated on load, isolation, restore, failure-injection, public benchmark, and independent-reproduction evidence. Inspect the claim policy ↗

Why audit matters

Storage is not the moat.

Run the same test.

Talk to us →