--- title: "Benchmarks" description: "BEAM accuracy at 100K, 500K, 1M and 10M tokens against every published system, the per-ability breakdown, LoCoMo on corrected keys, and the evaluation method." canonical: https://past.dev/benchmarks last-updated: 2026-10-06 --- # Benchmark results and methodology > past.dev ranks first on BEAM at every size tested, on complete splits from 100K to 10M tokens. Current public results report answer accuracy on BEAM and LoCoMo. The evaluation method treats retrieval sufficiency, cost, and latency as separate metrics. Source: https://past.dev/benchmarks ## BEAM accuracy by history size past.dev figures from the run of 29 September 2026. Every other figure is that vendor's own published number, linked in the table. - 92.08% at 100K tokens, 15.2 points above Exabase M-1, the next published result at this size. - 89.63% at 500K tokens, 18.5 points above Hindsight, the next published result at this size. - 90.65% at 1M tokens, 15.7 points above Exabase M-1, the next published result at this size. - 85.03% at 10M tokens, 17.0 points above Exabase M-1, the next published result at this size. | System | 100K | 500K | 1M | 10M | Source | | --- | --- | --- | --- | --- | --- | | past.dev | 92.08% | 89.63% | 90.65% | 85.03% | [past.dev](/benchmarks) | | Exabase M-1 | 76.9% | not published | 75.0% | 68.0% | [exabase.io](https://exabase.io/blog/exabase-m1-achieves-state-of-the-art-on-beam-benchmark) | | Hindsight | 75.0% | 71.1% | 73.9% | 64.1% | [benchmarks.hindsight.vectorize.io](https://benchmarks.hindsight.vectorize.io/) | | Honcho | 63.0% | 64.9% | 63.1% | 40.6% | [plasticlabs.ai](https://plasticlabs.ai/blog/research/Benchmarking-Honcho) | | mem0 | not published | not published | 64.1% | 48.6% | [mem0.ai](https://mem0.ai/blog/ai-memory-benchmarks-in-2026) | Every competitor figure is that vendor's own published result, linked in the table. These are separately published runs rather than one harness, so models, judges and configurations differ between rows, and the past.dev figures are complete splits at every size. [The BEAM paper, published at ICLR 2026](https://arxiv.org/abs/2510.27246) [Reproduce the BEAM run, on GitHub](https://github.com/pastdotdev/benchmarks/tree/main/beam) [The BEAM leaderboard, ranked at each size](https://past.dev/benchmarks/beam) ## BEAM results by ability From the same run as the figures above. ### 100K tokens: 92.08% Overall accuracy over 400 questions across 20 conversations. Complete split. | Ability | Accuracy | n | | --- | --- | --- | | Abstention | 98.75% | 40 | | Contradiction resolution | 78.13% | 40 | | Event ordering | 100.00% | 40 | | Information extraction | 95.00% | 40 | | Instruction following | 96.25% | 40 | | Knowledge update | 89.38% | 40 | | Multi-session reasoning | 68.95% | 40 | | Preference following | 100.00% | 40 | | Summarization | 99.38% | 40 | | Temporal reasoning | 95.00% | 40 | ### 500K tokens: 89.63% Overall accuracy over 700 questions across 35 conversations. Complete split. | Ability | Accuracy | n | | --- | --- | --- | | Abstention | 92.14% | 70 | | Contradiction resolution | 78.57% | 70 | | Event ordering | 100.00% | 70 | | Information extraction | 86.32% | 70 | | Instruction following | 92.02% | 70 | | Knowledge update | 86.07% | 70 | | Multi-session reasoning | 71.39% | 70 | | Preference following | 99.05% | 70 | | Summarization | 97.25% | 70 | | Temporal reasoning | 93.45% | 70 | ### 1M tokens: 90.65% Overall accuracy over 700 questions across 35 conversations. Complete split. | Ability | Accuracy | n | | --- | --- | --- | | Abstention | 97.14% | 70 | | Contradiction resolution | 75.36% | 70 | | Event ordering | 99.81% | 70 | | Information extraction | 83.98% | 70 | | Instruction following | 96.43% | 70 | | Knowledge update | 85.71% | 70 | | Multi-session reasoning | 72.82% | 70 | | Preference following | 98.87% | 70 | | Summarization | 99.60% | 70 | | Temporal reasoning | 96.79% | 70 | ### 10M tokens: 85.03% Overall accuracy over 200 questions across 10 conversations. Complete split. | Ability | Accuracy | n | | --- | --- | --- | | Abstention | 100.00% | 20 | | Contradiction resolution | 73.13% | 20 | | Event ordering | 100.00% | 20 | | Information extraction | 68.75% | 20 | | Instruction following | 95.00% | 20 | | Knowledge update | 92.50% | 20 | | Multi-session reasoning | 28.17% | 20 | | Preference following | 100.00% | 20 | | Summarization | 99.00% | 20 | | Temporal reasoning | 93.75% | 20 | ## LoCoMo 93.12% over 1,540 questions across 10 conversations, using the community-corrected answer keys. The macro average across the four categories is 89.18%. Run of 16 September 2026. | Category | Accuracy | n | | --- | --- | --- | | Single-hop | 95.01% | 841 | | Temporal | 93.46% | 321 | | Multi-hop | 93.26% | 282 | | Open-domain | 75.00% | 96 | ## Evaluation scope ### Datasets - **BEAM**: Complete splits at 100K, 500K, 1M and 10M tokens, graded across ten abilities. The results above. - **LoCoMo**: Run with the community-corrected answer keys. An audit found errors in roughly 6% of the original answers, and using those keys counts a label error as a memory error. ### How a result is reported - **Baselines**: Full context and plain RAG, both run alongside. - **Judges**: Each judge configuration is recorded with its false-accept rate. - **Metrics**: Accuracy, cost and latency reported separately rather than as one score. [How to reproduce this](https://past.dev/benchmarks/methodology) Get API key: https://sso.past.dev/sign-up