No memory
The no-context control. Correct answers are not credited to a memory system.
Agentic Memory
A controlled protocol for estimating what a memory architecture adds while the reader, judge, corpus, prompts, budgets, and run design remain fixed.
One reader traversed under four controlled memory conditions.
The subject is the memory architecture—not the model.
2026The no-context control. Correct answers are not credited to a memory system.
Shared automatic capture is enabled while the bespoke agentic memory remains absent.
Both layers operate together, exposing positive or negative interaction effects.
The bespoke agentic memory operates without reference capture; T4−T1 is the primary contrast.
Measure inventory
These abilities describe the behaviors a dependable memory system should make possible, from declining unsupported answers to preserving provenance, chronology, and set integrity. AMBIENT tests the observable result of each ability—not the internal mechanism a system uses to produce it.
Declines when the record cannot support an answer instead of inventing one.
Routes new memories into the governed store so they remain available later.
Establishes that one record existed before another without inventing time evidence.
Keeps confidence proportional to evidence and resists confidence laundering.
Preserves simultaneous writes without silent loss, corruption, or cross-talk.
Surfaces incompatible claims instead of choosing a convenient side as fact.
Returns the complete requested set without omissions, duplicates, or unrelated members.
Stops treating a claim as current after its declared or semantic validity ends.
Combines independent stores while preserving origin and exposing cross-store conflicts.
Detects cyclic ordering contradictions such as A before B before C before A.
Keeps hypothetical, proposed, and actual events from being mistaken for one another.
Preserves where a remembered claim came from and which evidence supports it.
Updates dependent conclusions when supporting memory changes or is superseded.
Shows the memory supplies facts the fixed reader could not recover alone.
Proves served items belong to the recorded set and that its history is append-only.
The repository is the complete benchmark and canonical protocol. This website certifies and places finished runs.
Clone the MIT-licensed repository to inspect the protocol, run locally, add an adapter, or submit an evidence bundle.
Open GitHub repository ↗Upload your evidence bundle with its frozen-corpus attestation. The automated certifier re-derives every number from your raw artifacts, and a clean pass places the row immediately.
Submit a runNo configuration or API credentials are collected by this website.
A useful system raises traced completion without increasing gullibility, unsupported correctness, or failures to serve needed evidence.
Scoring examples
AMBIENT asks whether memory caused the correct answer—and whether bad memory caused harm.
Distribution
Participants run the benchmark on their own infrastructure against their own memory system through the AMBIENT adapter contract. The reader is fixed, answers are checked by exact mechanical oracles, and the harness changes only the memory condition.
Run scopes contain 10, 100, 200, or 400 unique questions sampled evenly across all ten BEAM abilities. Ten is explicitly a smoke test; the complete 400-question scope is the first hosted option near a ±5-point worst-case single-tier margin. Repeats never count as new questions.
The primary result is attributed memory lift: T4 memory-on completion minus T1 no-memory completion. The report also separates gullible answers, answers without traced support, and cases where the needed evidence was not served. This is not a model leaderboard.
Every run produces an evidence bundle, and results are never posted automatically from a run. A row appears only when the participant uploads that bundle here and it clears the automated certifier, which re-derives every published number from the raw artifacts and verifies the frozen-corpus attestation.