Open benchmark

Release
0.1 pre-release
Scope
Memory architecture
Execution
Hosted or local
License
MIT

Agentic Memory

Baseline Isolated Evaluation, Normalized Tiers.

A controlled protocol for estimating what a memory architecture adds while the reader, judge, corpus, prompts, budgets, and run design remain fixed.

T1 / Control Capture off · Memory off
Figure 01

One reader traversed under four controlled memory conditions.

Protocol2 × 2 isolation

The subject is the memory architecture—not the model.

2026
01 / T1

No memory

The no-context control. Correct answers are not credited to a memory system.

Capture off · Bespoke memory off
02 / T2

Reference capture

Shared automatic capture is enabled while the bespoke agentic memory remains absent.

Capture on · Bespoke memory off
03 / T3

Full composition

Both layers operate together, exposing positive or negative interaction effects.

Capture on · Bespoke memory on
04 / T4

Bespoke agentic memory isolated

The bespoke agentic memory operates without reference capture; T4−T1 is the primary contrast.

Capture off · Bespoke memory on
01

Measure inventory

Fifteen reported abilities

These abilities describe the behaviors a dependable memory system should make possible, from declining unsupported answers to preserving provenance, chronology, and set integrity. AMBIENT tests the observable result of each ability—not the internal mechanism a system uses to produce it.

01

Abstention

Declines when the record cannot support an answer instead of inventing one.

02

Adoption

Routes new memories into the governed store so they remain available later.

03

Anteriority

Establishes that one record existed before another without inventing time evidence.

04

Calibration

Keeps confidence proportional to evidence and resists confidence laundering.

05

Concurrency

Preserves simultaneous writes without silent loss, corruption, or cross-talk.

06

Contradiction resolution

Surfaces incompatible claims instead of choosing a convenient side as fact.

07

Enumeration

Returns the complete requested set without omissions, duplicates, or unrelated members.

08

Expiry

Stops treating a claim as current after its declared or semantic validity ends.

09

Federation

Combines independent stores while preserving origin and exposing cross-store conflicts.

10

Holonomy

Detects cyclic ordering contradictions such as A before B before C before A.

11

Modality

Keeps hypothetical, proposed, and actual events from being mistaken for one another.

12

Provenance

Preserves where a remembered claim came from and which evidence supports it.

13

Reactivity

Updates dependent conclusions when supporting memory changes or is superseded.

14

Reader independence

Shows the memory supplies facts the fixed reader could not recover alone.

15

Set integrity

Proves served items belong to the recorded set and that its history is append-only.

Read the ability guide
Run AMBIENT

Download it.
Run it.
Submit it here.

The repository is the complete benchmark and canonical protocol. This website certifies and places finished runs.

01

Download from GitHub

Clone the MIT-licensed repository to inspect the protocol, run locally, add an adapter, or submit an evidence bundle.

Open GitHub repository ↗
02

Submit on this site

Upload your evidence bundle with its frozen-corpus attestation. The automated certifier re-derives every number from your raw artifacts, and a clean pass places the row immediately.

Submit a run

No configuration or API credentials are collected by this website.

A robotic tape library selecting one barcode-indexed record from a dense archival store.
Figure 02

Does the memory help?

A useful system raises traced completion without increasing gullibility, unsupported correctness, or failures to serve needed evidence.

02

Scoring examples

Memory attribution

A correct answer is not enough.

AMBIENT asks whether memory caused the correct answer—and whether bad memory caused harm.

Memory deserves credit

Correct because memory served it

01Schedule correction
Question
When is the Atlas review?
Served memory
“The review moved from 12 May to 18 May.”
Answer
18 May
Completed · traced memory credit
02Policy change
Question
Which login method is now required?
Served memory
“Password login was replaced by passkeys on 4 June.”
Answer
Passkeys
Completed · traced memory credit
03Ownership handoff
Question
Who is the current incident lead?
Served memory
“Alex handed incident command to Noor at 14:20.”
Answer
Noor
Completed · traced memory credit
Memory does not deserve credit

Wrong—or correct for another reason

01Outdated memory
Question
When is the Atlas review?
Served memory
The superseded “12 May” record.
Answer
12 May
Wrong / gullible · no memory credit
02Poisoned memory
Question
Who owns the Orion release?
Served memory
An unverified note says Rowan; the governed record says Priya.
Answer
Rowan
Gullible · no memory credit
03Correct but untraced
Question
Which region hosts staging?
Served memory
No supporting record was served.
Answer
eu-west-1
Correct but untraced · zero memory credit
Integrity requirements and capability probes
04

Distribution

Choose the memory.

Participants run the benchmark on their own infrastructure against their own memory system through the AMBIENT adapter contract. The reader is fixed, answers are checked by exact mechanical oracles, and the harness changes only the memory condition.

Run scopes contain 10, 100, 200, or 400 unique questions sampled evenly across all ten BEAM abilities. Ten is explicitly a smoke test; the complete 400-question scope is the first hosted option near a ±5-point worst-case single-tier margin. Repeats never count as new questions.

The primary result is attributed memory lift: T4 memory-on completion minus T1 no-memory completion. The report also separates gullible answers, answers without traced support, and cases where the needed evidence was not served. This is not a model leaderboard.

Every run produces an evidence bundle, and results are never posted automatically from a run. A row appears only when the participant uploads that bundle here and it clears the automated certifier, which re-derives every published number from the raw artifacts and verifies the frozen-corpus attestation.