Monarchic / Independent AI R&D

Now / Agent enhancements Next / Agent workflows

ExplicitMem / Published evaluations

Memory evaluation report

Four separately scoped evaluations cover answer synthesis, LoCoMo retrieval, cross-dataset candidate generation, and generic-answer support. Every result keeps its date, sample, broad method, and interpretation limit beside it.

Study register

Metric labels are intentionally different because each evaluation asks a different question.

01 / External benchmark

LongMemEval-S

Fixed answer synthesis

99.8%

Answer accuracy

499 / 500 answers correct / generated 1 August 2026

Method

The complete 500-question set was scored with deterministic normalized answer matching under a fixed evaluation profile.

Observed failure

One temporal item differed from the upstream gold answer, so the strict perfect-score gate remained unmet.

Interpretation limit

Benchmark-scoped result. It is not an LLM-judged score, a cross-dataset synthesis result, or a hosted latency claim.

Retained records

Versioned result and failure records retained for technical review.

02 / External benchmark

LoCoMo

Retrieval evaluation

93.52%

Expected-memory recall

1,986 runtime cases / generated 3 June 2026

Method

A fixed retrieval profile was evaluated against 1,986 cases using expected-memory recall as the primary metric.

Observed failure

Secondary recall measures were lower than the primary measure and remain outside this headline claim.

Interpretation limit

LoCoMo-specific retrieval and context recall. It does not measure generated-answer accuracy and is not the default generic runtime.

Retained records

Versioned result and failure records retained for technical review.

03 / Monarchic evaluation on external datasets

BEAM, ConvoMem, LoCoMo, LongMemEval

Cross-dataset candidate generation

98.93%

Held-out recall@400

740 held-out source cases / generated 3 June 2026

Method

A fixed candidate-generation profile was evaluated on held-out cases drawn from five external datasets.

Observed failure

Smaller candidate pools performed below the published recall@400 result.

Interpretation limit

High-recall candidate generation for downstream ranking. It is not final top-50 ranking quality or answer accuracy.

Retained records

Versioned result and split records retained for technical review.

04 / Monarchic evaluation

Generic and non-LongMemEval fixtures

Generic answer support

100%

Deterministic accuracy

256 / 256 cases passed / generated 6 June 2026

Method

A fixed set of local, external, production-shaped, unseen-generic, and generated cases was scored with deterministic support checks.

Observed failure

No scored failure appears in this artifact. That limits failure analysis and increases the importance of broader adversarial and provider-judged evaluation.

Interpretation limit

Deterministic extractive answer support. It is not LLM-judged provider scoring or proof of open-ended answer quality.

Retained records

Versioned result records retained for technical review.

LongMemEval-S detail

What the synthesis run measured

Questions

500

Answer accuracy

99.8%

Answer faithfulness

100%

Answer source hit rate

78.2%

Retrieval recall@K

100%

Retrieval receipts

100%

Where the system was strongest and weakest

Question type Questions Answer accuracy Scored result
Knowledge update 78 100% 78 / 78
Multi-session 133 100% 133 / 133
Single-session assistant 56 100% 56 / 56
Single-session preference 30 100% 30 / 30
Single-session user 70 100% 70 / 70
Temporal reasoning 133 99.2% 132 / 133

Fixed evaluation protocol

The run used the complete LongMemEval-S question set, the promoted evaluation profile, and deterministic normalized answer matching.

One upstream-gold mismatch remains

The only scored miss occurs on one temporal item where the result differs from the upstream gold answer. The strict 100% threshold remains unmet.

No cross-system comparison

Mem0 and Supermemory are excluded because no comparable run used this exact dataset, question order, scoring method, context budget, and latency boundary.

Disclosure policy

What we publish

Published

Dataset, evaluation date, sample size, metric, broad method, observed failure, and interpretation limit.

Controlled review

Detailed implementation, internal artifact paths, validation commands, model dependencies, and deployment configuration remain private. Qualified customers and auditors can request a technical review.

Source and interpretation

Read the result within its benchmark

Evaluation records are versioned by study date and metric. LongMemEval-S, LoCoMo, cross-dataset retrieval, and generic-answer support remain separate claims. None establishes hosted latency or direct superiority over another memory product.

View ExplicitMem