01 / External benchmark
LongMemEval-S
Fixed answer synthesis
99.8%
Answer accuracy
499 / 500 answers correct / generated 1 August 2026
Method
The complete 500-question set was scored with deterministic normalized answer matching under a fixed evaluation profile.
Observed failure
One temporal item differed from the upstream gold answer, so the strict perfect-score gate remained unmet.
Interpretation limit
Benchmark-scoped result. It is not an LLM-judged score, a cross-dataset synthesis result, or a hosted latency claim.
Retained records
Versioned result and failure records retained for technical review.