TL;DR — MemER scene-mem v2 LoRA, held-out val: final-query 0.0% (0/30); true compressed-memory recall (chunked) 0.0%
0.0%
PRIMARY final-query
0/30
4.16%
memory-turn match
28/673
0.0%
CHUNKED (true mem)
0/20
0.0%
NOT-chunked floor
0/10
Eval set (honest split)
Held-out val only: 30 test samples across
10 scenarios (split-val total
57). Train leak:
NONE. MemER did NOT train on these scenarios.
base = gemini-2.5-flash
adapter = None
max_new_tokens = 2048
(2) Per-memory-turn narration match rate
| bucket | match rate | count |
| overall | 4.16% |
28/673 |
| fine segments | 4.65% |
28/602 |
| composite segments (verbatim) | 0.0% |
0/71 |
(3) Keyframe selection quality
keyframe-count distribution:
{0: 126, 1: 918, 2: 1003, 3: 182, 4: 124, 5: 58, 6: 14, 7: 4, 8: 7}
mean keyframes/turn: 1.816
final-query emitted-a-keyframe rate:
96.67%
(29/30)
(4) Chunked-vs-not memory-recall decomposition
| bucket | final-query acc | count |
| CHUNKED (true compressed-memory recall) |
0.0% |
0/20 |
| NOT-chunked (learned-prior floor) |
0.0% |
0/10 |
chunked = the test's target task instruction matches the _PNP pick-and-place regex reused verbatim from scripts/analyze_tier3b_vs_tier3a.py (object genuinely entered the bounded FIFO). chunked = true compressed-memory recall; notchunked = learned-prior floor.
Run
parse failures: 10 ·
duration: 26824.9s