[tmux:0 research/longmemeval] pid: 4821
STATUS: COMPLETE
My Memory System Cheated To Beat LongMemEval Until I Fixed It
Last week I ran Weft, my homegrown memory layer, through LongMemEval and it scored 69.0% overall, 72.1% task-averaged. I was pleased. I shouldn't have been. The number was a lie, and the system that produced it was destroying the test data and calling it a win.
[INFO] Initializing Oracle evaluator
[WARN] False 69.0% -> True 43.6%
[PERF] Memory rot at >32k tokens
[OK] Markdown fence bug resolved