{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"Tech Stories Tech Brief By HackerNoon","title":"Experimental Results from a Self-Improving Retrieval System for Conversational Memory","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/3ad0c965\"></iframe>","width":"100%","height":180,"duration":2671,"description":"\n        This story was originally published on HackerNoon at: https://hackernoon.com/experimental-results-from-a-self-improving-retrieval-system-for-conversational-memory.\nEighteen retrieval experiments on agent memory: why BM25 dominates, what clustered retrieval-induced forgetting actually does, and the Rust port that shipped.\nCheck more stories related to tech-stories at: https://hackernoon.com/c/tech-stories.\n            You can also check exclusive content about #agent-memory, #rag, #bm25, #retrieval-systems, #cross-encoder-reranking, #longmemeval, #faiss, #hackernoon-top-story,  and more.\nThis story was written by: @teimurjan. Learn more about this writer by checking @teimurjan's about page,\n            and for more stories, please visit hackernoon.com.\nThe biology-inspired mutation layer didn't work. A learned MLP adapter and segmentation mutation both produced ~zero NDCG lift on LongMemEval. The control loop was sound; the perturbations weren't load-bearing.\n\nA recall diagnostic reframed the project: 78% of relevant entries never reached the cross-encoder. Bi-encoder recall was the ceiling, not the mutation layer.\n\nStandard IR wins compounded: 0.95-cosine dedup plus BM25 alongside vector plus cross-encoder rerank took NDCG@10 from 0.22 to 0.34. BM25 alone beat pretrained embeddings by 76% on this corpus.\n\nClustered retrieval-induced forgetting (Anderson 1994, ported as far as I can tell for the first time) added +1.9pp NDCG with p=0.0001 on LongMemEval. Regresses on NFCorpus: the mechanism is scoped to single-user long-term conversation memory, not general IR.\n\nWrite-time LLM enrichment (gist plus anticipated queries via Haiku) was the biggest single lever: +8.3pp NDCG on covered queries.\n\nA regex-tokenizer fix that BM25 had been missing was worth +1.4pp NDCG on the headline benchmark.\n\nSix independent ablations (reranker swap, BGE bi-encoder, multi-field BM25, field-boosted BM25, late chunking on a GPU, k_deep sweep) all bounced off the same ceiling:...","thumbnail_url":"https://img.transistorcdn.com/IuqXIpaNNuezY7jNfIDnL5gqB1iL_SEndwUUzLGdljY/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9zaG93/LzQxNDI5LzE2ODM1/ODM0NjQtYXJ0d29y/ay5qcGc.webp","thumbnail_width":300,"thumbnail_height":300}