{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"Tech Stories Tech Brief By HackerNoon","title":"Why Local LLMs Suddenly Slow Down at Long Context","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/ecca952f\"></iframe>","width":"100%","height":180,"duration":331,"description":"\n        This story was originally published on HackerNoon at: https://hackernoon.com/why-local-llms-suddenly-slow-down-at-long-context.\nYour local LLM runs fine until it doesn't. A look at KV cache spilling from VRAM into shared memory, and why it happens silently on Windows.\nCheck more stories related to tech-stories at: https://hackernoon.com/c/tech-stories.\n            You can also check exclusive content about #local-llms, #llama.cpp, #kv-cache, #vram, #gpu, #local-inference, #machine-learning, #hackernoon-top-story,  and more.\nThis story was written by: @speederx. Learn more about this writer by checking @speederx's about page,\n            and for more stories, please visit hackernoon.com.\nYour local LLM runs fine until the context fills up past a certain point - then generation speed can drop by ~50%. The cause is the KV cache spilling out of VRAM into slower shared memory. On Windows it happens silently, with no out-of-memory error to warn you.\n        \n        ","thumbnail_url":"https://img.transistorcdn.com/IuqXIpaNNuezY7jNfIDnL5gqB1iL_SEndwUUzLGdljY/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9zaG93/LzQxNDI5LzE2ODM1/ODM0NjQtYXJ0d29y/ay5qcGc.webp","thumbnail_width":300,"thumbnail_height":300}