{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"Tech Stories Tech Brief By HackerNoon","title":"How Fast Can DeepSeek Run on 8GB VRAM?","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/cce7de3b\"></iframe>","width":"100%","height":180,"duration":195,"description":"\n        This story was originally published on HackerNoon at: https://hackernoon.com/how-fast-can-deepseek-run-on-8gb-vram.\nRethinking local LLM inference as a full resource path across disk, RAM, PCIe, VRAM and compute—and why residency is only one part of the problem.\nCheck more stories related to tech-stories at: https://hackernoon.com/c/tech-stories.\n            You can also check exclusive content about #local-llm, #llm-inference, #consumer-hardware, #vram, #gpu, #performance-optimization, #deepseek, #open-source,  and more.\nThis story was written by: @speederx. Learn more about this writer by checking @speederx's about page,\n            and for more stories, please visit hackernoon.com.\nAfter realizing that expert residency alone wasn’t the full performance wall, I started looking at local LLM inference as a complete resource path across disk, RAM, PCIe, VRAM and compute. Recent work like DwarfStar4 and FreeToken is converging on parts of the same problem. My angle is to measure the real bottleneck first, then decide what to optimize.","thumbnail_url":"https://img.transistorcdn.com/IuqXIpaNNuezY7jNfIDnL5gqB1iL_SEndwUUzLGdljY/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9zaG93/LzQxNDI5LzE2ODM1/ODM0NjQtYXJ0d29y/ay5qcGc.webp","thumbnail_width":300,"thumbnail_height":300}