{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"DEV","title":"Compressing Transformer Models With Weight Clustering","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/5ce5d2d0\"></iframe>","width":"100%","height":180,"duration":495,"description":"Transformer models are powerful but notoriously heavy — weight clustering offers a practical path to shrinking them without gutting their performance. This episode breaks down how the technique works, how to implement it, and why it belongs in every AI engineer's deployment toolkit.","thumbnail_url":"https://img.transistorcdn.com/FjCd-OuusfvO3o_XEB1lBI9M3jCiMFpn2OICEsvCyrs/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9kYzVl/MjVhMjFhZGZhOTg4/Zjc1YTFlMGNkZWE1/ZmVhMi5wbmc.webp","thumbnail_width":300,"thumbnail_height":300}