Pretrained

Kimi's serving architecture, mooncake to offload GPU memory to other chipsets, the ubiquity of vllm, and the growing standard LLM stack

What is Pretrained?

Everyone's talking about AI but most of it is hype, jargon, or someone trying to sell you something. Pierce Freeman and Richard Diehl Martinez are two Stanford friends who actually work in this space with nothing to sell you but good vibes. We'll make you smarter - and you might actually enjoy it.