fp16.in is about running language models in production — inference serving, throughput and latency, quantisation, and the GPU infrastructure holding all of it up. What the papers claim, and what actually survives real traffic.
▌
I work on the core infrastructure behind language models — taking generative AI systems from prototype to production and keeping them healthy once real traffic arrives. Most of my time goes to inference serving and throughput, and to managing and scaling the GPU fleets underneath: memory budgets, batching and KV cache behaviour, multi-GPU placement, utilisation and cost. This is where I write down how that actually works, in detail, minus the parts that don’t hold up.
No schedule, no series. Just a note when something’s worth reading.