Inference Engineering: A Practical Tutorial from GPU Fundamentals to Production LLM ServingLearn inference engineering from the ground up — GPU architecture, LLM inference mechanics, quantization, speculative decoding, KV caching, vLLM, and production deployment. Hands-on tutorial with practical examples.