1. An Inference Engine in 1,500 Lines of CUDA: The Mapjmaczan/tiny-vllm is an inference engine for Llama 3.2 1B Instruct, written from scratch in C++ and CUDA. No PyTorch, no Hugging Face, not … September 17, 2026 10 min