Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs
-
Updated
Aug 20, 2026 - HTML
Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs
Modular C++ Toolkit for Performance Analysis and Logging. Profiling API and Tools for C, C++, CUDA, Fortran, and Python. The C++ template API is essentially a framework to creating tools: it is designed to provide a unifying interface for recording various performance measurements alongside data logging and interfaces to other tools.
Analyze LLM inference: FLOPs, memory, Roofline model. Supports GQA, MoE, MLA, RoPE, SwiGLU. 19 models × 20+ hardware platforms.
Why is LLM inference slow — and how do you make it fast? A hands-on, first-principles course: roofline → KV cache → quantization → parallelism → vLLM/SGLang, with GPU labs on open models.
Hand-written CUDA vs Mojo GPU kernels benchmarked on consumer Ampere (RTX 3090, sm_86), with roofline analysis
Interactive theoretical Kimi-K3 inference roofline calculator for H200, B300, and GB300
The physics before the engines: measured arithmetic intensity, prefill compute-bound vs decode memory-bandwidth-bound, with the roofline plot.
Hand-written CUDA linear algebra library. A GEMM driven from a naive baseline to a compute-bound tensor-core kernel at ~90% of cuBLAS on an RTX 5070, profiled with Nsight Compute, plus GEMV, TRSM, CSR SpMV, cuSOLVER, a measured roofline, and NVML telemetry.
Interactive 3D visualization of dense decoder-only LLM inference. Companion to the AI Inference Engineer 2026 course.
Is your llama.cpp decode memory-bandwidth-bound? Find out in one command.
Interactive macOS GPU benchmark suite for Apple Silicon - 25 Metal compute benchmarks (compute, memory, ML/tensor, ray tracing) with a full measurement harness and empirical roofline.
Three hand-written Triton kernels for LLM inference (fused RMSNorm plus residual, online softmax, INT4 g128 GEMV) benchmarked on NVIDIA Blackwell against PyTorch eager and torch.compile, with every raw CUDA-event sample, measured device ceiling, and Nsight Compute report committed and CI-verified.
Pipeline-aware Roofline & Inference Sweep Model — wall-clock + TCO prediction for NPU architecture exploration, calibrated on Ascend 910B4 via msprof. Other architectures are under consideration.
Measurement-driven study of LLM training (FSDP2) and serving (vLLM) on 8×A100 — 16 reproducible incident write-ups that predict the bound, control the confound, and report the negative.
Where does the next watt go? Choosing an AI-accelerator budget allocation when the workload mix is unknown -- max-expected vs minimax-regret over the whole forecast space. Normalized model, stdlib only.
GPU performance optimization labs covering Roofline Analysis, LLM Decode Optimization, and CUDA Graphs using PyTorch
Qwen3.5-2B BF16 Roofline and Nsight profiling on Jetson Orin Nano
Double-precision dense GEMM across sequential, OpenMP, MPI and CUDA on an i9-12900K, Tesla T4 and Tesla P100. The T4 loses to the CPU in FP64 (Turing runs it at 1:32); a P100 at 1:2 explains why. Raw data, Intel Advisor and Nsight profiles, notebooks and a 33-page report.
Add a description, image, and links to the roofline topic page so that developers can more easily learn about it.
To associate your repository with the roofline topic, visit your repo's landing page and select "manage topics."