Theory first, then code, then both at once
For the first four weeks the lectures come in pairs. A theory lecture on Monday explains an idea, for example how a GPU moves data from memory to its cores. On Tuesday, a coding lecture has you write a kernel that depends on that idea and measure whether it behaved as predicted. Thursday and Friday follow the same pattern. The one exception is the coding lecture on October 13, on parallelism on the CPU, which stands on its own. From November 9 the material is frontier work, where the idea and the kernel are hard to separate, so each lecture covers both: six deep-dives into the modern stack and two on kernels written by AI. Every lecture is recorded, and all code lives in the course repository.

Dr. Raj Dandekar
MIT PhD, founder of Vizuara. Designed and taught Vizuara's 5D Parallelism and Inference Engineering workshops.

Shubham Panchal
Creator of SmolChat. Works across C, C++ and Rust, deploying ML models on-device.
When
7:00 to 9:00 AM IST, every Monday, Tuesday, Thursday and Friday from October 12 to November 20, then demo day on December 7.
GPUs
No GPU of your own is needed. Sessions run on cloud GPUs, and the setup is covered in the first week. Hopper and Blackwell sessions use H100 and B200 machines.
Guest lecture
A session with the Crusoe team, in the middle or towards the end of the cohort. The date is announced inside the cohort.
The schedule
Readings are optional before the session and useful after it. Folder names under a session refer to the course repository. Dates for later sessions may move by a day or two; changes will appear here first.
| Date | Session | Readings and code | Deadlines |
|---|---|---|---|
| Week 1Oct 12 to Oct 16Part I · Why GPUs, and how they run your code | |||
| Mon Oct 12 | Lecture 1 · Theory How fast can this go? Qwen3-8B on one H100 writes about 149 tokens per second, far below what the chip's arithmetic allows. We find out why with the roofline model: arithmetic intensity, memory-bound and compute-bound work. In the second hour, the class becomes the batch and we watch throughput change as more users arrive. |
code · Lecture 1 practical kit |
|
| Tue Oct 13 | Lecture 2 · Coding Parallelism on the CPU Instruction-level parallelism: pipelining, branches and out-of-order execution. Threads in the Linux kernel and the 1:1 and M:N threading models of Java, Go and Rust. SIMD vector instructions. Why graphics, and then matrix multiplication, pushed computing towards GPUs. This lecture stands on its own; it is the CPU baseline that every later lecture compares against. |
code · 00_cpu_parallelism |
Out: Assignment 1 |
| Thu Oct 15 | Lecture 3 · Theory How does a GPU run your code? We follow one image through an H100. Grids, blocks, threads and warps; streaming multiprocessors and the warp scheduler; registers, shared memory and occupancy; and why 32 threads move as one. |
||
| Fri Oct 16 | Lecture 4 · Coding Your first CUDA kernels Vector addition and 2D max-pooling in CUDA C++. Compute capability and what nvidia-smi reports. How nvcc turns a .cu file into PTX and then SASS, and how to read both for the vecadd kernel. |
code · 02_kernels_ptx_sass |
Out: Assignment 2 |
| Week 2Oct 19 to Oct 23Part I · Memory, and Part II · The GEMM worklog begins | |||
| Mon Oct 19 | Lecture 5 · Theory The memory hierarchy Registers, shared memory, L1 and L2, and HBM. Cache lines and 32-byte sectors. Coalesced and uncoalesced loads. Shared-memory banks and bank conflicts. Why most kernels spend their time waiting for data rather than doing arithmetic. |
Due: Assignment 1 | |
| Tue Oct 20 | Lecture 6 · Coding The transpose ladder, and reading Nsight Compute A naive transpose, then a shared-memory tile that makes the writes coalesced, then padding that removes the bank conflicts. Along the way we learn Nsight Compute: sectors per request, L2 hit rate, bank conflicts, and a speed-of-light comparison of a memory-bound ReLU against a compute-bound matmul. |
code · 03_matrix_transpose · 01_benchmark_pytorch_kernels |
Out: Assignment 3 |
| Thu Oct 22 | Lecture 7 · Theory GEMM worklog I Why matrix multiplication is the operation AI runs more than any other. The naive kernel and the small fraction of the GPU it uses. Warp states and stall reasons. Shared-memory tiling and giving each thread more than one output. |
||
| Fri Oct 23 | Lecture 8 · Coding GEMM by hand I Four kernels, from the naive version to shared-memory tiles where each thread computes a 4 by 4 block with float4 loads. Each one is profiled in Nsight Compute and timed against cuBLAS. |
code · 04_gemm_A |
Due: Assignment 2 |
| Week 3Oct 26 to Oct 30Part II · GEMM and tensor cores | |||
| Mon Oct 26 | Lecture 9 · Theory GEMM worklog II 2D block tiling, register reuse, vectorised memory access, warp tiling and autotuning: the steps that take a hand-written FP32 kernel close to cuBLAS. |
||
| Tue Oct 27 | Lecture 10 · Coding GEMM by hand II We build the 2D-tiled and warp-tiled kernels step by step and keep a worklog: one change, one measurement, one profile, then the next change. |
Due: Assignment 3Out: Assignment 4 | |
| Thu Oct 29 | Lecture 11 · Theory Tensor cores What a tensor core computes in a single instruction, and why it changes the GEMM ladder. FP16, BF16 and FP8. WMMA and mma.sync fragments, ldmatrix and shared-memory swizzling. |
||
| Fri Oct 30 | Lecture 12 · Coding A WMMA tensor-core GEMM We rewrite our best GEMM to use tensor cores through the WMMA API, check it against cuBLAS, and look at what the profiler says has become the new bottleneck. |
||
| Week 4Nov 2 to Nov 6Part III · Profiling and attention | |||
| Mon Nov 2 | Lecture 13 · Theory Profiling and debugging like a pro Reading an Nsight Compute report from top to bottom. Stall reasons and the source view with SASS next to your code. compute-sanitizer for races and out-of-bounds accesses. A repeatable method for finding out why a kernel is slow. |
||
| Tue Nov 3 | Lecture 14 · Coding Debug three sabotaged kernels Three kernels that are wrong or slow on purpose. You find each problem with the profiler and the sanitizer before you are allowed to read the source. |
||
| Thu Nov 5 | Lecture 15 · Theory Attention: the kernel that ate the world Attention as two matrix multiplications and a softmax. Why the N by N score matrix stops fitting. Safe softmax and online softmax. Tiling and IO-awareness, the idea behind FlashAttention. The lecture ends with the capstone kickoff. |
Capstone kickoff | |
| Fri Nov 6 | Lecture 16 · Coding Build FlashAttention live Vanilla attention in PyTorch until it runs out of memory, then FlashAttention written in Triton and benchmarked from N = 2,048 to N = 65,536. |
code · flash_attention |
Due: Assignment 4Out: Assignment 5 |
| Week 5Nov 9 to Nov 13Part IV · The modern frontier | |||
| Mon Nov 9 | Lecture 17 FlashAttention from scratch: FA2 and FA3 We rebuild FlashAttention properly, then see what changed in FlashAttention 2 (better work partitioning across warps and blocks) and FlashAttention 3 (Hopper's asynchronous tensor cores and FP8). |
code · flash_attention |
|
| Tue Nov 10 | Lecture 18 Beating cuBLAS on an H100 What Hopper added and why it matters for GEMM: the Tensor Memory Accelerator (TMA), wgmma, thread-block clusters, asynchronous pipelines and warp specialisation. We use them to build a matrix multiply that beats NVIDIA's own library. |
||
| Thu Nov 12 | Lecture 19 Triton, CUTLASS and CuTe-DSL The abstraction ladder above raw CUDA. Triton's blocks, CUTLASS templates, CuTe layouts, and CuTe-DSL for writing CUTLASS-class kernels in Python. We write one fused kernel in each and compare the code and the speed. |
Due: Assignment 5 | |
| Fri Nov 13 | Lecture 20 Inference-serving kernels Prefill and decode behave like two different workloads. The KV cache and PagedAttention. Fused RMSNorm and rotary embeddings. Quantised GEMMs. What speculative decoding asks of the kernels underneath it. |
Capstone proposal due | |
| Week 6Nov 16 to Nov 20Part IV · Blackwell, and Part V · AI-written kernels | |||
| Mon Nov 16 | Lecture 21 Blackwell and NVFP4 What changed from Hopper to Blackwell, and the new tensor-core instructions that come with it. FP4 numbers with block scaling, and what NVFP4 does to throughput and to accuracy. We run an NVFP4 GEMM on a B200. |
||
| Tue Nov 17 | Lecture 22 Flash Attention 4 From FlashAttention 3 on Hopper to FlashAttention 4 on Blackwell. Why, on the newest chips, the exponentials in softmax rather than the matrix multiplications become the limit, and the tricks FA4 uses to work around it. |
code · flash_attention/papers |
|
| Thu Nov 19 | Lecture 23 DeepSeek: FlashMLA and DeepGEMM Multi-head latent attention and why it needs its own decode kernel. FP8 GEMMs with fine-grained scaling for mixture-of-experts models. What a small, focused kernel library can do against large general ones. |
||
| Fri Nov 20 | Lecture 24 Kernels written by AI: the agent and profiler loop KernelBench and its fast_p metric. An LLM proposes a kernel, a harness checks correctness and times it, and the profiler output goes back into the next prompt. How to tell a generated kernel that is actually faster from one that only looks fast, and where the kernel engineer still decides. |
Capstone checkpoint 1: baseline | |
| Week 7 and 8Nov 23 to Dec 4Part VI · Capstone build | |||
| Fri Nov 27 | Capstone Capstone checkpoint 2 No lectures in these two weeks. You build your capstone. By this date, share a first optimised version with a profile that explains the speed-up. |
Capstone checkpoint 2 | |
| Sun Dec 6 | Capstone Capstone code and worklog due Final code, correctness tests, benchmark harness and the worklog. |
Final submission | |
| Week 9Dec 7Part VI · Capstone | |||
| Mon Dec 7 | Lecture 25 · Capstone Capstone demo day Every team presents its kernel, its baseline, its measured speed-up and its worklog. Crusoe joins the review. |
Demo day | |
Assignments
Each assignment follows the lectures it depends on. Most are released on the day of a coding lecture and are due about a week later; the GEMM worklog gets ten days because it is the longest. After Assignment 5, the capstone takes over.
Roofline by hand, and the CPU baseline
Compute the arithmetic intensity of five operations (vector add, ReLU, a softmax row, a matrix-vector product and a 4096 by 4096 GEMM) and place each on the H100 roofline. Then parallelise a CPU loop with threads and with SIMD and measure how it scales with core count.
You hand in: A short report with your numbers, your plot and one paragraph on what surprised you.
First kernels
Write SAXPY, an RGB-to-grayscale kernel and a 2D max-pool in CUDA. Sweep the block size from 32 to 1,024 threads and plot the runtime. Dump the PTX and SASS of one kernel and mark where the loads, the arithmetic and the stores happen. Work out the occupancy for three register counts.
You hand in: Code, the plot and the annotated SASS.
Memory and measurement
Get a matrix transpose to at least 80% of the bandwidth of a plain device-to-device copy and show with Nsight Compute why it is fast. Then take six PyTorch operations, predict from the roofline whether each is memory-bound or compute-bound, and check your prediction with the profiler.
You hand in: Code, profiles and a table of predictions against measurements.
Your own GEMM worklog
Take an FP32 GEMM from the naive kernel to at least 70% of cuBLAS at 4096 by 4096 by 4096, and add one tensor-core version with WMMA. Every step in the worklog needs a measurement and a profiler screenshot that explains it.
You hand in: Code and a worklog in the style of the GEMM lectures.
FlashAttention forward pass
Implement a causal FlashAttention forward pass in Triton. Match PyTorch's scaled_dot_product_attention within BF16 tolerance, benchmark it from N = 1,024 to N = 32,768, and report peak memory against vanilla attention. The backward pass is an optional extension.
You hand in: Code, a correctness test and the benchmark table.
The capstone project
The capstone is a real kernel problem solved from start to finish: a baseline you can trust, a kernel that beats it, and a worklog that explains every step. Crusoe shares project ideas at the start of the cohort, so the problems match what an AI-infrastructure company actually works on. The ideas below are examples of the scale we have in mind.
- Mon Oct 12Project ideas published, including problems from Crusoe.
- Thu Nov 5Capstone kickoff at the end of the attention lecture. Work solo or in pairs.
- Fri Nov 13Proposal due: one page with the problem, the baseline, the target metric and the GPU you need.
- Fri Nov 20Checkpoint 1: a correctness test, a benchmark harness and a measured baseline.
- Nov 23 to Dec 4Two weeks with no lectures, kept free for the capstone build.
- Fri Nov 27Checkpoint 2: a first optimised version, with a profile that explains the speed-up.
- Sun Dec 6Final code and worklog due.
- Mon Dec 7Demo day, Lecture 25.
A decode attention kernel for a paged KV cache
Beat the vLLM kernel on one model's shapes, at batch sizes that matter in serving.
A fused RMSNorm, rotary and quantise kernel
Replace three launches in a real transformer block with one, and measure the end-to-end effect.
FP8 or NVFP4 GEMMs for one model's layer shapes
Tune for the exact matrix shapes of an open model such as Qwen3-8B on H100 or B200.
A grouped GEMM for mixture-of-experts routing
Handle uneven expert loads without padding every expert to the largest one.
You against the machine
Optimise the same kernel by hand and with an agent loop, and write up what each found and missed.
Speed up a served model
Profile a vLLM deployment on Modal, find the slowest kernel, replace it, and show the change in tokens per second.
- Correct. It passes a test against a trusted reference, including awkward sizes.
- Faster, and measured honestly. Warm-up, synchronisation and fixed clocks, against a fair baseline.
- Explained. A worklog where every change comes with a measurement and a profile.
- Defended. At demo day you can answer why each step worked, which is what an on-site interview will ask.
