Vizuara Kernel Engineering
Founding cohort · October 12 to December 7, 2026

Schedule

The workshop takes you from parallelism on a CPU to kernels written by AI, in 25 lectures. The first 16 come in theory and coding pairs. The next eight cover the 2025 and 2026 frontier, with theory and code together in each lecture. The last is capstone demo day on December 7. Lectures run on Monday, Tuesday, Thursday and Friday until November 20, and the two weeks after that are kept free for the capstone. Alongside the lectures you work through five assignments.

Tentative · last updated 2 October 2026
How the workshop is taught

Theory first, then code, then both at once

For the first four weeks the lectures come in pairs. A theory lecture on Monday explains an idea, for example how a GPU moves data from memory to its cores. On Tuesday, a coding lecture has you write a kernel that depends on that idea and measure whether it behaved as predicted. Thursday and Friday follow the same pattern. The one exception is the coding lecture on October 13, on parallelism on the CPU, which stands on its own. From November 9 the material is frontier work, where the idea and the kernel are hard to separate, so each lecture covers both: six deep-dives into the modern stack and two on kernels written by AI. Every lecture is recorded, and all code lives in the course repository.

Dr. Raj Dandekar

Dr. Raj Dandekar

Instructor

MIT PhD, founder of Vizuara. Designed and taught Vizuara's 5D Parallelism and Inference Engineering workshops.

Shubham Panchal

Shubham Panchal

Instructor

Creator of SmolChat. Works across C, C++ and Rust, deploying ML models on-device.

When

7:00 to 9:00 AM IST, every Monday, Tuesday, Thursday and Friday from October 12 to November 20, then demo day on December 7.

GPUs

No GPU of your own is needed. Sessions run on cloud GPUs, and the setup is covered in the first week. Hopper and Blackwell sessions use H100 and B200 machines.

Guest lecture

A session with the Crusoe team, in the middle or towards the end of the cohort. The date is announced inside the cohort.

Session by session

The schedule

Readings are optional before the session and useful after it. Folder names under a session refer to the course repository. Dates for later sessions may move by a day or two; changes will appear here first.

Theory ideas and derivations Coding you write the kernel Lecture from Nov 9, both together Assignment released or due Capstone milestone
DateSessionReadings and codeDeadlines
Week 1Oct 12 to Oct 16Part I · Why GPUs, and how they run your code
Mon Oct 12 Lecture 1 · Theory
How fast can this go?

Qwen3-8B on one H100 writes about 149 tokens per second, far below what the chip's arithmetic allows. We find out why with the roofline model: arithmetic intensity, memory-bound and compute-bound work. In the second hour, the class becomes the batch and we watch throughput change as more users arrive.

code · Lecture 1 practical kit
Tue Oct 13 Lecture 2 · Coding
Parallelism on the CPU

Instruction-level parallelism: pipelining, branches and out-of-order execution. Threads in the Linux kernel and the 1:1 and M:N threading models of Java, Go and Rust. SIMD vector instructions. Why graphics, and then matrix multiplication, pushed computing towards GPUs. This lecture stands on its own; it is the CPU baseline that every later lecture compares against.

code · 00_cpu_parallelism
Out: Assignment 1
Thu Oct 15 Lecture 3 · Theory
How does a GPU run your code?

We follow one image through an H100. Grids, blocks, threads and warps; streaming multiprocessors and the warp scheduler; registers, shared memory and occupancy; and why 32 threads move as one.

Fri Oct 16 Lecture 4 · Coding
Your first CUDA kernels

Vector addition and 2D max-pooling in CUDA C++. Compute capability and what nvidia-smi reports. How nvcc turns a .cu file into PTX and then SASS, and how to read both for the vecadd kernel.

code · 02_kernels_ptx_sass
Out: Assignment 2
Week 2Oct 19 to Oct 23Part I · Memory, and Part II · The GEMM worklog begins
Mon Oct 19 Lecture 5 · Theory
The memory hierarchy

Registers, shared memory, L1 and L2, and HBM. Cache lines and 32-byte sectors. Coalesced and uncoalesced loads. Shared-memory banks and bank conflicts. Why most kernels spend their time waiting for data rather than doing arithmetic.

Due: Assignment 1
Tue Oct 20 Lecture 6 · Coding
The transpose ladder, and reading Nsight Compute

A naive transpose, then a shared-memory tile that makes the writes coalesced, then padding that removes the bank conflicts. Along the way we learn Nsight Compute: sectors per request, L2 hit rate, bank conflicts, and a speed-of-light comparison of a memory-bound ReLU against a compute-bound matmul.

code · 03_matrix_transpose · 01_benchmark_pytorch_kernels
Out: Assignment 3
Thu Oct 22 Lecture 7 · Theory
GEMM worklog I

Why matrix multiplication is the operation AI runs more than any other. The naive kernel and the small fraction of the GPU it uses. Warp states and stall reasons. Shared-memory tiling and giving each thread more than one output.

Fri Oct 23 Lecture 8 · Coding
GEMM by hand I

Four kernels, from the naive version to shared-memory tiles where each thread computes a 4 by 4 block with float4 loads. Each one is profiled in Nsight Compute and timed against cuBLAS.

code · 04_gemm_A
Due: Assignment 2
Week 3Oct 26 to Oct 30Part II · GEMM and tensor cores
Mon Oct 26 Lecture 9 · Theory
GEMM worklog II

2D block tiling, register reuse, vectorised memory access, warp tiling and autotuning: the steps that take a hand-written FP32 kernel close to cuBLAS.

Tue Oct 27 Lecture 10 · Coding
GEMM by hand II

We build the 2D-tiled and warp-tiled kernels step by step and keep a worklog: one change, one measurement, one profile, then the next change.

Due: Assignment 3Out: Assignment 4
Thu Oct 29 Lecture 11 · Theory
Tensor cores

What a tensor core computes in a single instruction, and why it changes the GEMM ladder. FP16, BF16 and FP8. WMMA and mma.sync fragments, ldmatrix and shared-memory swizzling.

Fri Oct 30 Lecture 12 · Coding
A WMMA tensor-core GEMM

We rewrite our best GEMM to use tensor cores through the WMMA API, check it against cuBLAS, and look at what the profiler says has become the new bottleneck.

Week 4Nov 2 to Nov 6Part III · Profiling and attention
Mon Nov 2 Lecture 13 · Theory
Profiling and debugging like a pro

Reading an Nsight Compute report from top to bottom. Stall reasons and the source view with SASS next to your code. compute-sanitizer for races and out-of-bounds accesses. A repeatable method for finding out why a kernel is slow.

Tue Nov 3 Lecture 14 · Coding
Debug three sabotaged kernels

Three kernels that are wrong or slow on purpose. You find each problem with the profiler and the sanitizer before you are allowed to read the source.

Thu Nov 5 Lecture 15 · Theory
Attention: the kernel that ate the world

Attention as two matrix multiplications and a softmax. Why the N by N score matrix stops fitting. Safe softmax and online softmax. Tiling and IO-awareness, the idea behind FlashAttention. The lecture ends with the capstone kickoff.

Capstone kickoff
Fri Nov 6 Lecture 16 · Coding
Build FlashAttention live

Vanilla attention in PyTorch until it runs out of memory, then FlashAttention written in Triton and benchmarked from N = 2,048 to N = 65,536.

code · flash_attention
Due: Assignment 4Out: Assignment 5
Week 5Nov 9 to Nov 13Part IV · The modern frontier
Mon Nov 9 Lecture 17
FlashAttention from scratch: FA2 and FA3

We rebuild FlashAttention properly, then see what changed in FlashAttention 2 (better work partitioning across warps and blocks) and FlashAttention 3 (Hopper's asynchronous tensor cores and FP8).

code · flash_attention
Tue Nov 10 Lecture 18
Beating cuBLAS on an H100

What Hopper added and why it matters for GEMM: the Tensor Memory Accelerator (TMA), wgmma, thread-block clusters, asynchronous pipelines and warp specialisation. We use them to build a matrix multiply that beats NVIDIA's own library.

Thu Nov 12 Lecture 19
Triton, CUTLASS and CuTe-DSL

The abstraction ladder above raw CUDA. Triton's blocks, CUTLASS templates, CuTe layouts, and CuTe-DSL for writing CUTLASS-class kernels in Python. We write one fused kernel in each and compare the code and the speed.

Due: Assignment 5
Fri Nov 13 Lecture 20
Inference-serving kernels

Prefill and decode behave like two different workloads. The KV cache and PagedAttention. Fused RMSNorm and rotary embeddings. Quantised GEMMs. What speculative decoding asks of the kernels underneath it.

Capstone proposal due
Week 6Nov 16 to Nov 20Part IV · Blackwell, and Part V · AI-written kernels
Mon Nov 16 Lecture 21
Blackwell and NVFP4

What changed from Hopper to Blackwell, and the new tensor-core instructions that come with it. FP4 numbers with block scaling, and what NVFP4 does to throughput and to accuracy. We run an NVFP4 GEMM on a B200.

Tue Nov 17 Lecture 22
Flash Attention 4

From FlashAttention 3 on Hopper to FlashAttention 4 on Blackwell. Why, on the newest chips, the exponentials in softmax rather than the matrix multiplications become the limit, and the tricks FA4 uses to work around it.

code · flash_attention/papers
Thu Nov 19 Lecture 23
DeepSeek: FlashMLA and DeepGEMM

Multi-head latent attention and why it needs its own decode kernel. FP8 GEMMs with fine-grained scaling for mixture-of-experts models. What a small, focused kernel library can do against large general ones.

Fri Nov 20 Lecture 24
Kernels written by AI: the agent and profiler loop

KernelBench and its fast_p metric. An LLM proposes a kernel, a harness checks correctness and times it, and the profiler output goes back into the next prompt. How to tell a generated kernel that is actually faster from one that only looks fast, and where the kernel engineer still decides.

Capstone checkpoint 1: baseline
Week 7 and 8Nov 23 to Dec 4Part VI · Capstone build
Fri Nov 27 Capstone
Capstone checkpoint 2

No lectures in these two weeks. You build your capstone. By this date, share a first optimised version with a profile that explains the speed-up.

Capstone checkpoint 2
Sun Dec 6 Capstone
Capstone code and worklog due

Final code, correctness tests, benchmark harness and the worklog.

Final submission
Week 9Dec 7Part VI · Capstone
Mon Dec 7 Lecture 25 · Capstone
Capstone demo day

Every team presents its kernel, its baseline, its measured speed-up and its worklog. Crusoe joins the review.

Demo day
Five assignments

Assignments

Each assignment follows the lectures it depends on. Most are released on the day of a coding lecture and are due about a week later; the GEMM worklog gets ten days because it is the longest. After Assignment 5, the capstone takes over.

Assignment 1out Tue Oct 13 · due Mon Oct 19

Roofline by hand, and the CPU baseline

Compute the arithmetic intensity of five operations (vector add, ReLU, a softmax row, a matrix-vector product and a 4096 by 4096 GEMM) and place each on the H100 roofline. Then parallelise a CPU loop with threads and with SIMD and measure how it scales with core count.

You hand in: A short report with your numbers, your plot and one paragraph on what surprised you.

Assignment 2out Fri Oct 16 · due Fri Oct 23

First kernels

Write SAXPY, an RGB-to-grayscale kernel and a 2D max-pool in CUDA. Sweep the block size from 32 to 1,024 threads and plot the runtime. Dump the PTX and SASS of one kernel and mark where the loads, the arithmetic and the stores happen. Work out the occupancy for three register counts.

You hand in: Code, the plot and the annotated SASS.

Assignment 3out Tue Oct 20 · due Tue Oct 27

Memory and measurement

Get a matrix transpose to at least 80% of the bandwidth of a plain device-to-device copy and show with Nsight Compute why it is fast. Then take six PyTorch operations, predict from the roofline whether each is memory-bound or compute-bound, and check your prediction with the profiler.

You hand in: Code, profiles and a table of predictions against measurements.

Assignment 4out Tue Oct 27 · due Fri Nov 6

Your own GEMM worklog

Take an FP32 GEMM from the naive kernel to at least 70% of cuBLAS at 4096 by 4096 by 4096, and add one tensor-core version with WMMA. Every step in the worklog needs a measurement and a profiler screenshot that explains it.

You hand in: Code and a worklog in the style of the GEMM lectures.

Assignment 5out Fri Nov 6 · due Thu Nov 12

FlashAttention forward pass

Implement a causal FlashAttention forward pass in Triton. Match PyTorch's scaled_dot_product_attention within BF16 tolerance, benchmark it from N = 1,024 to N = 32,768, and report peak memory against vanilla attention. The backward pass is an optional extension.

You hand in: Code, a correctness test and the benchmark table.

Part VI · With Crusoe

The capstone project

The capstone is a real kernel problem solved from start to finish: a baseline you can trust, a kernel that beats it, and a worklog that explains every step. Crusoe shares project ideas at the start of the cohort, so the problems match what an AI-infrastructure company actually works on. The ideas below are examples of the scale we have in mind.

Milestones
  • Mon Oct 12Project ideas published, including problems from Crusoe.
  • Thu Nov 5Capstone kickoff at the end of the attention lecture. Work solo or in pairs.
  • Fri Nov 13Proposal due: one page with the problem, the baseline, the target metric and the GPU you need.
  • Fri Nov 20Checkpoint 1: a correctness test, a benchmark harness and a measured baseline.
  • Nov 23 to Dec 4Two weeks with no lectures, kept free for the capstone build.
  • Fri Nov 27Checkpoint 2: a first optimised version, with a profile that explains the speed-up.
  • Sun Dec 6Final code and worklog due.
  • Mon Dec 7Demo day, Lecture 25.
Example projects

A decode attention kernel for a paged KV cache

Beat the vLLM kernel on one model's shapes, at batch sizes that matter in serving.

A fused RMSNorm, rotary and quantise kernel

Replace three launches in a real transformer block with one, and measure the end-to-end effect.

FP8 or NVFP4 GEMMs for one model's layer shapes

Tune for the exact matrix shapes of an open model such as Qwen3-8B on H100 or B200.

A grouped GEMM for mixture-of-experts routing

Handle uneven expert loads without padding every expert to the largest one.

You against the machine

Optimise the same kernel by hand and with an agent loop, and write up what each found and missed.

Speed up a served model

Profile a vLLM deployment on Modal, find the slowest kernel, replace it, and show the change in tokens per second.

How a capstone is judged