Vizuara Kernel Engineering
Cohort begins Oct 12, 2026

partnerThe most modern GPU kernel-engineering program anywhere, built over two years, taught live, and explained from the ground up, from the silicon up to Flash Attention 4, Blackwell, and kernels written by AI. If you want to actually be ready for GPU performance-engineering roles at the frontier, this is it.
The third and final chapter of Vizuara's GPU trilogy, after 5D Parallelism and the Inference Engineering Workshop.
Most GPU courses stop at 2022, and the few modern ones are written for people who already know everything. This one lives on the 2025-2026 frontier but explains each idea from first principles, with, for every topic, why it matters, where it's used, and which companies use it.
Every session opens with the idea in simple words before the jargon, you'll leave knowing what the terms actually mean.
Flash Attention 4, Blackwell, NVFP4, DeepSeek, AI-written kernels, the newest material, not a 2022 rerun.
Every lecture is paired with a live-coding session. You write the kernels yourself, step by step.
Maps 1:1 to what frontier companies like Wafer, NVIDIA and the top labs actually hire for.
A single, deliberate arc. Each concept lecture (Dr. Raj Dandekar) is paired with a live-coding session (Shubham Panchal) so theory and practice go hand in hand. Click any session, each one explains the idea simply, plus why it matters, where it's used, and who uses it.
Get comfortable with "doing many things at once" on a CPU first, then meet the GPU and how you program it, so every kernel that follows sits on solid foundations.
Take the single most important operation in AI, matrix multiply, from a version that uses ~1% of the GPU to one that beats NVIDIA's own library, one improvement at a time.
Learn to find why a kernel is slow like a pro, then build the operation at the heart of every LLM, FlashAttention, and kick off your capstone.
The 2025-2026 kernels the best labs ship right now: Hopper, Blackwell, DeepSeek, and Flash Attention 4, each one explained from the ground up, not assumed.
The newest twist of all: tiny hand-crafted kernels that beat huge libraries, and AI models that now write GPU kernels themselves, and how a kernel engineer guides them.
A real kernel-engineering project, built on capstone project ideas from Crusoe and graded at demo day, the closest thing to an on-site interview before the on-site interview.
Every session is live, concepts built from first principles, then written into working kernels in front of you.
Has designed and taught Vizuara's 5D Parallelism and Inference Engineering workshops, the first two chapters of this GPU trilogy, and is known for taking hard systems and ML material and making it click from first principles. He teaches the concept lectures.
An Android developer since 2017, Shubham specializes in deploying ML models on-device in Android apps, the creator of SmolChat (run LLMs locally, on-device). An active StackOverflow contributor and frequent Medium writer on ML, math and Android, he now works across Rust, C and C++ alongside Java/Kotlin backends. He leads the live-coding, where you write every kernel with him.
This cohort runs in partnership with Crusoe, the AI-infrastructure company. The partnership is built into the workshop itself.
A session with the Crusoe team during the cohort, scheduled around the middle or the end of the nine weeks.
Crusoe shares capstone project ideas as the cohort starts, so your final project tracks problems an AI-infrastructure company actually cares about.
One price, the whole system. The live cohort is the centre, and around it sits everything we've built for kernel engineers, the book, the capstone projects, the interview-preparation platform, and a bonus we think you'll love.
Every concept lecture paired with a live-coding session, Mon · Wed · Fri, 7:00–9:00 AM, October 12 → December 7.
Full access to our illustrated Kernel Engineering book, the written companion to every session. Unlocked when you register.
A real kernel-engineering project, with project ideas from Crusoe, solved end to end and graded at demo day.
Access to the Kernel Engineering Interview Preparation website, practice for the exact roles this workshop gears you for.
Every kernel you build in the cohort, plus session recordings, yours to keep.
One month of access to Vizz-AI and Vizuara AI Pods, an additional bonus on top of the workshop.
GPU performance engineering is one of the most sought-after, hardest-to-fill skills in AI. This isn't a toy syllabus, it maps section-for-section onto what companies like Wafer, NVIDIA, and the top labs actually hire for.
The core screen for every GPU-perf role, tiling, vectorization, tensor cores, reasoning about speed byte-for-byte.
FlashAttention 1-4, PagedAttention, KV-cache, speculative decoding, the exact stack in frontier inference roles.
The modern kernel-authoring layer, warp-specialized GEMM, and epilogue fusion.
Nsight/NCU, reading the compiled assembly, and the agent + profiler loop that frontier GPU startups productionize.
These roles are among the highest-paid, most in-demand engineering jobs in the industry, the kind where a single week's pay dwarfs the cost of this workshop. You leave able to point at each line of a job description and say: built that, graded on it, wrote the chapter.
Nobody else sells a single live, graded, book-backed path through all of this. Against what these skills unlock, $3,000 is a fraction of what the job it prepares you for pays in a single week.
Seats are limited; the cohort is kept small enough for live review of every participant's kernels. Questions before enrolling? Email team@vizuara.com.
The cohort begins Monday, October 12, 2026 and finishes on Monday, December 7, 2026, 25 sessions of two hours each. Every week runs Monday, Wednesday and Friday, 7:00–9:00 AM, and every session is recorded.
Capstone demo day follows the final session; the exact date is announced inside the cohort.
One plan, $3,000, everything included, and the launch price holds only until August 15, 2026.