EulerFold
AI Infrastructure

LLM Inference Internals: KV Cache, PagedAttention, and Continuous Batching

7 weeks
0 Learners
Jul 23

Deep dive into making LLMs fast. Understand the KV cache, memory management, and scaling inference. This comprehensive curriculum is designed to give you hands-on experience with LLM Inference Internals: KV Cache, PagedAttention, and Continuous Batching.

Share:

About this Course

Deep dive into making LLMs fast. Understand the KV cache, memory management, and scaling inference. This comprehensive curriculum is designed to give you hands-on experience with LLM Inference Internals: KV Cache, PagedAttention, and Continuous Batching. This AI Infrastructure curriculum is designed to give you hands-on experience and deep conceptual understanding. Across 7 intensive modules, you'll tackle real-world challenges and build practical projects that reinforce your learning. By the end of this journey, you'll have the skills and proof of work to demonstrate your expertise.

What you'll learn

Understand the bottleneck in autoregressive decoding and the role of the KV cache.
Learn how vLLM optimizes memory using operating system paging concepts.
Explore methods to maximize GPU utilization across multiple requests.
Learn techniques to run models in lower precision without significantly losing accuracy.

Prerequisites

advanced Level

This is a rigorous course requiring prior technical background.

  • Python scripting
  • Basic linear algebra & tensors

Ideal for

Anyone interested in AI Infrastructure who wants to learn about LLM Inference Internals: KV Cache, PagedAttention, and Continuous Batching.

AI/ML Engineers
W1

Transformer Inference Fundamentals

Understand the bottleneck in autoregressive decoding and the role of the KV cache.

3 videos74m
3 readings
3 topics
1 homework
Learn
W2

Memory Management & PagedAttention

Learn how vLLM optimizes memory using operating system paging concepts.

3 videos51m
3 readings
3 topics
1 homework
Learn
W3

Batching Strategies

Explore methods to maximize GPU utilization across multiple requests.

3 videos58m
3 readings
3 topics
1 homework
Learn
W4

Quantization and Low Precision Inference

Learn techniques to run models in lower precision without significantly losing accuracy.

3 videos60m
3 readings
3 topics
1 homework
Learn
W5

Serving at Scale

Deploy optimized models and benchmark their performance.

3 videos39m
3 readings
3 topics
1 homework
Learn
W6

PagedAttention & Memory Management

Implement vLLM-style PagedAttention to eliminate memory fragmentation in the KV cache.

3 videos108m
3 readings
3 topics
1 homework
Learn
W7

Continuous Batching & Serving

Build a continuous batching scheduler to maximize GPU utilization across concurrent requests.

3 videos50m
3 readings
3 topics
1 homework
Learn
01

Learn

Watch curated videos and read study resources

02

Practice

Practice what you learned

03

Build Projects

Build projects using your new gained knowledge

04

Submit & Verify

Submit your project and get verified by our system

Rate this course

0.0
0 reviews

Help the community find verified technical paths.

Community Insights

0

Join the discussion

Sign in to share your thoughts and technical insights.

Loading insights...