EulerFold
HPC & ML Infrastructure Engineers

GPU Programming with CUDA

6 weeks
0 Learners
Jul 23

Direct GPU control for custom kernels: essential when PyTorch abstractions aren't enough. This comprehensive curriculum is designed to give you hands-on experience with GPU Programming with CUDA.

Share:

About this Course

Direct GPU control for custom kernels: essential when PyTorch abstractions aren't enough. This comprehensive curriculum is designed to give you hands-on experience with GPU Programming with CUDA. This HPC & ML Infrastructure Engineers curriculum is designed to give you hands-on experience and deep conceptual understanding. Across 6 intensive modules, you'll tackle real-world challenges and build practical projects that reinforce your learning. By the end of this journey, you'll have the skills and proof of work to demonstrate your expertise.

What you'll learn

Master the core concepts of cuda architecture & execution model.
Gain hands-on experience with cuda memory hierarchy: shared, global & constant.
Understand the architecture behind warp divergence, parallel reduction & atomic operations.
Implement production-grade building custom pytorch c++/cuda extensions.

Prerequisites

intermediate Level

Requires basic familiarity with the tech stack.

  • Familiarity with core concepts

Ideal for

HPC & ML Infrastructure Engineers

HPC & ML Infrastructure Engineers Professionals
Tech Enthusiasts
W1

CUDA Architecture & Execution Model

Master the core concepts of cuda architecture & execution model.

3 videos25m
2 readings
3 topics
1 homework
Learn

Topics

1.1
GPU Hardware Architecture & Streaming Multiprocessors
8 minutes
1.2
Thread Hierarchy & Grid Launch Configuration
9 minutes
1.3
Memory Allocation & Data Transfers
8 minutes
W2

CUDA Memory Hierarchy: Shared, Global & Constant

Gain hands-on experience with cuda memory hierarchy: shared, global & constant.

3 videos56m
3 readings
3 topics
1 homework
Learn
W3

Warp Divergence, Parallel Reduction & Atomic Operations

Understand the architecture behind warp divergence, parallel reduction & atomic operations.

3 videos29m
3 readings
3 topics
1 homework
Learn
W4

Building Custom PyTorch C++/CUDA Extensions

Implement production-grade building custom pytorch c++/cuda extensions.

3 videos68m
3 readings
3 topics
1 homework
Learn
W5

Shared Memory and Synchronization

Optimize CUDA kernels by utilizing shared memory and managing thread synchronization.

3 videos40m
3 readings
3 topics
1 homework
Learn
W6

Streams, Warps, and Occupancy

Maximize GPU utilization using concurrent streams and understanding warp execution.

3 videos58m
3 readings
3 topics
1 homework
Learn
01

Learn

Watch curated videos and read study resources

02

Practice

Practice what you learned

03

Build Projects

Build projects using your new gained knowledge

04

Submit & Verify

Submit your project and get verified by our system

Rate this course

0.0
0 reviews

Help the community find verified technical paths.

Community Insights

0

Join the discussion

Sign in to share your thoughts and technical insights.

Loading insights...