Fine-tuning Vision-Language Models

Adapt multimodal models like LLaVA and PaliGemma for domain-specific visual understanding, document parsing, and image captioning.

Created Byeulerfoldeulerfold
4 weeks
Jul 23
to start learning
Curriculum

Share:

About this Course

Adapt multimodal models like LLaVA and PaliGemma for domain-specific visual understanding, document parsing, and image captioning. This AI/ML & Computer Vision curriculum is designed to give you hands-on experience and deep conceptual understanding. Across 4 intensive modules, you'll tackle real-world challenges and build practical projects that reinforce your learning. By the end of this journey, you'll have the skills and proof of work to demonstrate your expertise.

What you'll learn

Master the core concepts of vision-language model architectures.
Gain hands-on experience with dataset preparation & augmentation for vlm.
Understand the architecture behind parameter-efficient fine-tuning (peft/lora).
Implement production-grade evaluation & deployment of custom vlms.

Prerequisites

intermediate Level

Requires basic familiarity with the tech stack.

  • Python scripting
  • Basic linear algebra & tensors

Ideal for

Machine Learning Engineers

AI/ML & Computer Vision Professionals
Tech Enthusiasts
W1

Vision-Language Model Architectures

Master the core concepts of vision-language model architectures.

4 videos51m
3 readings
4 topics
1 homework
Learn

Topics

1.1
Image Encoders (ViT, SigLIP) & Projection Layers
How AI 'Understands' Images (CLIP) - Computerphile
18 minutes
1.2
Multimodal Fusion Strategies
Vision-Language Models Explained | CLIP, DALL·E, Florence & Multimodal AI
8 minutes
1.3
LLaVA vs PaliGemma Architectures
Vision-Language Models (VLMs) Explained | GPT-4V, LLaVA & CLIP
8 minutes
1.4
Tokenizing Images and Text
Stop Using GPUs for LLMs : Fine-Tune Gemma 3 with JAX on a TPU (Fast & Cheap)
17 minutes
W2

Dataset Preparation & Augmentation for VLM

Gain hands-on experience with dataset preparation & augmentation for vlm.

4 videos76m
3 readings
4 topics
1 homework
Learn
W3

Parameter-Efficient Fine-Tuning (PEFT/LoRA)

Understand the architecture behind parameter-efficient fine-tuning (peft/lora).

4 videos124m
3 readings
4 topics
1 homework
Learn
W4

Evaluation & Deployment of Custom VLMs

Implement production-grade evaluation & deployment of custom vlms.

4 videos98m
3 readings
4 topics
1 homework
Learn
Rate this course
0.0
0 reviews

Community Insights

0

Join the discussion

Sign in to share your thoughts and technical insights.

Loading insights...