HomeCourses › Multimodal AI Systems

Multimodal AI Systems

From CLIP to GPT-4V. Build models that see, hear and read

5 phases. 25 lessons. 25 labs. 1 capstone. Multimodal AI from vision-language models to end-to-end systems — multimodal foundations (CLIP, BLIP, image-text alignment), vision-language models (LLaVA, GPT-4V architecture, visual instruction tuning), multimodal generation (text-to-image, image editing, audio-visual generation), and production multimodal systems (RAG with images, multimodal agents, evaluation & deployment). You build a multimodal AI system that processes text, images, and audio together.

5 phases · 25 lessons · 25 labs · 1 capstone

Take Transformers & LLMs from Scratch first — this course builds on it.

Outcomes you will have by the end

What you will be able to do

CLIP & Contrastive Image-Text Learning · BLIP & BLIP-2 for Vision-Language · Cross-Modal Retrieval · LLaVA Architecture (Vision Encoder + LLM) · Visual Instruction Tuning · GPT-4V-Style Multimodal Reasoning · Multimodal Generation (Text-to-Image, Image Editing) · Multimodal RAG · Multimodal Agents · Multimodal System Evaluation

Every phase, every lesson, every project

The technologies you will use

PyTorch · Hugging Face Transformers · CLIP · LLaVA · Diffusers

Roles this course prepares you for

What multimodal AI systems actually is

It's building AI that processes multiple modalities — text, images, audio — together. From CLIP that aligns vision and language to LLaVA that connects vision encoders to LLMs to multimodal agents that use tools across modalities, you build the systems that represent the frontier of AI.

What you do every day

You build multimodal AI systems — connecting vision encoders to LLMs, building cross-modal retrieval, and creating agents that reason across modalities. When a VLM gives wrong answers about an image, you know whether it's the vision encoder, the projection layer, or the LLM reasoning.

Why companies hire for this

Multimodal AI is the frontier — GPT-4V, Gemini, Claude with vision. Companies building AI products need engineers who understand how to combine modalities, not just use single-modality models. The engineer who built LLaVA from scratch is the one who can customize VLMs for specific domains and build production multimodal systems.

What this course is not

It is not a tutorial on using GPT-4V's API. It is not about prompt engineering for vision models. It is about building the multimodal systems themselves — CLIP, LLaVA, multimodal RAG, multimodal agents — from scratch. If you want to understand how AI combines seeing, hearing, and reading, this is the course.

Common questions

Do I need Transformers & LLMs from Scratch first?

Yes — this course uses transformers, LLMs, and attention mechanisms throughout. Transformers & LLMs from Scratch builds those foundations; this course extends them to multimodal settings.

Should I take Computer Vision Engineering and NLP & Speech Processing first?

Not required, but helpful. This course focuses on the intersection of modalities — how vision and language are aligned and combined. If you want deep expertise in a single modality, take the specialized courses first. If you want to understand how modalities work together, this course is self-contained with the transformers prerequisite.

How is this different from the multimodal content in other courses?

Computer Vision Engineering covers CLIP and ViTs from a vision perspective. NLP & Speech Processing covers text and audio from a language perspective. This course focuses on the intersection — how vision encoders connect to LLMs (LLaVA), how multimodal RAG works, and how to build production systems that process multiple modalities together.

What do I end up with?

A multimodal AI system that processes text, images, and audio together — with cross-modal retrieval, vision-language reasoning, and multimodal generation. Plus the ability to architect any multimodal AI system from research papers.

Key terms in this course

RAG (Retrieval-Augmented Generation)

Continue your learning path

Transformers & LLMs from Scratch · Computer Vision Engineering · NLP & Speech Processing · Generative AI Fundamentals · ML & AI Engineering

Start the Multimodal AI Systems course

Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.

All courses · Pricing · About · FAQ · Glossary