Computer Vision
Master the full computer vision pipeline: image processing, CNNs, object detection, segmentation, vision transformers, and CLIP. For engineers who want to build production vision systems, not just use pre-trained models.
The route
- CNNs from Scratch — Deep Learning from Scratch (phase dl-01). Implement convolutions, pooling, and backprop through convolution layers. Build a CNN that classifies images — understanding exactly what every layer does.
- Computer Vision Engineering (full course) — Computer Vision Engineering. 4 phases, 16 lessons: image processing, CNNs, object detection (YOLO, DETR), segmentation (U-Net, Mask R-CNN), and modern vision (ViT, CLIP, DINOv2, diffusion). Build a medical imaging classifier as the capstone.
- CLIP & Multimodal Alignment — Multimodal AI Systems (phase mm-00). Implement CLIP from scratch and understand contrastive learning — the foundation of modern multimodal AI that aligns vision and language.
What you build
A medical imaging classifier, an object detection system, and a CLIP-based cross-modal retrieval system — the full computer vision toolkit.
Every learning path
Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.
All courses · Pricing · About · FAQ · Glossary