Computer Vision

Master the full computer vision pipeline: image processing, CNNs, object detection, segmentation, vision transformers, and CLIP. For engineers who want to build production vision systems, not just use pre-trained models.

The route

  1. CNNs from ScratchDeep Learning from Scratch (phase dl-01). Implement convolutions, pooling, and backprop through convolution layers. Build a CNN that classifies images — understanding exactly what every layer does.
  2. Computer Vision Engineering (full course)Computer Vision Engineering. 4 phases, 16 lessons: image processing, CNNs, object detection (YOLO, DETR), segmentation (U-Net, Mask R-CNN), and modern vision (ViT, CLIP, DINOv2, diffusion). Build a medical imaging classifier as the capstone.
  3. CLIP & Multimodal AlignmentMultimodal AI Systems (phase mm-00). Implement CLIP from scratch and understand contrastive learning — the foundation of modern multimodal AI that aligns vision and language.

What you build

A medical imaging classifier, an object detection system, and a CLIP-based cross-modal retrieval system — the full computer vision toolkit.

Every learning path

Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.

All courses · Pricing · About · FAQ · Glossary