Computer Vision Engineering
From convolutions to ViTs — build models that see
4 phases. 16 lessons. 16 labs. 1 capstone. Computer vision from convolutions to cutting-edge — CNNs from scratch (convolution, pooling, architectures), object detection (YOLO, DETR, Faster R-CNN), segmentation (U-Net, Mask R-CNN, SAM), and modern vision (Vision Transformers, CLIP, DINOv2, diffusion for vision). You build a vision model that detects, segments, and classifies objects in real images.
- Lessons: —
- Labs: —
- Projects: —
- Level: Beginner
Curriculum
- Vision Foundations — Image processing, convolutions, and edge detection from scratch.
- Convolutional Neural Networks — CNNs, pooling, and classic architectures from scratch.
- Object Detection & Segmentation — YOLO, R-CNN, and segmentation from scratch.
- Generative Vision Models — Autoencoders, GANs, and diffusion models for images.
- Vision Transformers — ViTs, CLIP, and multimodal vision models.
- CV Deployment & Applications — Model optimization, deployment, and real-world CV applications.
Skills You Will Learn
- 2D Convolutions from Scratch
- CNN Architectures (LeNet to ResNet)
- Transfer Learning for Vision
- Object Detection (YOLO, Faster R-CNN, DETR)
- Semantic & Instance Segmentation (U-Net, Mask R-CNN, SAM)
- Vision Transformers (ViT)
- CLIP & Image-Text Matching
- Diffusion Models for Image Generation
Related Courses
Browse all courses · View pricing · DeVenture Academy