How do robots and self-driving cars actually see? Computer vision for physical AI, explained simply with animations — from how a camera turns the world into numbers to how a sidewalk robot detects, tracks, and measures depth. Perfect as a concept refresher or as a starting point to dive deeper.
What you'll learn:
• How images, cameras & stereo depth work
• Perceptrons → CNNs → ResNets → Vision Transformers
• Object detection (YOLO, Faster R-CNN, DETR), IoU, NMS, mAP
• Segmentation (U-Net) and multi-object tracking
• 3D perception: LiDAR, point clouds, bird's-eye view, sensor fusion, SLAM
• Foundation models: self-supervised learning, MAE, CLIP, SAM
• Shipping it: data flywheels, latency, quantization, drift
KEY PAPERS FOR FURTHER READING
ResNet: https://arxiv.org/abs/1512.03385
Vision Transformer: https://arxiv.org/abs/2010.11929
Faster R-CNN: https://arxiv.org/abs/1506.01497
YOLO: https://arxiv.org/abs/1506.02640
Focal loss / RetinaNet: https://arxiv.org/abs/1708.02002
DETR: https://arxiv.org/abs/2005.12872
U-Net: https://arxiv.org/abs/1505.04597
PointNet: https://arxiv.org/abs/1612.00593
BEVFusion: https://arxiv.org/abs/2205.13542
MAE: https://arxiv.org/abs/2111.06377
CLIP: https://arxiv.org/abs/2103.00020
Segment Anything: https://arxiv.org/abs/2304.02643
Learn more: Stanford CS231n (cs231n.stanford.edu) · Szeliski's free CV book (szeliski.org/Book)
How do robots and self-driving cars actually see? Computer vision for physical AI, explained simply with animations — from how a camera turns the world into numbers to how a sidewalk robot detects, tracks, and measures depth. Perfect as a concept refresher or as a starting point to dive deeper.
What you'll learn:
• How images, cameras & stereo depth work
• Perceptrons → CNNs → ResNets → Vision Transformers
• Object detection (YOLO, Faster R-CNN, DETR), IoU, NMS, mAP
• Segmentation (U-Net) and multi-object tracking
• 3D perception: LiDAR, point clouds, bird's-eye view, sensor fusion, SLAM
• Foundation models: self-supervised learning, MAE, CLIP, SAM
• Shipping it: data flywheels, latency, quantization, drift
KEY PAPERS FOR FURTHER READING
ResNet: https://arxiv.org/abs/1512.03385
Vision Transformer: https://arxiv.org/abs/2010.11929
Faster R-CNN: https://arxiv.org/abs/1506.01497
YOLO: https://arxiv.org/abs/1506.02640
Focal loss / RetinaNet: https://arxiv.org/abs/1708.02002
DETR: https://arxiv.org/abs/2005.12872
U-Net: https://arxiv.org/abs/1505.04597
PointNet: https://arxiv.org/abs/1612.00593
BEVFusion: https://arxiv.org/abs/2205.13542
MAE: https://arxiv.org/abs/2111.06377
CLIP: https://arxiv.org/abs/2103.00020
Segment Anything: https://arxiv.org/abs/2304.02643
Learn more: Stanford CS231n (cs231n.stanford.edu) · Szeliski's free CV book (szeliski.org/Book)